About

About

I’m Qian Cheng — I work on large-model training and inference systems.

Most of what I write here comes from the floor of real jobs: why a run got slower after we stretched context, which attention or MoE tweak actually moved loss vs. wall-clock, how to estimate memory before you blow a GPU, and where speculative decoding, distillation, or RL training quietly fall over. You’ll also find notes on agents, kernels (Triton / fused ops), serving stacks (vLLM and friends), and the homelab tooling I use to keep the rest of the workflow sane.

I prefer short feedback loops: measure first, keep the negative results, and write down the decision — not only the win.

If something here is useful (or wrong), find me on GitHub or drop a line at im.qiancheng@gmail.com.