Placeholder: the mini-game for this improvement is being built separately. Playing it is how the improvement gets earned.
Kimi K3, 2.8 trillion parameters, on new hardware (a node of eight MI355X) with a new inference engine (TokenSpeed). Customers want it, at about 29x the price per token of GPT-OSS.
Everything you tuned was tuned for GPT-OSS on MI300X under vLLM. None of it carries: the kernels, the cache policy, the tile tables are specific to that model, chip and engine. Throughput per node starts far lower (the published number is 161.7 tok/s at sixteen concurrent streams), and MakerMaker has to find the bottlenecks again.
What carries over: your money and the floors you have unlocked. What resets: the improvements and the demand, which starts small and grows. The first node comes with the launch program (no rental, you pay power); more nodes rent at $20/hr and cannot pay for themselves at these prices, so the price slider is your capacity lever.
The story ends here for now. Demand keeps rising if you keep serving.
Play thru this simulator to learn all the real improvements that the MakerMaker improvement engine found to improve inference serving. Your starting levers: the price you charge and how many GPUs you rent. You start with one GPU.
You die when too many of your requests time out, or you run out of money.