money
$500
hour 0
revenue / hr
$0
at $0.28 / 1M tok
costs / hr
$0
rental ยท power ยท refunds
profit / hr
$0
margin 0%
cost per 1M tokens
โ€“
you charge $0.28
power
0 W
tokens per joule โ€“
GPU utilization
0%
seats busy / seats
HBM bandwidth
0%
of peak, modelled
compute eff.
0%
FLOPs / peak, modelled
KV cache in use
0%
0 sequences ยท 0 tokens
tokens / s
0
system throughput
per query
0
tok/s each ยท wait โ€“ ยท prefill โ€“
time to first token
โ€“
p50, last 30 s ยท wait + prefill
waiting
0
arriving 0/s ยท SLA 20 s ยท 0 timeouts

YOUR LEVERS GPT-OSS-120B

$0.28
demand vs capacity
what if: hover RENT

INCOMING 0 waiting

SERVINGclock: seconds MODELclock: milliseconds โ€” ร—10ยณ faster KERNELSclock: microseconds SILICONclock: nanoseconds โ€” ร—10ยณ again FLEETclock: days โ€” the slow loop that owns the fast ones BEDROCK โ€” published numbers: 1,611 default ยท 2,418 tuned ยท 2,919 with three improvements (GPT-OSS-120B, one MI300X, vLLM) INTAKE SCHEDULER KV ALLOC fused bookkeeping โœ“ KV-fusion BATCH B(14) C(30) max batch: 256 max-batch embed attention router experts combine sample โŸฒ AUTOREGRESSIVE STEP โ€” the output token becomes the next input โ†ฉ ร—k tokens / lap ร—36 EMBED ATTENTION ROUTER EXPERTS COMBINE SAMPLER draft ร—k โœ“ EAGLE-3 CARTRIDGE: GPT-OSS-120B โ€” the model this machine serves DISPATCH ENV=โ€ฆ ร—7 torch.compile env-flags ATTN KERNEL AITER GEMM your 3 hand entries launch drawer: 437 shapes masked tail tile tables FUSED OPS fused top-k CU ARRAY MM@2: L1 experts parked โœ“ MM@2 ATTN โ—‚ โ–ธ MOE IF LINKS 8 replicas โ€” each GPU serves its own requests; one scheduler fans them out fleet ร—8 A2A โ–ธ to the experts' hosts, and home ร—8 โณ barrier: waiting for GPU 3 OP LAUNCH TILES CORES requests ยท s 0 tokens ยท ms 0 launches ยท ยตs 0 FLOPs ยท ns โ€” a blur 0 discoveries ยท days 0
no bottleneck right now
TIME ร—1.0

Mini-game

Placeholder: the mini-game for this improvement is being built separately. Playing it is how the improvement gets earned.

Machine down

A new model just dropped.

Kimi K3, 2.8 trillion parameters, on new hardware (a node of eight MI355X) with a new inference engine (TokenSpeed). Customers want it, at about 29x the price per token of GPT-OSS.

Everything you tuned was tuned for GPT-OSS on MI300X under vLLM. None of it carries: the kernels, the cache policy, the tile tables are specific to that model, chip and engine. Throughput per node starts far lower (the published number is 161.7 tok/s at sixteen concurrent streams), and MakerMaker has to find the bottlenecks again.

What carries over: your money and the floors you have unlocked. What resets: the improvements and the demand, which starts small and grows. The first node comes with the launch program (no rental, you pay power); more nodes rent at $20/hr and cannot pay for themselves at these prices, so the price slider is your capacity lever.

Rent another GPU?

Kimi K3 is tuned.

The story ends here for now. Demand keeps rising if you keep serving.

Oh wow, nice GPU you got there. Would be a shame if they were not utilized properly.

Play thru this simulator to learn all the real improvements that the MakerMaker improvement engine found to improve inference serving. Your starting levers: the price you charge and how many GPUs you rent. You start with one GPU.

You die when too many of your requests time out, or you run out of money.