The wall
Batch-1 decode re-reads every weight, every token. tok/s ≈ bandwidth / (params × bytes). If you are not on that line, the kernel is the bug, not the model.
Decode is a memory problem. A guild for GPU performance engineering of inference: the math, the kernels, the quants, and the recipes that make a named card actually fast. Not another engine.
Batch-1 decode re-reads every weight, every token. tok/s ≈ bandwidth / (params × bytes). If you are not on that line, the kernel is the bug, not the model.
Flags, screenshots, and unnamed GPUs. NVFP4, EXL3, GGUF, MTP, DFlash: real tools, written down as threads. A 3090 recipe and a Spark recipe are different objects.
The org that owns the math and the named-card notebook will outlast any single runtime, the way a roofline outlasts a compiler flag.
The algebra
Decode, prefill, KV, roofline, speculation, occupancy. Worked as if a 3090 were on the desk.
Decode wall
tok/s ≈ B / (N · q)
B = HBM bandwidth, N = parameters, q = bytes per weight
Autoregressive decode re-reads the weights every token. On a 3090 (936 GB/s) a 27B model at 4-bit is ~87 tok/s if you actually hit the bus. If you measure less, you are leaving bandwidth on the table.
Prefill
FLOPs ≈ 2 · N · S
S = sequence length; tok/s ≈ peak FLOP/s / FLOPs per token
Prefill is compute-bound until the sequence is short or the quant is extreme. Tensor Cores matter here. They barely matter at batch-1 decode.
KV cache
bytes = 2 · L · n_kv · d_h · S · q_kv
L layers, n_kv heads, d_h head dim, S sequence, q_kv bytes per element
Long context is a RAM tax, not a FLOP tax. GQA, quantized KV, and a smaller n_kv buy sequence. They do not buy a smarter model.
Roofline
I = FLOPs / bytes; achieved = min(peak_FLOP, I · B)
I = arithmetic intensity. Ridge point = peak_FLOP / B
If I is below the ridge, more Tensor Cores do nothing. Fuse, tile, or raise batch until I climbs, or buy bandwidth.
Speculation
E[tokens] = (1 − α^{γ+1}) / (1 − α)
α = draft accept rate, γ = draft tokens per step
A fast draft on a slow target wins. A slow draft on a fast 7B is a tax. Measure α on your prompt, not on HumanEval.
Occupancy
occ = active_warps / max_warps
Limited by registers, shared memory, and the block size you picked
nvidia-smi at 100% can still be stalled on the scoreboard. Nsight is the instrument. Occupancy is the first number it should teach you.
The stack
Click a layer. Engines stay engines. InferenceOSS writes the math and the notes between them.
Layer 02 · we specify
Arithmetic intensity, HBM bytes per token, KV footprint, ridge point. If the equation does not move, the flag will not either.
Incubator
The equations first: decode wall, prefill FLOPs, KV bytes, arithmetic intensity, occupancy. Worked examples on named cards.
For this GPU, this model, this quant, this engine: measured prefill, decode, and concurrent streams. Commands, not screenshots.
GEMM tiling, Tensor Core shapes, FlashAttention, fused RMSNorm, warp stalls. The notes CUDA docs should have started with.
NVFP4, MXFP4, EXL3, GGUF, AWQ as a trade: bits, error, Tensor Core paths, and the context you buy back.
MTP, DFlash, draft models, acceptance curves. When speculation is 1.8× and when it is 0.27×.
How to squeeze one card with llama.cpp, LocalAI, ExLlama, vLLM, and SGLang. Honest flags. No thirteenth engine.
Prefill, decode, prose, concurrent streams. Pinned builds. Named GPUs. Anyone can rerun the command.
How to read Nsight Compute, Nsight Systems, and rocprof for inference. Occupancy, warp stalls, L2, tensor pipe.
90 days
No keynotes. Equations, four recipes, a harness, and a first public session.
01
Publish this document. Invite twelve interim TSC nominees from kernel, local-inference, and operator backgrounds. Seat seven. Publish meeting notes.
02
Freeze the six equations. Work them on a 3090, a 5090, and a Spark. One calculator page.
03
Four recipes, same 27B-class model, two quants. Prefill and decode split. Commands in the repo.
04
CLI that prints prefill, decode, VRAM, power, and the exact argv. Pin one model. Do not declare a winner.
05
Memory hierarchy and the decode wall, with one annotated Nsight trace. No keynote.
06
A three-hour public session. Math and recipes only. Notes on the site within a day.