Guild 0.1

tok/s ≈ B / (N · q)

Decode is a memory problem. A guild for GPU performance engineering of inference: the math, the kernels, the quants, and the recipes that make a named card actually fast. Not another engine.

The wall

Batch-1 decode re-reads every weight, every token. tok/s ≈ bandwidth / (params × bytes). If you are not on that line, the kernel is the bug, not the model.

The folklore

Flags, screenshots, and unnamed GPUs. NVFP4, EXL3, GGUF, MTP, DFlash: real tools, written down as threads. A 3090 recipe and a Spark recipe are different objects.

The bet

The org that owns the math and the named-card notebook will outlast any single runtime, the way a roofline outlasts a compiler flag.

The algebra

Six equations. Then the flags.

Decode, prefill, KV, roofline, speculation, occupancy. Worked as if a 3090 were on the desk.

  1. Decode wall

    tok/s ≈ B / (N · q)

    B = HBM bandwidth, N = parameters, q = bytes per weight

    Autoregressive decode re-reads the weights every token. On a 3090 (936 GB/s) a 27B model at 4-bit is ~87 tok/s if you actually hit the bus. If you measure less, you are leaving bandwidth on the table.

  2. Prefill

    FLOPs ≈ 2 · N · S

    S = sequence length; tok/s ≈ peak FLOP/s / FLOPs per token

    Prefill is compute-bound until the sequence is short or the quant is extreme. Tensor Cores matter here. They barely matter at batch-1 decode.

  3. KV cache

    bytes = 2 · L · n_kv · d_h · S · q_kv

    L layers, n_kv heads, d_h head dim, S sequence, q_kv bytes per element

    Long context is a RAM tax, not a FLOP tax. GQA, quantized KV, and a smaller n_kv buy sequence. They do not buy a smarter model.

  4. Roofline

    I = FLOPs / bytes; achieved = min(peak_FLOP, I · B)

    I = arithmetic intensity. Ridge point = peak_FLOP / B

    If I is below the ridge, more Tensor Cores do nothing. Fuse, tile, or raise batch until I climbs, or buy bandwidth.

  5. Speculation

    E[tokens] = (1 − α^{γ+1}) / (1 − α)

    α = draft accept rate, γ = draft tokens per step

    A fast draft on a slow target wins. A slow draft on a fast 7B is a tax. Measure α on your prompt, not on HumanEval.

  6. Occupancy

    occ = active_warps / max_warps

    Limited by registers, shared memory, and the block size you picked

    nvidia-smi at 100% can still be stalled on the scoreboard. Nsight is the instrument. Occupancy is the first number it should teach you.

The stack

From FLOPs to a named card.

Click a layer. Engines stay engines. InferenceOSS writes the math and the notes between them.

Layer 02 · we specify

Roofline & KV

Arithmetic intensity, HBM bytes per token, KV footprint, ridge point. If the equation does not move, the flag will not either.

Incubator

Eight labs. Zero engines.

All projects

90 days

What we measure before the year ends.

No keynotes. Equations, four recipes, a harness, and a first public session.

  1. 01

    Guild and TSC

    Publish this document. Invite twelve interim TSC nominees from kernel, local-inference, and operator backgrounds. Seat seven. Publish meeting notes.

  2. 02

    Roofline notebook 0.1

    Freeze the six equations. Work them on a 3090, a 5090, and a Spark. One calculator page.

  3. 03

    CardBook seed

    Four recipes, same 27B-class model, two quants. Prefill and decode split. Commands in the repo.

  4. 04

    Bench harness

    CLI that prints prefill, decode, VRAM, power, and the exact argv. Pin one model. Do not declare a winner.

  5. 05

    Kernel note 1

    Memory hierarchy and the decode wall, with one annotated Nsight trace. No keynote.

  6. 06

    First working session

    A three-hour public session. Math and recipes only. Notes on the site within a day.

Join the guild