inferenceoss
GPU inference, measured.
This is the founding note for an open guild of GPU performance engineers. It is a group, not a petition. 0 people have joined so far.
01
The problem
Inference is where models become products, and it is the layer with the least shared math. Training has papers. Serving has flags. The gap between a 27B model and a usable tok/s on a named card is performance engineering: roofline, kernels, quantization, and a recipe someone else can rerun.
Batch-1 decode is memory-bound. Prefill is compute-bound. Long context is a KV allocation. Speculation is an expected-value problem with an acceptance rate almost nobody publishes. llama.cpp, LocalAI, ExLlama, vLLM, and SGLang are excellent, and they still get tuned by folklore.
The useful knowledge currently lives in threads: which card for which quant, whether NVFP4 beats EXL3 on a 3090, how a Spark differs from a 5090 at batch 1 versus 4. That work is real. It is not a constitution.
02
Mission
Make GPU inference fast, measured, and teachable.
InferenceOSS incubates equations, kernel notes, named-card recipes, and a public harness. It does not incubate another engine. A stack that hits the decode wall on a 16 GB card is a citizen, whether it runs under llama.cpp or vLLM.
03
The math
Six identities. If a change cannot be stated in one of them, it is not a performance change.
Decode wall
tok/s ≈ B / (N · q)
B = HBM bandwidth, N = parameters, q = bytes per weight
Autoregressive decode re-reads the weights every token. On a 3090 (936 GB/s) a 27B model at 4-bit is ~87 tok/s if you actually hit the bus. If you measure less, you are leaving bandwidth on the table.
Prefill
FLOPs ≈ 2 · N · S
S = sequence length; tok/s ≈ peak FLOP/s / FLOPs per token
Prefill is compute-bound until the sequence is short or the quant is extreme. Tensor Cores matter here. They barely matter at batch-1 decode.
KV cache
bytes = 2 · L · n_kv · d_h · S · q_kv
L layers, n_kv heads, d_h head dim, S sequence, q_kv bytes per element
Long context is a RAM tax, not a FLOP tax. GQA, quantized KV, and a smaller n_kv buy sequence. They do not buy a smarter model.
Roofline
I = FLOPs / bytes; achieved = min(peak_FLOP, I · B)
I = arithmetic intensity. Ridge point = peak_FLOP / B
If I is below the ridge, more Tensor Cores do nothing. Fuse, tile, or raise batch until I climbs, or buy bandwidth.
Speculation
E[tokens] = (1 − α^{γ+1}) / (1 − α)
α = draft accept rate, γ = draft tokens per step
A fast draft on a slow target wins. A slow draft on a fast 7B is a tax. Measure α on your prompt, not on HumanEval.
Occupancy
occ = active_warps / max_warps
Limited by registers, shared memory, and the block size you picked
nvidia-smi at 100% can still be stalled on the scoreboard. Nsight is the instrument. Occupancy is the first number it should teach you.
04
Principles
Six constraints. Endorse the ones you will actually defend in a working group.
- 01
Write the equation before the flag
A tok/s number without a roofline is a vibe. Decode is memory-bound, prefill is compute-bound, long context is a KV tax. If a change cannot be stated as FLOPs, bytes, or occupancy, it is not a performance change.
- 02
Name the card, the build, and the prompt
Recipes are first-class. Model, quant, engine, driver, batch, context, and the GPU on the desk. A number that cannot be reproduced on a 3090, a 5090, a Spark, or a Strix Halo is marketing.
- 03
Local and cluster are one problem
llama.cpp, LocalAI, ExLlama, vLLM, and SGLang are citizens. A note that only works behind an H100 is half a note. The same algebra governs a 16 GB gamer card and a rack.
- 04
The kernel is the product
GEMM tiles, Tensor Core shapes, FlashAttention, fused RMSNorm, and warp stalls are the work. Wrappers are not. We publish notes and patches, not another serving stack.
- 05
Quantization is a trade, not a trick
NVFP4, MXFP4, EXL3, GGUF, and AWQ buy bandwidth and context. They also buy error. Publish the quality drop next to the tok/s, or do not publish the tok/s.
- 06
Apache 2.0, DCO, no CLA tax
Default license is Apache 2.0 with an implicit patent grant. Contributions are Developer Certificate of Origin, not a lawyer-gated CLA. Trademarks stay with the org; copyright stays with authors.
05
The program
Eight incubator projects. All start as drafts. Graduation is two independent measurements on named cards, published as commands, not a slide.
01 · Roofline
draftInference math notebook
The equations first: decode wall, prefill FLOPs, KV bytes, arithmetic intensity, occupancy. Worked examples on named cards.
Open the brief02 · CardBook
draftNamed-card recipe atlas
For this GPU, this model, this quant, this engine: measured prefill, decode, and concurrent streams. Commands, not screenshots.
Open the brief03 · Kernels
draftKernel notes
GEMM tiling, Tensor Core shapes, FlashAttention, fused RMSNorm, warp stalls. The notes CUDA docs should have started with.
Open the brief04 · Quant
draftQuantization lab
NVFP4, MXFP4, EXL3, GGUF, AWQ as a trade: bits, error, Tensor Core paths, and the context you buy back.
Open the brief05 · Speculate
draftSpeculative decoding notes
MTP, DFlash, draft models, acceptance curves. When speculation is 1.8× and when it is 0.27×.
Open the brief06 · Local
draftLocal stack notes
How to squeeze one card with llama.cpp, LocalAI, ExLlama, vLLM, and SGLang. Honest flags. No thirteenth engine.
Open the brief07 · Bench
draftPublic inference harness
Prefill, decode, prose, concurrent streams. Pinned builds. Named GPUs. Anyone can rerun the command.
Open the brief08 · Trace
draftProfiler literacy
How to read Nsight Compute, Nsight Systems, and rocprof for inference. Occupancy, warp stalls, L2, tensor pipe.
Open the brief
06
Working groups
Six rooms. Join the ones you will attend.
Math
Roofline, KV algebra, occupancy. The notebook and the calculator.
Kernels
GEMM, attention, fusion, speculation, profiler literacy. Notes and upstream patches.
Quant
NVFP4, EXL3, GGUF, AWQ. Bits, error, and Tensor Core paths.
Hardware
CardBook. Named GPUs, VRAM, bandwidth, recipes that fit.
Local
llama.cpp, LocalAI, ExLlama, one-box serving. The card on the desk.
Bench
The harness. Prefill, decode, concurrency, published commands.
07
The first 90 days
- 01
Guild and TSC
Publish this document. Invite twelve interim TSC nominees from kernel, local-inference, and operator backgrounds. Seat seven. Publish meeting notes.
- 02
Roofline notebook 0.1
Freeze the six equations. Work them on a 3090, a 5090, and a Spark. One calculator page.
- 03
CardBook seed
Four recipes, same 27B-class model, two quants. Prefill and decode split. Commands in the repo.
- 04
Bench harness
CLI that prints prefill, decode, VRAM, power, and the exact argv. Pin one model. Do not declare a winner.
- 05
Kernel note 1
Memory hierarchy and the decode wall, with one annotated Nsight trace. No keynote.
- 06
First working session
A three-hour public session. Math and recipes only. Notes on the site within a day.
08
Governance
- Form
- Independent nonprofit (or unincorporated association on the way there), with a path to a directed fund if that later serves the guild. The guild is the constraint, not the parent org.
- TSC
- Seven seats, two-year terms, odd number on purpose. No more than two seats from any one employer. TSC votes on incubation, graduation, and note releases. Working groups propose; they do not ship alone.
- Membership
- Individuals join the guild for free. Academic labs join at cost. Industry membership funds cards, a tiny staff, and the bench log. Dues do not buy a TSC seat. One organization, one vote in TSC elections.
- IP
- Apache 2.0 default. DCO on every commit. Notes and harnesses are dual-licensed as code. The InferenceOSS mark is held by the org and licensed to conforming recipes.
- Why not NVIDIA / vLLM / LocalAI
- NVIDIA ships the hardware and many of the kernels. vLLM, SGLang, llama.cpp, ExLlama, and LocalAI ship engines. Those are the right homes for those projects. InferenceOSS exists because the math, the named-card recipe, and the public measurement have no owner that is not also selling a GPU or a runtime.
09
Success
If these are not true in 2029, the org failed, however many stars the repos have.
- The decode-wall equation is cited in two engine docs that we did not write.
- CardBook holds recipes for at least six named GPUs, each with commands a stranger can rerun.
- Bench is used in place of a screenshot at least once in a serious venue.
- A kernel note leads to an upstream patch in llama.cpp, vLLM, or SGLang.
- A 16 GB builder and a cluster operator can talk in intensity, KV bytes, and α.
- No single employer holds a TSC majority, including in the fifth year.
Guild clerk
Ask the guild
The clerk is paused while membership is email interest only.
Guild 0.1 · 6 principles · 8 projects