inferenceoss

GPU inference, measured.

This is the founding note for an open guild of GPU performance engineers. It is a group, not a petition. 0 people have joined so far.

01

The problem

Inference is where models become products, and it is the layer with the least shared math. Training has papers. Serving has flags. The gap between a 27B model and a usable tok/s on a named card is performance engineering: roofline, kernels, quantization, and a recipe someone else can rerun.

Batch-1 decode is memory-bound. Prefill is compute-bound. Long context is a KV allocation. Speculation is an expected-value problem with an acceptance rate almost nobody publishes. llama.cpp, LocalAI, ExLlama, vLLM, and SGLang are excellent, and they still get tuned by folklore.

The useful knowledge currently lives in threads: which card for which quant, whether NVFP4 beats EXL3 on a 3090, how a Spark differs from a 5090 at batch 1 versus 4. That work is real. It is not a constitution.

02

Mission

Make GPU inference fast, measured, and teachable.

InferenceOSS incubates equations, kernel notes, named-card recipes, and a public harness. It does not incubate another engine. A stack that hits the decode wall on a 16 GB card is a citizen, whether it runs under llama.cpp or vLLM.

03

The math

Six identities. If a change cannot be stated in one of them, it is not a performance change.

  1. Decode wall

    tok/s ≈ B / (N · q)

    B = HBM bandwidth, N = parameters, q = bytes per weight

    Autoregressive decode re-reads the weights every token. On a 3090 (936 GB/s) a 27B model at 4-bit is ~87 tok/s if you actually hit the bus. If you measure less, you are leaving bandwidth on the table.

  2. Prefill

    FLOPs ≈ 2 · N · S

    S = sequence length; tok/s ≈ peak FLOP/s / FLOPs per token

    Prefill is compute-bound until the sequence is short or the quant is extreme. Tensor Cores matter here. They barely matter at batch-1 decode.

  3. KV cache

    bytes = 2 · L · n_kv · d_h · S · q_kv

    L layers, n_kv heads, d_h head dim, S sequence, q_kv bytes per element

    Long context is a RAM tax, not a FLOP tax. GQA, quantized KV, and a smaller n_kv buy sequence. They do not buy a smarter model.

  4. Roofline

    I = FLOPs / bytes; achieved = min(peak_FLOP, I · B)

    I = arithmetic intensity. Ridge point = peak_FLOP / B

    If I is below the ridge, more Tensor Cores do nothing. Fuse, tile, or raise batch until I climbs, or buy bandwidth.

  5. Speculation

    E[tokens] = (1 − α^{γ+1}) / (1 − α)

    α = draft accept rate, γ = draft tokens per step

    A fast draft on a slow target wins. A slow draft on a fast 7B is a tax. Measure α on your prompt, not on HumanEval.

  6. Occupancy

    occ = active_warps / max_warps

    Limited by registers, shared memory, and the block size you picked

    nvidia-smi at 100% can still be stalled on the scoreboard. Nsight is the instrument. Occupancy is the first number it should teach you.

04

Principles

Six constraints. Endorse the ones you will actually defend in a working group.

  1. 01

    Write the equation before the flag

    A tok/s number without a roofline is a vibe. Decode is memory-bound, prefill is compute-bound, long context is a KV tax. If a change cannot be stated as FLOPs, bytes, or occupancy, it is not a performance change.

  2. 02

    Name the card, the build, and the prompt

    Recipes are first-class. Model, quant, engine, driver, batch, context, and the GPU on the desk. A number that cannot be reproduced on a 3090, a 5090, a Spark, or a Strix Halo is marketing.

  3. 03

    Local and cluster are one problem

    llama.cpp, LocalAI, ExLlama, vLLM, and SGLang are citizens. A note that only works behind an H100 is half a note. The same algebra governs a 16 GB gamer card and a rack.

  4. 04

    The kernel is the product

    GEMM tiles, Tensor Core shapes, FlashAttention, fused RMSNorm, and warp stalls are the work. Wrappers are not. We publish notes and patches, not another serving stack.

  5. 05

    Quantization is a trade, not a trick

    NVFP4, MXFP4, EXL3, GGUF, and AWQ buy bandwidth and context. They also buy error. Publish the quality drop next to the tok/s, or do not publish the tok/s.

  6. 06

    Apache 2.0, DCO, no CLA tax

    Default license is Apache 2.0 with an implicit patent grant. Contributions are Developer Certificate of Origin, not a lawyer-gated CLA. Trademarks stay with the org; copyright stays with authors.

05

The program

Eight incubator projects. All start as drafts. Graduation is two independent measurements on named cards, published as commands, not a slide.

  1. 01 · Roofline

    draft

    Inference math notebook

    The equations first: decode wall, prefill FLOPs, KV bytes, arithmetic intensity, occupancy. Worked examples on named cards.

    Open the brief
  2. 02 · CardBook

    draft

    Named-card recipe atlas

    For this GPU, this model, this quant, this engine: measured prefill, decode, and concurrent streams. Commands, not screenshots.

    Open the brief
  3. 03 · Kernels

    draft

    Kernel notes

    GEMM tiling, Tensor Core shapes, FlashAttention, fused RMSNorm, warp stalls. The notes CUDA docs should have started with.

    Open the brief
  4. 04 · Quant

    draft

    Quantization lab

    NVFP4, MXFP4, EXL3, GGUF, AWQ as a trade: bits, error, Tensor Core paths, and the context you buy back.

    Open the brief
  5. 05 · Speculate

    draft

    Speculative decoding notes

    MTP, DFlash, draft models, acceptance curves. When speculation is 1.8× and when it is 0.27×.

    Open the brief
  6. 06 · Local

    draft

    Local stack notes

    How to squeeze one card with llama.cpp, LocalAI, ExLlama, vLLM, and SGLang. Honest flags. No thirteenth engine.

    Open the brief
  7. 07 · Bench

    draft

    Public inference harness

    Prefill, decode, prose, concurrent streams. Pinned builds. Named GPUs. Anyone can rerun the command.

    Open the brief
  8. 08 · Trace

    draft

    Profiler literacy

    How to read Nsight Compute, Nsight Systems, and rocprof for inference. Occupancy, warp stalls, L2, tensor pipe.

    Open the brief

06

Working groups

Six rooms. Join the ones you will attend.

  • Math

    Roofline, KV algebra, occupancy. The notebook and the calculator.

  • Kernels

    GEMM, attention, fusion, speculation, profiler literacy. Notes and upstream patches.

  • Quant

    NVFP4, EXL3, GGUF, AWQ. Bits, error, and Tensor Core paths.

  • Hardware

    CardBook. Named GPUs, VRAM, bandwidth, recipes that fit.

  • Local

    llama.cpp, LocalAI, ExLlama, one-box serving. The card on the desk.

  • Bench

    The harness. Prefill, decode, concurrency, published commands.

07

The first 90 days

  1. 01

    Guild and TSC

    Publish this document. Invite twelve interim TSC nominees from kernel, local-inference, and operator backgrounds. Seat seven. Publish meeting notes.

  2. 02

    Roofline notebook 0.1

    Freeze the six equations. Work them on a 3090, a 5090, and a Spark. One calculator page.

  3. 03

    CardBook seed

    Four recipes, same 27B-class model, two quants. Prefill and decode split. Commands in the repo.

  4. 04

    Bench harness

    CLI that prints prefill, decode, VRAM, power, and the exact argv. Pin one model. Do not declare a winner.

  5. 05

    Kernel note 1

    Memory hierarchy and the decode wall, with one annotated Nsight trace. No keynote.

  6. 06

    First working session

    A three-hour public session. Math and recipes only. Notes on the site within a day.

08

Governance

Form
Independent nonprofit (or unincorporated association on the way there), with a path to a directed fund if that later serves the guild. The guild is the constraint, not the parent org.
TSC
Seven seats, two-year terms, odd number on purpose. No more than two seats from any one employer. TSC votes on incubation, graduation, and note releases. Working groups propose; they do not ship alone.
Membership
Individuals join the guild for free. Academic labs join at cost. Industry membership funds cards, a tiny staff, and the bench log. Dues do not buy a TSC seat. One organization, one vote in TSC elections.
IP
Apache 2.0 default. DCO on every commit. Notes and harnesses are dual-licensed as code. The InferenceOSS mark is held by the org and licensed to conforming recipes.
Why not NVIDIA / vLLM / LocalAI
NVIDIA ships the hardware and many of the kernels. vLLM, SGLang, llama.cpp, ExLlama, and LocalAI ship engines. Those are the right homes for those projects. InferenceOSS exists because the math, the named-card recipe, and the public measurement have no owner that is not also selling a GPU or a runtime.

09

Success

If these are not true in 2029, the org failed, however many stars the repos have.

  • The decode-wall equation is cited in two engine docs that we did not write.
  • CardBook holds recipes for at least six named GPUs, each with commands a stranger can rerun.
  • Bench is used in place of a screenshot at least once in a serious venue.
  • A kernel note leads to an upstream patch in llama.cpp, vLLM, or SGLang.
  • A 16 GB builder and a cluster operator can talk in intensity, KV bytes, and α.
  • No single employer holds a TSC majority, including in the fifth year.

Guild clerk

Ask the guild

The clerk is paused while membership is email interest only.

Guild 0.1 · 6 principles · 8 projects