Incubator
Labs, not products.
Everything starts as a draft. Graduation is two independent measurements on named cards. InferenceOSS does not ship an engine.
- Rooflinedraft
Inference math notebook
The equations first: decode wall, prefill FLOPs, KV bytes, arithmetic intensity, occupancy. Worked examples on named cards.
WG · Math
- CardBookdraft
Named-card recipe atlas
For this GPU, this model, this quant, this engine: measured prefill, decode, and concurrent streams. Commands, not screenshots.
WG · Hardware
- Kernelsdraft
Kernel notes
GEMM tiling, Tensor Core shapes, FlashAttention, fused RMSNorm, warp stalls. The notes CUDA docs should have started with.
WG · Kernels
- Quantdraft
Quantization lab
NVFP4, MXFP4, EXL3, GGUF, AWQ as a trade: bits, error, Tensor Core paths, and the context you buy back.
WG · Quant
- Speculatedraft
Speculative decoding notes
MTP, DFlash, draft models, acceptance curves. When speculation is 1.8× and when it is 0.27×.
WG · Kernels
- Localdraft
Local stack notes
How to squeeze one card with llama.cpp, LocalAI, ExLlama, vLLM, and SGLang. Honest flags. No thirteenth engine.
WG · Local
- Benchdraft
Public inference harness
Prefill, decode, prose, concurrent streams. Pinned builds. Named GPUs. Anyone can rerun the command.
WG · Bench
- Tracedraft
Profiler literacy
How to read Nsight Compute, Nsight Systems, and rocprof for inference. Occupancy, warp stalls, L2, tensor pipe.
WG · Kernels