Charter 0.1

The last mile of AI is ungoverned.

A vendor-neutral organization for the protocols, interchange formats, and reference implementations that make inference engines interchangeable. Not another engine.

The fracture

vLLM, SGLang, llama.cpp, TGI, TensorRT-LLM, Ollama, llm-d, Dynamo — each a fiefdom. The OpenAI API is IPv4: de facto, underspecified, silently incompatible.

The gap

Training has PyTorch Foundation. Kubernetes serving has KServe. The interfaces between engines — protocol, KV cache, capabilities, telemetry — have no owner that is not also a vendor.

The bet

The org that owns the contracts of inference will matter more than any single runtime, the way Kubernetes mattered more than any container daemon.

The stack

We specify the layers engines currently improvise.

Click a layer. Engines stay engines. InferenceOSS writes the sockets between them.

Layer 02 · we specify

Open Inference Protocol (OIP-G)

Streaming chat, tools, structured output, multimodal parts, and cancellation. Capability-aware: a client asks what the engine can do instead of guessing from undocumented OpenAI extras.

Roadmap

Four years to boring.

  1. Phase 0 · Aug–Dec 2026

    Found

    Ratify Charter 0.1 and stand up an interim TSC of seven, no employer majority.

  2. Phase 1 · 2027 H1

    Interfaces

    OIP-G v0.2: tools, structured output, abort, capabilities.

  3. Phase 2 · 2027 H2

    Interchange

    KVX adapters on four engines.

  4. Phase 3 · 2028

    1.0

    OIP-G 1.0. CapSpec and InferTrace graduate.

  5. Phase 4 · 2029

    Default

    “Speaks OIP-G” is a normal checkbox, the way “speaks OpenAPI” is for HTTP services.

Full roadmap

90 days

What we do before the year ends.

No keynotes. RFCs, a schema, a harness, and a first public session.

  1. 01

    Charter and TSC

    Publish this document. Invite twelve interim TSC nominees from independent engine, academic, and operator backgrounds. Seat seven. Publish meeting notes.

  2. 02

    CapSpec 0.1 RFC

    Freeze a JSON schema. Ship generators for vLLM, SGLang, and llama.cpp. One page of docs, no more.

  3. 03

    InferTrace draft

    Name TTFT, ITL, queue, cache, batch, speculation, abort. Align with OpenTelemetry. Send to the OTel SIG.

  4. 04

    SPEC/I harness

    Pin one model, one GPU type, four engines, three workloads. Publish the exact commands. Do not declare a winner.

  5. 05

    Legal and identity

    Apache 2.0, DCO, trademark filing for InferenceOSS, domain, GitHub org. No CLA.

  6. 06

    First working session

    A three-hour public session. RFCs only. No keynotes. Notes on the site within a day.

Discipline

What we will not do.

  • Write another full serving engine.
  • Operate a hosted inference cloud. The org is not a vendor.
  • Bless a single accelerator, quantization, or kernel family in a core spec.
  • Replace PyTorch, vLLM, KServe, or OpenTelemetry. We interoperate.
  • Sell a leaderboard, or let a member-submitted number stand in for SPEC/I.
  • Require a CLA, a contributor license assignment, or an employment gate.

Founding roll

Sign in public.

Individuals sign the charter for free. Names, roles, and intent are published. Dues do not buy a vote.

The roll is empty.

Founding contributors sign the charter in public. Be first.

Sign the charter