RAIF is a repairable format for model output

RAIF is a wire format for one JSON object produced by a language model. The model emits RAIF instead of JSON. A deterministic decoder repairs known syntax mistakes and returns ordinary JSON to the application.

RAIF, Repairable AI Interchange Format

What RAIF changes

JSON expects a deterministic writer. A language model can add a markdown fence, use the wrong separator, repeat a key, or stop before the response is complete. One mistake can make the whole JSON document fail to parse.

RAIF expects a probabilistic writer and a deterministic interpreter. It keeps values in small independent leaves. The decoder can report a damaged leaf without throwing away every intact value. It only repairs unambiguous syntax and records every repair.

Why it is useful

RAIF is usually smaller than JSON because it removes repeated structure. Uniform records can use a table where field names appear once.

AspectJSONRAIF
Token cost on real function callsbaselineabout 9 to 10 percent lower
Repeated recordsevery key is repeatedfield names can appear once
Known syntax mistakesparse failsrepaired and reported
Cut off outputparse failsintact leaves can be returned
Application valueJSONthe same JSON after decode

Across 10,677 real function call payloads, RAIF used 9.2 percent fewer cl100k tokens and 10.2 percent fewer o200k tokens. The median saving was 11 percent. RAIF was larger in about 3 percent of cases.

The saving comes from the shape of the data. A single flat object can be close to JSON. Tables and repeated records save much more.

How it works

  1. A language model emits RAIF instead of JSON. The published adapters train this behavior directly.
  2. The codec reads each leaf and applies only deterministic syntax repairs. Ambiguous input is rejected.
  3. The codec returns the JSON value expected by the application. The lenient decoder also returns intact leaves and named errors when output is cut off.

A small example

Scalar values use one line. Nested values use paths or inline objects. Arrays can use paths, array literals, inline objects, or tables.

The canonical encoder builds the legal representations and keeps the shortest one by byte length. It is the reference conversion from known JSON to RAIF.

TypeScript
import { encode, decode, decodeLenient } from "raif-format";

const input = { user: { name: "Ada", email: "ada@x.io" }, active: true };
const raif = encode(input);
// active=true
// user={email=ada@x.io,name=Ada}

decode(raif).value; // the same JSON value as input

The decoder strips markdown fences and repairs a stray separator when the intended structure is unambiguous. It keeps a record of both changes.

TypeScript
const repaired = decode("```\nactive=true\nuser.name:Ada\n```");
repaired.value;
// { active: true, user: { name: "Ada" } }
repaired.repairs;
// [{ kind: "markdown_stripped" }, { kind: "separator_coerced", ... }]

decodeLenient("<raif>\ncity=Oslo\nlat"); // stream cut off
// { value: { city: "Oslo" }, truncated: true, errors: [{ line: 2, … }] }

A larger example

Uniform arrays can use a table. Field names appear once, then each record uses one line.

This transaction result is 590 bytes in minified JSON and 356 bytes in RAIF. The RAIF version is 39.7 percent smaller and decodes to the same data.

JSON 590 B
{
  "query": "status:settled",
  "page": 1,
  "pageSize": 6,
  "total": 6,
  "results": [
    {
      "id": "tx_1001",
      "amount": 4200,
      "currency": "USD",
      "status": "settled",
      "vendor": "Acme Cloud"
    },
    {
      "id": "tx_1002",
      "amount": 1875,
      "currency": "USD",
      "status": "settled",
      "vendor": "Northwind"
    },
    {
      "id": "tx_1003",
      "amount": 990,
      "currency": "EUR",
      "status": "settled",
      "vendor": "Globex"
    },
    {
      "id": "tx_1004",
      "amount": 15400,
      "currency": "USD",
      "status": "settled",
      "vendor": "Initech"
    },
    {
      "id": "tx_1005",
      "amount": 320,
      "currency": "GBP",
      "status": "settled",
      "vendor": "Soylent"
    },
    {
      "id": "tx_1006",
      "amount": 7600,
      "currency": "USD",
      "status": "settled",
      "vendor": "Umbrella"
    }
  ]
}
RAIF 356 B, −39.7%
page=1
pageSize=6
query=status:settled
results::amount,currency,id,status,vendor
results[0]=4200,USD,tx_1001,settled,Acme Cloud
results[1]=1875,USD,tx_1002,settled,Northwind
results[2]=990,EUR,tx_1003,settled,Globex
results[3]=15400,USD,tx_1004,settled,Initech
results[4]=320,GBP,tx_1005,settled,Soylent
results[5]=7600,USD,tx_1006,settled,Umbrella
total=6

Try it yourself

Paste a JSON object and convert it in the browser. Switch direction to decode RAIF back to JSON. The playground uses the codec from the official distribution.

This shows the reference round trip. In normal model use, the model produces RAIF and the application calls decode.

JSON editable
RAIF 590 → 356 B, −39.7%
page=1
pageSize=6
query=status:settled
results::amount,currency,id,status,vendor
results[0]=4200,USD,tx_1001,settled,Acme Cloud
results[1]=1875,USD,tx_1002,settled,Northwind
results[2]=990,EUR,tx_1003,settled,Globex
results[3]=15400,USD,tx_1004,settled,Initech
results[4]=320,GBP,tx_1005,settled,Soylent
results[5]=7600,USD,tx_1006,settled,Umbrella
total=6

Models that emit RAIF

The published adapters make small self hosted models emit RAIF directly. Three adapters are available on Hugging Face.

ModelBaseFootprintResult
raif-llama-3.2-3b-loraLlama 3.2 3Breference adapter (~195 MB)100% parse / 95% fidelity
raif-qwen3-4b-loraQwen3 4Bagent-grade (~14 GB VRAM)98% parse / 95% fidelity
raif-qwen2.5-0.5b-loraQwen2.5 0.5Btiny / edge97% parse / 81% fidelity

The parse score measures valid RAIF output. The fidelity score measures whether the decoded JSON matches the expected value.

Use the package

The package name is raif-format on npm and PyPI. The core codec has no runtime dependencies.

Shell
npm install raif-format    # also works with bun and pnpm
pip install raif-format    # imports as raif
TypeScript
import { decode } from "raif-format";

const result = decode(modelOutput, schema);
if (result.ok) {
  useJson(result.value);
}

The model path starts with RAIF output and decode. The encode function is the reference JSON to RAIF conversion. It is useful for testing, canonicalization, and known JSON data.

The full API is small.

TypeScript
encode(obj, opts?)            // JSON object → RAIF
decode(raif, schema?)         // → { ok, value, repairs }
decodeLenient(raif, schema?)  // → { value, errors, truncated, repairs }
fix(raif, schema?)            // → canonical RAIF
validate(raif, schema?)       // read-only canonicality check

The canonical profile chooses the shortest legal representation. The generation profile uses stable forms that models are trained to emit.

Use it with vLLM

raif-vllm adds RAIF support to a standard OpenAI compatible vLLM endpoint. The model emits RAIF. The plugin decodes it at the server boundary and returns JSON to existing clients.

Shell
pip install raif-vllm

VLLM_PLUGINS=raif vllm serve unsloth/Llama-3.2-3B-Instruct \
  --enable-lora --lora-modules raif=skrrt-sh/raif-llama-3.2-3b-lora \
  --chat-template "$(raif-vllm-chat-template llama-3b)" \
  --reasoning-parser raif --enable-auto-tool-choice --tool-call-parser raif

Scope

RAIF covers one JSON object produced by a language model. It supports strings, numbers, booleans, null values, arrays, and nested objects.

It is not a general interchange format. It is not compression, a schema language, or an input format for a model. Its job is to make structured model output smaller and easier to recover.

Resources

RAIF uses the Apache 2.0 license. It includes a patent grant for anyone who implements or builds on the format.