RAIF is a repairable format for model output
RAIF is a wire format for one JSON object produced by a language model. The model emits RAIF instead of JSON. A deterministic decoder repairs known syntax mistakes and returns ordinary JSON to the application.

What RAIF changes
JSON expects a deterministic writer. A language model can add a markdown fence, use the wrong separator, repeat a key, or stop before the response is complete. One mistake can make the whole JSON document fail to parse.
RAIF expects a probabilistic writer and a deterministic interpreter. It keeps values in small independent leaves. The decoder can report a damaged leaf without throwing away every intact value. It only repairs unambiguous syntax and records every repair.
Why it is useful
RAIF is usually smaller than JSON because it removes repeated structure. Uniform records can use a table where field names appear once.
| Aspect | JSON | RAIF |
|---|---|---|
| Token cost on real function calls | baseline | about 9 to 10 percent lower |
| Repeated records | every key is repeated | field names can appear once |
| Known syntax mistakes | parse fails | repaired and reported |
| Cut off output | parse fails | intact leaves can be returned |
| Application value | JSON | the same JSON after decode |
Across 10,677 real function call payloads, RAIF used 9.2 percent fewer cl100k tokens and 10.2 percent fewer o200k tokens. The median saving was 11 percent. RAIF was larger in about 3 percent of cases.
The saving comes from the shape of the data. A single flat object can be close to JSON. Tables and repeated records save much more.
How it works
- A language model emits RAIF instead of JSON. The published adapters train this behavior directly.
- The codec reads each leaf and applies only deterministic syntax repairs. Ambiguous input is rejected.
- The codec returns the JSON value expected by the application. The lenient decoder also returns intact leaves and named errors when output is cut off.
A small example
Scalar values use one line. Nested values use paths or inline objects. Arrays can use paths, array literals, inline objects, or tables.
The canonical encoder builds the legal representations and keeps the shortest one by byte length. It is the reference conversion from known JSON to RAIF.
The decoder strips markdown fences and repairs a stray separator when the intended structure is unambiguous. It keeps a record of both changes.
A larger example
Uniform arrays can use a table. Field names appear once, then each record uses one line.
This transaction result is 590 bytes in minified JSON and 356 bytes in RAIF. The RAIF version is 39.7 percent smaller and decodes to the same data.
Try it yourself
Paste a JSON object and convert it in the browser. Switch direction to decode RAIF back to JSON. The playground uses the codec from the official distribution.
This shows the reference round trip. In normal model use, the model produces RAIF and the application calls decode.
page=1
pageSize=6
query=status:settled
results::amount,currency,id,status,vendor
results[0]=4200,USD,tx_1001,settled,Acme Cloud
results[1]=1875,USD,tx_1002,settled,Northwind
results[2]=990,EUR,tx_1003,settled,Globex
results[3]=15400,USD,tx_1004,settled,Initech
results[4]=320,GBP,tx_1005,settled,Soylent
results[5]=7600,USD,tx_1006,settled,Umbrella
total=6Models that emit RAIF
The published adapters make small self hosted models emit RAIF directly. Three adapters are available on Hugging Face.
| Model | Base | Footprint | Result |
|---|---|---|---|
| raif-llama-3.2-3b-lora | Llama 3.2 3B | reference adapter (~195 MB) | 100% parse / 95% fidelity |
| raif-qwen3-4b-lora | Qwen3 4B | agent-grade (~14 GB VRAM) | 98% parse / 95% fidelity |
| raif-qwen2.5-0.5b-lora | Qwen2.5 0.5B | tiny / edge | 97% parse / 81% fidelity |
The parse score measures valid RAIF output. The fidelity score measures whether the decoded JSON matches the expected value.
Use the package
The package name is raif-format on npm and PyPI. The core codec has no runtime dependencies.
The model path starts with RAIF output and decode. The encode function is the reference JSON to RAIF conversion. It is useful for testing, canonicalization, and known JSON data.
The full API is small.
The canonical profile chooses the shortest legal representation. The generation profile uses stable forms that models are trained to emit.
Use it with vLLM
raif-vllm adds RAIF support to a standard OpenAI compatible vLLM endpoint. The model emits RAIF. The plugin decodes it at the server boundary and returns JSON to existing clients.
Scope
RAIF covers one JSON object produced by a language model. It supports strings, numbers, booleans, null values, arrays, and nested objects.
It is not a general interchange format. It is not compression, a schema language, or an input format for a model. Its job is to make structured model output smaller and easier to recover.
Resources
- raif-standard has the specification, decisions, and reference libraries.
- raif-lora has the adapter training work.
- raif-vllm has the vLLM plugin.
- Packages are available on npm, PyPI, and PyPI for vLLM.
- Models are available for Llama 3.2 3B, Qwen 3 4B, and Qwen 2.5 0.5B.
RAIF uses the Apache 2.0 license. It includes a patent grant for anyone who implements or builds on the format.