docs/guide/benchmarks.md
The benchmarks on this page measure TOON's performance across two key dimensions:
Benchmarks are organized into two tracks to ensure fair comparisons:
Benchmarks test LLM comprehension across different input formats using 244 data retrieval questions on 4 models.
<details> <summary><strong>Show Dataset Catalog</strong></summary>| Dataset | Rows | Structure | CSV Support | Eligibility |
|---|---|---|---|---|
| Uniform employee records | 100 | uniform | โ | 100% |
| E-commerce orders with nested structures | 50 | nested | โ | 33% |
| Time-series analytics data | 60 | uniform | โ | 100% |
| Top 100 GitHub repositories | 100 | uniform | โ | 100% |
| Semi-uniform event logs | 75 | semi-uniform | โ | 50% |
| Deeply nested configuration | 1 | deep | โ | 0% |
| Valid complete dataset (control) | 20 | uniform | โ | 100% |
| Array truncated: 3 rows removed from end | 20 | uniform | โ | 100% |
| Extra rows added beyond declared length | 20 | uniform | โ | 100% |
| Inconsistent field count (missing salary in row 10) | 20 | uniform | โ | 100% |
| Missing required fields (no email in multiple rows) | 20 | uniform | โ | 100% |
| Feature flags keyed by name | 40 | uniform | โ | 100% |
| Contacts with nested address and plan groups | 50 | nested | โ | 100% |
Structure classes:
CSV Support: โ (supported), โ (not supported โ would require lossy flattening)
Eligibility: Percentage of arrays that qualify for TOON's tabular form (uniform objects with primitive values)
</details>Each format ranked by efficiency (accuracy percentage per 1,000 tokens):
TOON โโโโโโโโโโโโโโโโโโโโ 29.2 acc%/1K tok โ 72.2% ยฑ2.8 acc โ 2,474 tokens
JSON compact โโโโโโโโโโโโโโโโโโโโ 23.8 acc%/1K tok โ 69.0% ยฑ2.9 acc โ 2,892 tokens
YAML โโโโโโโโโโโโโโโโโโโโ 20.1 acc%/1K tok โ 70.1% ยฑ2.9 acc โ 3,487 tokens
JSON โโโโโโโโโโโโโโโโโโโโ 16.6 acc%/1K tok โ 71.4% ยฑ2.8 acc โ 4,308 tokens
XML โโโโโโโโโโโโโโโโโโโโ 14.4 acc%/1K tok โ 70.7% ยฑ2.9 acc โ 4,909 tokens
Efficiency score = (Accuracy % รท Tokens) ร 1,000. Higher is better.
[!TIP] TOON achieves 72.2% accuracy (vs JSON's 71.4%) while using 42.6% fewer tokens.
[!NOTE] CSV is excluded from the ranking as it only supports 109 of 244 questions (flat tabular data only). While CSV is highly token-efficient for simple tabular data, it cannot represent nested structures that other formats handle.
Every format answers the same 109 flat-dataset questions per model, so CSV can be compared on equal footing here.
| Format | Accuracy | Correct/Total | Avg Tokens |
|---|---|---|---|
toon | 63.1% ยฑ4.5 | 275/436 | 1,994 |
csv | 62.2% ยฑ4.5 | 271/436 | 1,851 |
json-pretty | 60.3% ยฑ4.6 | 263/436 | 3,950 |
xml | 60.1% ยฑ4.6 | 262/436 | 4,516 |
yaml | 59.9% ยฑ4.6 | 261/436 | 3,270 |
json-compact | 58.0% ยฑ4.6 | 253/436 | 2,718 |
Accuracy across 4 LLMs on 244 data retrieval questions:
claude-haiku-4-5-20251001
โ TOON โโโโโโโโโโโโโโโโโโโโ 65.6% ยฑ5.9 (160/244)
JSON โโโโโโโโโโโโโโโโโโโโ 63.5% ยฑ6.0 (155/244)
XML โโโโโโโโโโโโโโโโโโโโ 62.3% ยฑ6.0 (152/244)
YAML โโโโโโโโโโโโโโโโโโโโ 62.3% ยฑ6.0 (152/244)
JSON compact โโโโโโโโโโโโโโโโโโโโ 61.9% ยฑ6.0 (151/244)
CSV โโโโโโโโโโโโโโโโโโโโ 49.5% ยฑ9.2 (54/109)
gemini-3.6-flash
โ TOON โโโโโโโโโโโโโโโโโโโโ 69.3% ยฑ5.8 (169/244)
JSON โโโโโโโโโโโโโโโโโโโโ 68.4% ยฑ5.8 (167/244)
YAML โโโโโโโโโโโโโโโโโโโโ 67.6% ยฑ5.8 (165/244)
XML โโโโโโโโโโโโโโโโโโโโ 65.2% ยฑ5.9 (159/244)
JSON compact โโโโโโโโโโโโโโโโโโโโ 63.5% ยฑ6.0 (155/244)
CSV โโโโโโโโโโโโโโโโโโโโ 57.8% ยฑ9.1 (63/109)
gpt-5.4-nano
XML โโโโโโโโโโโโโโโโโโโโ 59.4% ยฑ6.1 (145/244)
JSON โโโโโโโโโโโโโโโโโโโโ 57.4% ยฑ6.2 (140/244)
โ TOON โโโโโโโโโโโโโโโโโโโโ 57.0% ยฑ6.2 (139/244)
JSON compact โโโโโโโโโโโโโโโโโโโโ 54.9% ยฑ6.2 (134/244)
YAML โโโโโโโโโโโโโโโโโโโโ 54.5% ยฑ6.2 (133/244)
CSV โโโโโโโโโโโโโโโโโโโโ 46.8% ยฑ9.2 (51/109)
grok-4.5
โ TOON โโโโโโโโโโโโโโโโโโโโ 97.1% ยฑ2.2 (237/244)
JSON โโโโโโโโโโโโโโโโโโโโ 96.3% ยฑ2.5 (235/244)
XML โโโโโโโโโโโโโโโโโโโโ 95.9% ยฑ2.6 (234/244)
YAML โโโโโโโโโโโโโโโโโโโโ 95.9% ยฑ2.6 (234/244)
JSON compact โโโโโโโโโโโโโโโโโโโโ 95.5% ยฑ2.7 (233/244)
CSV โโโโโโโโโโโโโโโโโโโโ 94.5% ยฑ4.5 (103/109)
<details> <summary><strong>Performance by dataset and question type</strong></summary>[!NOTE] Accuracy figures include Wilson 95% confidence intervals (ยฑ); when two formats' intervals overlap, the difference between them is not statistically meaningful. CSV answers only the 109 flat-dataset questions, so its per-model cells cover a smaller, easier population than the other formats.
| Question Type | TOON | JSON | XML | YAML | JSON compact | CSV |
|---|---|---|---|---|---|---|
| Field Retrieval | 97.8% | 99.2% | 99.2% | 99.7% | 98.9% | 100.0% |
| Aggregation | 48.4% | 48.4% | 46.0% | 46.0% | 45.2% | 32.8% |
| Filtering | 38.0% | 41.1% | 37.5% | 40.1% | 38.0% | 33.3% |
| Structure Awareness | 90.3% | 84.0% | 84.0% | 79.2% | 78.5% | 82.8% |
| Structural Validation | 100.0% | 50.0% | 80.0% | 50.0% | 45.0% | 80.0% |
| Format | Accuracy | Tokens | Correct/Total |
|---|---|---|---|
csv | 64.6% | 2,336 | 106/164 |
toon | 62.8% | 2,537 | 103/164 |
json-compact | 62.2% | 3,919 | 102/164 |
yaml | 64.0% | 4,982 | 105/164 |
json-pretty | 62.2% | 6,326 | 102/164 |
xml | 61.0% | 7,286 | 100/164 |
| Format | Accuracy | Tokens | Correct/Total |
|---|---|---|---|
json-compact | 70.7% | 6,875 | 116/164 |
toon | 71.3% | 7,344 | 117/164 |
yaml | 72.0% | 8,456 | 118/164 |
json-pretty | 71.3% | 10,842 | 117/164 |
xml | 74.4% | 12,180 | 122/164 |
| Format | Accuracy | Tokens | Correct/Total |
|---|---|---|---|
csv | 64.2% | 1,408 | 77/120 |
toon | 63.3% | 1,595 | 76/120 |
json-compact | 59.2% | 2,351 | 71/120 |
yaml | 62.5% | 2,951 | 75/120 |
json-pretty | 65.0% | 3,678 | 78/120 |
xml | 62.5% | 4,386 | 75/120 |
| Format | Accuracy | Tokens | Correct/Total |
|---|---|---|---|
toon | 57.6% | 9,017 | 76/132 |
csv | 54.5% | 8,726 | 72/132 |
json-compact | 53.8% | 11,650 | 71/132 |
yaml | 53.8% | 13,350 | 71/132 |
json-pretty | 55.3% | 15,350 | 73/132 |
xml | 53.8% | 17,304 | 71/132 |
| Format | Accuracy | Tokens | Correct/Total |
|---|---|---|---|
json-compact | 56.7% | 4,793 | 68/120 |
toon | 60.8% | 5,814 | 73/120 |
json-pretty | 60.0% | 6,759 | 72/120 |
yaml | 55.0% | 5,798 | 66/120 |
xml | 50.8% | 7,668 | 61/120 |
| Format | Accuracy | Tokens | Correct/Total |
|---|---|---|---|
json-compact | 91.4% | 562 | 106/116 |
yaml | 93.1% | 675 | 108/116 |
toon | 91.4% | 669 | 106/116 |
json-pretty | 94.8% | 918 | 110/116 |
xml | 94.0% | 1,007 | 109/116 |
| Format | Accuracy | Tokens | Correct/Total |
|---|---|---|---|
toon | 100.0% | 566 | 4/4 |
json-compact | 100.0% | 772 | 4/4 |
yaml | 100.0% | 984 | 4/4 |
json-pretty | 100.0% | 1,259 | 4/4 |
xml | 0.0% | 1,441 | 0/4 |
csv | 0.0% | 473 | 0/4 |
| Format | Accuracy | Tokens | Correct/Total |
|---|---|---|---|
csv | 100.0% | 408 | 4/4 |
toon | 100.0% | 498 | 4/4 |
xml | 100.0% | 1,229 | 4/4 |
json-pretty | 0.0% | 1,075 | 0/4 |
yaml | 0.0% | 841 | 0/4 |
json-compact | 0.0% | 660 | 0/4 |
| Format | Accuracy | Tokens | Correct/Total |
|---|---|---|---|
csv | 100.0% | 547 | 4/4 |
toon | 100.0% | 644 | 4/4 |
xml | 100.0% | 1,663 | 4/4 |
json-pretty | 0.0% | 1,452 | 0/4 |
yaml | 0.0% | 1,135 | 0/4 |
json-compact | 0.0% | 893 | 0/4 |
| Format | Accuracy | Tokens | Correct/Total |
|---|---|---|---|
csv | 100.0% | 470 | 4/4 |
toon | 100.0% | 563 | 4/4 |
json-compact | 75.0% | 767 | 3/4 |
xml | 100.0% | 1,432 | 4/4 |
yaml | 75.0% | 977 | 3/4 |
json-pretty | 75.0% | 1,251 | 3/4 |
| Format | Accuracy | Tokens | Correct/Total |
|---|---|---|---|
csv | 100.0% | 442 | 4/4 |
toon | 100.0% | 535 | 4/4 |
xml | 100.0% | 1,386 | 4/4 |
yaml | 75.0% | 941 | 3/4 |
json-pretty | 75.0% | 1,207 | 3/4 |
json-compact | 50.0% | 732 | 2/4 |
| Format | Accuracy | Tokens | Correct/Total |
|---|---|---|---|
toon | 97.1% | 931 | 66/68 |
json-compact | 94.1% | 1,264 | 64/68 |
yaml | 92.6% | 1,443 | 63/68 |
json-pretty | 95.6% | 1,873 | 65/68 |
xml | 95.6% | 2,306 | 65/68 |
| Format | Accuracy | Tokens | Correct/Total |
|---|---|---|---|
toon | 94.4% | 1,444 | 68/72 |
json-compact | 91.7% | 2,357 | 66/72 |
yaml | 94.4% | 2,797 | 68/72 |
json-pretty | 97.2% | 4,014 | 70/72 |
xml | 98.6% | 4,534 | 71/72 |
claude-haiku-4-5-20251001, gemini-3.6-flash, gpt-5.4-nano, grok-4.5gpt-tokenizer with o200k_base encoding (GPT-5 tokenizer). Other providers tokenize differently, so absolute counts are tokenizer-specific; relative differences between formats hold directionally.reasoning: 'none' (Gemini 3 floors at minimal thinking, grok-4.5 at low)What the datasets contain, how the questions are generated, and how answers are validated is documented in the benchmark README.
<!-- /automd -->Token counts are measured using the GPT-5 o200k_base tokenizer via gpt-tokenizer. Savings are calculated against formatted JSON (2-space indentation) as the primary baseline, with additional comparisons to compact JSON (minified), YAML, and XML. Actual savings vary by model and tokenizer.
The benchmarks test datasets across different structural patterns (uniform, semi-uniform, nested, deeply nested) to show where TOON excels and where other formats may be better.
<!-- automd:file src="../../benchmarks/results/token-efficiency.md" -->Datasets with nested or semi-uniform structures. CSV excluded as it cannot properly represent these structures.
๐ E-commerce orders with nested structures โ Tabular: 33%
โ
TOON โโโโโโโโโโโโโโโโโโโโ 72,832 tokens
โโ vs JSON (โ32.9%) 108,611 tokens
โโ vs JSON compact (+5.6%) 68,944 tokens
โโ vs YAML (โ14.0%) 84,701 tokens
โโ vs XML (โ40.4%) 122,119 tokens
๐งพ Semi-uniform event logs โ Tabular: 50%
โ
TOON โโโโโโโโโโโโโโโโโโโโ 154,084 tokens
โโ vs JSON (โ15.0%) 181,201 tokens
โโ vs JSON compact (+19.9%) 128,529 tokens
โโ vs YAML (โ0.8%) 155,397 tokens
โโ vs XML (โ25.2%) 205,859 tokens
๐งฉ Deeply nested configuration โ Tabular: 0%
โ
TOON โโโโโโโโโโโโโโโโโโโโ 589 tokens
โโ vs JSON (โ34.9%) 905 tokens
โโ vs JSON compact (+6.7%) 552 tokens
โโ vs YAML (โ11.0%) 662 tokens
โโ vs XML (โ40.9%) 997 tokens
๐ Feature flags keyed by name โ Tabular: 100%
โ
TOON โโโโโโโโโโโโโโโโโโโโ 10,503 tokens
โโ vs JSON (โ54.6%) 23,141 tokens
โโ vs JSON compact (โ32.8%) 15,635 tokens
โโ vs YAML (โ41.3%) 17,905 tokens
โโ vs XML (โ63.3%) 28,655 tokens
๐ Contacts with nested address and plan groups โ Tabular: 100%
โ
TOON โโโโโโโโโโโโโโโโโโโโ 26,726 tokens
โโ vs JSON (โ66.5%) 79,779 tokens
โโ vs JSON compact (โ42.9%) 46,791 tokens
โโ vs YAML (โ51.8%) 55,475 tokens
โโ vs XML (โ70.4%) 90,306 tokens
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Total โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
TOON โโโโโโโโโโโโโโโโโโโโ 264,734 tokens
โโ vs JSON (โ32.7%) 393,637 tokens
โโ vs JSON compact (+1.6%) 260,451 tokens
โโ vs YAML (โ15.7%) 314,140 tokens
โโ vs XML (โ40.9%) 447,936 tokens
Datasets with flat, fully tabular-eligible data where CSV is applicable.
๐ฅ Uniform employee records โ Tabular: 100%
โ
CSV โโโโโโโโโโโโโโโโโโโโ 47,153 tokens
TOON โโโโโโโโโโโโโโโโโโโโ 49,978 tokens (+6.0% vs CSV)
โโ vs JSON (โ60.7%) 127,061 tokens
โโ vs JSON compact (โ36.8%) 79,057 tokens
โโ vs YAML (โ50.0%) 100,054 tokens
โโ vs XML (โ65.9%) 146,605 tokens
๐ Time-series analytics data โ Tabular: 100%
โ
CSV โโโโโโโโโโโโโโโโโโโโ 8,383 tokens
TOON โโโโโโโโโโโโโโโโโโโโ 9,115 tokens (+8.7% vs CSV)
โโ vs JSON (โ59.0%) 22,245 tokens
โโ vs JSON compact (โ35.9%) 14,211 tokens
โโ vs YAML (โ49.0%) 17,858 tokens
โโ vs XML (โ65.8%) 26,616 tokens
โญ Top 100 GitHub repositories โ Tabular: 100%
โ
CSV โโโโโโโโโโโโโโโโโโโโ 8,711 tokens
TOON โโโโโโโโโโโโโโโโโโโโ 8,937 tokens (+2.6% vs CSV)
โโ vs JSON (โ41.7%) 15,337 tokens
โโ vs JSON compact (โ23.2%) 11,640 tokens
โโ vs YAML (โ33.0%) 13,337 tokens
โโ vs XML (โ48.3%) 17,294 tokens
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Total โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
CSV โโโโโโโโโโโโโโโโโโโโ 64,247 tokens
TOON โโโโโโโโโโโโโโโโโโโโ 68,030 tokens (+5.9% vs CSV)
โโ vs JSON (โ58.7%) 164,643 tokens
โโ vs JSON compact (โ35.2%) 104,908 tokens
โโ vs YAML (โ48.2%) 131,249 tokens
โโ vs XML (โ64.3%) 190,515 tokens
Token counts use gpt-tokenizer with o200k_base encoding (GPT-5 tokenizer). Other providers tokenize differently, so absolute counts are tokenizer-specific; relative differences between formats hold directionally.
Savings: 13,130 tokens (59.0% reduction vs JSON)
JSON (22,245 tokens):
{
"metrics": [
{
"date": "2025-01-01",
"views": 6138,
"clicks": 174,
"conversions": 12,
"revenue": 2712.49,
"bounceRate": 0.35
},
{
"date": "2025-01-02",
"views": 4616,
"clicks": 274,
"conversions": 34,
"revenue": 9156.29,
"bounceRate": 0.56
},
{
"date": "2025-01-03",
"views": 4460,
"clicks": 143,
"conversions": 8,
"revenue": 1317.98,
"bounceRate": 0.59
},
{
"date": "2025-01-04",
"views": 4740,
"clicks": 125,
"conversions": 13,
"revenue": 2934.77,
"bounceRate": 0.37
},
{
"date": "2025-01-05",
"views": 6428,
"clicks": 369,
"conversions": 19,
"revenue": 1317.24,
"bounceRate": 0.3
}
]
}
TOON (9,115 tokens):
metrics[5]{date,views,clicks,conversions,revenue,bounceRate}:
2025-01-01,6138,174,12,2712.49,0.35
2025-01-02,4616,274,34,9156.29,0.56
2025-01-03,4460,143,8,1317.98,0.59
2025-01-04,4740,125,13,2934.77,0.37
2025-01-05,6428,369,19,1317.24,0.3
Savings: 6,400 tokens (41.7% reduction vs JSON)
JSON (15,337 tokens):
{
"repositories": [
{
"id": 132750724,
"name": "build-your-own-x",
"repo": "codecrafters-io/build-your-own-x",
"description": "Master programming by recreating your favorite technologies from scratch.",
"createdAt": "2018-05-09T12:03:18Z",
"updatedAt": "2026-07-23T18:57:15Z",
"pushedAt": "2026-07-14T19:25:58Z",
"stars": 530712,
"watchers": 6778,
"forks": 50205,
"defaultBranch": "master"
},
{
"id": 21737465,
"name": "awesome",
"repo": "sindresorhus/awesome",
"description": "๐ Awesome lists about all kinds of interesting topics",
"createdAt": "2014-07-11T13:42:37Z",
"updatedAt": "2026-07-23T18:57:24Z",
"pushedAt": "2026-06-30T18:21:16Z",
"stars": 488074,
"watchers": 8292,
"forks": 36010,
"defaultBranch": "main"
},
{
"id": 28457823,
"name": "freeCodeCamp",
"repo": "freeCodeCamp/freeCodeCamp",
"description": "freeCodeCamp.org's open-source codebase and curriculum. Learn math, programming,โฆ",
"createdAt": "2014-12-24T17:49:19Z",
"updatedAt": "2026-07-22T07:01:33Z",
"pushedAt": "2026-07-21T18:00:51Z",
"stars": 452380,
"watchers": 8590,
"forks": 45624,
"defaultBranch": "main"
}
]
}
TOON (8,937 tokens):
repositories[3]{id,name,repo,description,createdAt,updatedAt,pushedAt,stars,watchers,forks,defaultBranch}:
132750724,build-your-own-x,codecrafters-io/build-your-own-x,Master programming by recreating your favorite technologies from scratch.,"2018-05-09T12:03:18Z","2026-07-23T18:57:15Z","2026-07-14T19:25:58Z",530712,6778,50205,master
21737465,awesome,sindresorhus/awesome,๐ Awesome lists about all kinds of interesting topics,"2014-07-11T13:42:37Z","2026-07-23T18:57:24Z","2026-06-30T18:21:16Z",488074,8292,36010,main
28457823,freeCodeCamp,freeCodeCamp/freeCodeCamp,"freeCodeCamp.org's open-source codebase and curriculum. Learn math, programming,โฆ","2014-12-24T17:49:19Z","2026-07-22T07:01:33Z","2026-07-21T18:00:51Z",452380,8590,45624,main