Note: the judge model is also one of the contestants
— scores may carry self-preference bias.
🏆 Quality winner: openrouter:deepseek/deepseek-v4-pro
DeepSeek-v4-pro (2eb734) delivers the best-balanced package: a fully complete, polished rewritten passage that is faithful to the source; a precise Cold War Tu-4/B-29 reverse-engineering analogy that genuinely shares distillation's 'closes the gap but never matches the original' mechanism; a full 10-question SAT-style quiz with clean evidence-pairing and no invented facts; and a strong glossary. Kimi (e2507b) and Sonnet (fe7a5d) are very close runners-up, but Sonnet invents specific statistics (16M exchanges, 24,000 accounts) not present in the source, and Kimi's passage—while excellent—is marginally less accessible. 2eb734 avoids fabrication while remaining engaging and complete, edging out the field.
💰 Best value: openrouter:deepseek/deepseek-v4-pro
— quality 9/10 at 0.62¢ → value 8.7/10
Value = 0.7·quality + 0.3·cheapness. Quality stays dominant, so a cheap-but-weak output can't win on price alone.
| # | Model | Facts | Analogy | Age-fit | Clarity | Quiz | Glossary | Passage | Overall | Cost | Value | Notes |
| 🥇 |
anthropic:claude-sonnet-4-6
view ↗ |
8 | 9 | 9 | 9 | 9 | 9 | 8 |
9 |
8.57¢ |
6.6 |
Excellent analogies (chef, band recordings) sharing distillation's mechanism. Strong 10-question SAT set with evidence pairing. Invents specific figures (16M exchanges, 24,000 accounts) not in source, a fact-grounding lapse. Passage truncated but strong. |
| 🥈 |
openrouter:deepseek/deepseek-v4-pro
view ↗ |
9 | 8 | 9 | 9 | 9 | 9 | 9 |
9 |
0.62¢ |
8.7 |
Faithful to source, complete polished passage. Cold War Tu-4/B-29 analogy is precise and educational. Full 10-question quiz with solid evidence pairing. Balanced and well-organized throughout. |
| 🥉 |
openrouter:moonshotai/kimi-k2.6
view ↗ |
9 | 8 | 9 | 9 | 9 | 8 | 9 |
9 |
9.62¢ |
6.6 |
Highly faithful, sophisticated passage incorporating McGuire quote accurately. Conservatory/pharma analogies capture the open-info-harvested-at-scale mechanism. Excellent quiz with careful evidence pairing. Well-calibrated prose. |
| 4 |
openrouter:mistralai/mistral-large-2512
view ↗ |
8 | 7 | 9 | 8 | 8 | 8 | 9 |
8 |
0.85¢ |
8.0 |
Complete, accurate passage. 'Save scumming' cheat-code analogy is a bit loose; marathon analogy fine. Includes vocab quiz questions on 'surreptitious' which appears in source. Good glossary but only 5 terms. |
| 5 |
openrouter:minimax/minimax-m2.7
view ↗ |
9 | 8 | 9 | 9 | 8 | 8 | 9 |
8 |
0.62¢ |
8.0 |
Very faithful, clean passage. University/professor analogy solid. Quiz well-constructed with fair evidence pairing. Splinternet framing adds value. Minor: 'leverage' vocab question is a touch trivial. |
| 6 |
openrouter:stepfun/step-3.7-flash
view ↗ |
7 | 8 | 9 | 9 | 9 | 8 | 8 |
8 |
2.09¢ |
6.8 |
Strong quiz with explicit trap taxonomy and SAT tips. Analogies (SAT tutor, fast-food) fit mechanism. Overstates 'state-linked' when source says 'principally based in China'; adds specifics not in source. Passage truncated. |
| 7 |
openrouter:z-ai/glm-5-turbo
view ↗ |
7 | 8 | 9 | 9 | 8 | 8 | 8 |
8 |
3.40¢ |
6.8 |
Strong quiz with trap taxonomy. However quiz asks vocab on 'sharp' which appears in its own rewritten passage not the source, and some evidence quotes are paraphrased/invented ('close the competitive gap...partly because'). Passage truncated. Good analogies. |
| 8 |
openrouter:openai/gpt-5.4-nano
view ↗ |
8 | 8 | 9 | 9 | 8 | 8 | 8 |
8 |
0.63¢ |
8.0 |
Thoughtful passage with textbook/study-guide analogy that fits well. Quiz solid with good evidence pairing and honest handling of what source does/doesn't say. Complete and well-calibrated. Slightly repetitive framing. |
| 9 |
anthropic:claude-haiku-4-5-20251001
view ↗ |
8 | 8 | 9 | 9 | 8 | 9 | 8 |
8 |
2.57¢ |
6.8 |
Strong nuclear-arms-race and doping-scandal analogies. Detailed glossary (6 terms). Full quiz with good evidence pairing. Accurate. Passage truncated but engaging. Solid all-around. |
| 10 |
openrouter:google/gemini-3.1-flash-lite
view ↗ |
8 | 7 | 8 | 8 | 8 | 7 | 8 |
7 |
0.59¢ |
7.3 |
Solid, complete passage and full quiz. Chef/capture-the-flag analogies decent. Glossary 4 terms. Competent throughout but less distinctive than top entries. |
| 11 |
openrouter:x-ai/grok-4.3
view ↗ |
8 | 6 | 8 | 8 | 7 | 8 | 8 |
7 |
1.64¢ |
6.7 |
Accurate, complete passage. Analogies (test copying, scouting playbook) somewhat generic. Full quiz reasonable but a couple questions overlap. Competent, mid-tier execution. |
| 12 |
openrouter:qwen/qwen3-max-thinking
view ↗ |
8 | 8 | 8 | 7 | 4 | 8 | 8 |
6 |
1.02¢ |
6.0 |
Good passage and analogies (answer key, counterfeit handbags). Fatal weakness: only ONE quiz question despite the rubric expecting a full set. Otherwise solid but incomplete deliverable. |
| 13 |
openrouter:amazon/nova-pro-v1
view ↗ |
8 | 5 | 7 | 7 | 7 | 6 | 7 |
6 |
1.34¢ |
6.0 |
Accurate but thin. Chess/doping analogies are generic and don't capture distillation's mechanism. Hook is one flat sentence. Glossary only 3 terms. Quiz complete and adequate but unremarkable. |
| 14 |
openrouter:meta-llama/llama-4-maverick
view ↗ |
7 | 6 | 7 | 7 | 6 | 6 | 6 |
6 |
0.26¢ |
7.2 |
Adequate but bland. Analogies are homework-copying and PED sports—generic. Passage accurate but plain. Quiz fair but some questions verge on trivial. Glossary 3 terms only. |
| 15 |
openrouter:bytedance-seed/seed-2.0-lite
view ↗ |
7 | 6 | 8 | 8 | 6 | 8 | 7 |
6 |
1.73¢ |
6.0 |
Complete deliverable. Quiz asks vocab on 'principal' and 'surreptitious' but 'principal' isn't in the passage text as used (source says 'principally'). Analogies (MIT exam key, Twitch scraping) fine. Some overreach in framing '$1 trillion industry' unsourced. |
| 16 |
openrouter:baidu/ernie-4.5-vl-424b-a47b
view ↗ |
7 | 6 | 7 | 7 | 5 | 6 | 7 |
5 |
0.43¢ |
6.5 |
Compact and accurate but only 3 quiz questions and 3 glossary terms—under-delivered. Analogies (chef recipe, counterfeit goods) generic. Passage is short but coherent. Weakest complete entry due to thin quiz. |
Scores 0–10 · 🟩 ≥8 · 🟨 4–7 · 🟥 <4 · sorted by Overall. Hover a column header for its full rubric name.
| # | Model | Quality | Cost | Value |
| 🥇 |
openrouter:deepseek/deepseek-v4-pro |
9 |
0.62¢ |
8.7 |
| 🥈 |
openrouter:mistralai/mistral-large-2512 |
8 |
0.85¢ |
8.0 |
| 🥉 |
openrouter:minimax/minimax-m2.7 |
8 |
0.62¢ |
8.0 |
| 4 |
openrouter:openai/gpt-5.4-nano |
8 |
0.63¢ |
8.0 |
| 5 |
openrouter:google/gemini-3.1-flash-lite |
7 |
0.59¢ |
7.3 |
| 6 |
openrouter:meta-llama/llama-4-maverick |
6 |
0.26¢ |
7.2 |
| 7 |
openrouter:stepfun/step-3.7-flash |
8 |
2.09¢ |
6.8 |
| 8 |
openrouter:z-ai/glm-5-turbo |
8 |
3.40¢ |
6.8 |
| 9 |
anthropic:claude-haiku-4-5-20251001 |
8 |
2.57¢ |
6.8 |
| 10 |
openrouter:x-ai/grok-4.3 |
7 |
1.64¢ |
6.7 |
| 11 |
anthropic:claude-sonnet-4-6 |
9 |
8.57¢ |
6.6 |
| 12 |
openrouter:moonshotai/kimi-k2.6 |
9 |
9.62¢ |
6.6 |
| 13 |
openrouter:baidu/ernie-4.5-vl-424b-a47b |
5 |
0.43¢ |
6.5 |
| 14 |
openrouter:qwen/qwen3-max-thinking |
6 |
1.02¢ |
6.0 |
| 15 |
openrouter:amazon/nova-pro-v1 |
6 |
1.34¢ |
6.0 |
| 16 |
openrouter:bytedance-seed/seed-2.0-lite |
6 |
1.73¢ |
6.0 |
| # | Model | Quality | Cost | Value |
| 1 |
openrouter:deepseek/deepseek-v4-pro |
9 |
0.62¢ |
8.7 |
| 2 |
openrouter:mistralai/mistral-large-2512 |
8 |
0.85¢ |
8.0 |
| 3 |
openrouter:minimax/minimax-m2.7 |
8 |
0.62¢ |
8.0 |
| 4 |
openrouter:openai/gpt-5.4-nano |
8 |
0.63¢ |
8.0 |
| 5 |
openrouter:stepfun/step-3.7-flash |
8 |
2.09¢ |
6.8 |
| 6 |
anthropic:claude-haiku-4-5-20251001 |
8 |
2.57¢ |
6.8 |
| 7 |
openrouter:google/gemini-3.1-flash-lite |
7 |
0.59¢ |
7.3 |
| 8 |
openrouter:x-ai/grok-4.3 |
7 |
1.64¢ |
6.7 |
| 9 |
openrouter:meta-llama/llama-4-maverick |
6 |
0.26¢ |
7.2 |
| 10 |
openrouter:qwen/qwen3-max-thinking |
6 |
1.02¢ |
6.0 |
| 11 |
openrouter:amazon/nova-pro-v1 |
6 |
1.34¢ |
6.0 |
| 12 |
openrouter:bytedance-seed/seed-2.0-lite |
6 |
1.73¢ |
6.0 |
| 13 |
openrouter:baidu/ernie-4.5-vl-424b-a47b |
5 |
0.43¢ |
6.5 |
Strong quality without paying flagship prices — the cheap-and-good picks.