Note: the judge model is also one of the contestants
— scores may carry self-preference bias.
🏆 Quality winner: anthropic:claude-opus-4-8
Run 91e0a8 (Opus) wins on the combination of flawless fact-grounding (correctly stresses distilled models are weaker yet still cheap), mechanism-true analogies (tasting a dish thousands of times; memorizing a grandmaster's recorded games), a complete and genuinely SAT-caliber rewritten passage that closes on a real intellectual open question, and a ten-item quiz with precise trap taxonomy and clean evidence-pairing tied to lines actually in its passage. Runners-up e2507b and 2eb734 are nearly as strong—e2507b has superb prose and precise conservatory/pharma analogies, 2eb734 has the standout Tu-4/B-29 reverse-engineering analogy—but 91e0a8 edges them on passage quality and consistency. Candidates that fabricated figures (fe7a5d's 16M/24k numbers, 0313fc's $1T, a95998's 'state-linked') or under-delivered quizzes (f95e70 with one question, a0132a and a-vl variant with three) rank lower.
| # | Model | Facts | Analogy | Age-fit | Clarity | Quiz | Glossary | Passage | Overall | Cost | Value | Notes |
| 🥇 |
anthropic:claude-sonnet-4-6
view ↗ |
8 | 9 | 9 | 9 | 9 | 9 | 8 |
9 |
8.57¢ |
6.6 |
Excellent depth with strong same-mechanism analogies (chef ordering every dish, recording performances). Invents specific figures (16 million exchanges, 24,000 accounts) not in source—a fact-grounding risk, though plausible. Ten well-crafted SAT-style questions with evidence pairs; passage truncated but engaging. |
| 🥈 |
openrouter:deepseek/deepseek-v4-pro
view ↗ |
9 | 8 | 9 | 9 | 9 | 8 | 9 |
9 |
0.62¢ |
8.7 |
Faithful, well-organized, with a complete polished rewritten passage. Tu-4/B-29 Cold War analogy is precise and shares the reverse-engineering mechanism. Strong 10-question quiz with good vocab-in-context and evidence pairing. Minor: quiz asks about 'How To Think About It' tone, relying on the rewritten sections. |
| 🥉 |
openrouter:minimax/minimax-m2.7
view ↗ |
9 | 8 | 9 | 9 | 8 | 8 | 9 |
9 |
0.62¢ |
8.7 |
Very faithful, careful to note distilled models don't match originals. University/answer-key analogy is apt. Complete strong passage. Quiz Q4 asks about 'leverage' which appears in source—good. Well-balanced and accurate throughout with clear structure. |
| 4 |
openrouter:moonshotai/kimi-k2.6
view ↗ |
9 | 9 | 9 | 9 | 9 | 8 | 9 |
9 |
9.62¢ |
6.6 |
Highly faithful, sophisticated prose pitched perfectly for older teens. Conservatory/pharmaceutical analogies precisely share the incentive-structure mechanism. Complete, richly written passage that integrates McGuire quotes accurately. Excellent 10-question quiz with careful evidence pairing. One of the strongest. |
| 5 |
anthropic:claude-opus-4-8
view ↗ |
9 | 9 | 9 | 9 | 9 | 9 | 10 |
9 |
15.59¢ |
6.6 |
Outstanding across the board: faithful, precise about distilled models being weaker, complete and beautifully written SAT-style passage ending on a genuine open question. Recipe-tasting and chess-game-memorizing analogies match the mechanism. Ten sharp questions with excellent trap explanations and evidence pairing. Best overall. |
| 6 |
openrouter:mistralai/mistral-large-2512
view ↗ |
8 | 7 | 8 | 8 | 8 | 8 | 8 |
8 |
0.85¢ |
8.0 |
Solid, faithful coverage with the marathon-sweatband analogy (decent mechanism match). Adds some speculative claims (US firms pulling back from collaborations) but flags them as bigger-picture. Complete passage, competent 10-question quiz. Some analogies (cheat code save-scumming) don't share the distillation mechanism well. |
| 7 |
openrouter:z-ai/glm-5-turbo
view ↗ |
9 | 8 | 9 | 9 | 8 | 8 | 9 |
8 |
3.40¢ |
6.8 |
Very faithful with strong music-producer/scouting analogies and detailed trap taxonomy in quiz explanations. However, quiz Q4 asks about 'sharp' and Q9/Q10 reference lines ('close the competitive gap...partly because of export controls') that appear in its own rewritten passage, not verbatim source—slight self-referential risk. Passage truncated but excellent. |
| 8 |
openrouter:openai/gpt-5.4-nano
view ↗ |
9 | 7 | 8 | 8 | 8 | 8 | 8 |
8 |
0.63¢ |
8.0 |
Notably careful and faithful, repeatedly attributing claims to 'the article' and preserving nuance (distilled models lag but still shift power). Export-controls-as-lock analogy is apt. Complete passage, strong 10-question quiz with good evidence pairing. Analogies solid though the sports one is a touch generic. |
| 9 |
openrouter:stepfun/step-3.7-flash
view ↗ |
6 | 8 | 8 | 8 | 8 | 8 | 7 |
7 |
2.09¢ |
6.1 |
Good analogies and structure but overstates facts: calls Chinese groups 'state-linked' and says memo 'specifically names three firms' (Anthropic named them, not the memo). Invents specific dates. Quiz is decent with SAT-style traps but the passage is truncated. Fact drift lowers overall. |
| 10 |
openrouter:google/gemini-3.1-flash-lite
view ↗ |
8 | 7 | 8 | 8 | 7 | 7 | 8 |
7 |
0.59¢ |
7.3 |
Solid, accurate, complete passage with chef/recipe and capture-the-flag analogies. Quiz Q5/Q6 on author tone rely on the rewritten passage rather than source, slightly weakening fairness. Only 4 glossary terms. Competent but not standout. |
| 11 |
openrouter:x-ai/grok-4.3
view ↗ |
8 | 6 | 8 | 8 | 7 | 8 | 8 |
7 |
1.64¢ |
6.7 |
Faithful and clean but analogies (reverse-engineering, hidden-camera scouting) are somewhat generic. Sections are brief. Complete passage. Quiz Q5 asks about 'industrial-scale' meaning—fair—but overall quiz is competent rather than distinguished. Solid mid-tier entry. |
| 12 |
openrouter:meta-llama/llama-4-maverick
view ↗ |
8 | 6 | 8 | 8 | 7 | 8 | 8 |
7 |
0.26¢ |
7.9 |
Careful, faithful reporting that appropriately hedges ('the article states'). Analogies (textbook/study-guide, film scouting) are okay but not especially mechanism-tight. Good glossary. Quiz is fair with vocab-in-context and evidence pairing, but slightly less punchy than the top tier. Complete passage. |
| 13 |
openrouter:qwen/qwen3-max-thinking
view ↗ |
8 | 7 | 8 | 8 | 5 | 8 | 8 |
6 |
1.02¢ |
6.0 |
Clean, accurate explainer with good analogies and a complete polished passage. Fatal weakness: only ONE quiz question provided despite requesting multiple MCQs. That single question is well-built, but the quiz dimension is severely underdelivered. |
| 14 |
openrouter:amazon/nova-pro-v1
view ↗ |
8 | 5 | 6 | 7 | 6 | 6 | 6 |
6 |
1.34¢ |
6.0 |
Accurate but thin and somewhat dumbed-down. Analogies (copying homework, chess, doping) are generic and don't capture the query-API distillation mechanism. Only 3 glossary terms. Quiz Q9/Q10 reference claims from the rewritten passage but feel circular. Serviceable but bland. |
| 15 |
anthropic:claude-haiku-4-5-20251001
view ↗ |
6 | 7 | 8 | 8 | 6 | 7 | 7 |
6 |
2.57¢ |
5.4 |
Vivid MIT-lab and Twitch-scraping analogies but introduces unsupported specifics: fabricates DeepSeek release timing details, asserts '$1 trillion global industry' figure not in source, and claims the memo named three firms (Anthropic did). Quiz Q4 asks about 'principal' but source says 'principally'—minor slip. Fact drift hurts it. |
| 16 |
openrouter:baidu/ernie-4.5-vl-424b-a47b
view ↗ |
7 | 6 | 7 | 7 | 5 | 6 | 7 |
5 |
0.43¢ |
6.5 |
Concise and accurate but thin: only 3 glossary terms and only 3 quiz questions (rubric expects a fuller set). Analogies (recipe heist, counterfeit goods, generic drug) are decent but not deeply mechanism-matched. Complete passage. Underdelivers on quiz breadth. |
| 17 |
view ↗ |
0 | 0 | 0 | 0 | 0 | 0 | 0 |
0 |
– |
– |
Placeholder—ignore. |
Scores 0–10 · 🟩 ≥8 · 🟨 4–7 · 🟥 <4 · sorted by Overall. Hover a column header for its full rubric name.
| # | Model | Quality | Cost | Value |
| 🥇 |
openrouter:deepseek/deepseek-v4-pro |
9 |
0.62¢ |
8.7 |
| 🥈 |
openrouter:minimax/minimax-m2.7 |
9 |
0.62¢ |
8.7 |
| 🥉 |
openrouter:mistralai/mistral-large-2512 |
8 |
0.85¢ |
8.0 |
| 4 |
openrouter:openai/gpt-5.4-nano |
8 |
0.63¢ |
8.0 |
| 5 |
openrouter:meta-llama/llama-4-maverick |
7 |
0.26¢ |
7.9 |
| 6 |
openrouter:google/gemini-3.1-flash-lite |
7 |
0.59¢ |
7.3 |
| 7 |
openrouter:z-ai/glm-5-turbo |
8 |
3.40¢ |
6.8 |
| 8 |
openrouter:x-ai/grok-4.3 |
7 |
1.64¢ |
6.7 |
| 9 |
anthropic:claude-sonnet-4-6 |
9 |
8.57¢ |
6.6 |
| 10 |
openrouter:moonshotai/kimi-k2.6 |
9 |
9.62¢ |
6.6 |
| 11 |
anthropic:claude-opus-4-8 |
9 |
15.59¢ |
6.6 |
| 12 |
openrouter:baidu/ernie-4.5-vl-424b-a47b |
5 |
0.43¢ |
6.5 |
| 13 |
openrouter:stepfun/step-3.7-flash |
7 |
2.09¢ |
6.1 |
| 14 |
openrouter:qwen/qwen3-max-thinking |
6 |
1.02¢ |
6.0 |
| 15 |
openrouter:amazon/nova-pro-v1 |
6 |
1.34¢ |
6.0 |
| 16 |
anthropic:claude-haiku-4-5-20251001 |
6 |
2.57¢ |
5.4 |
Strong quality without paying flagship prices — the cheap-and-good picks.