Note: the judge model is also one of the contestants
— scores may carry self-preference bias.
🏆 Quality winner: openrouter:deepseek/deepseek-v4-pro
DeepSeek-v4-pro (2eb734) edges out the strong Opus and Sonnet entries by combining a COMPLETE, non-truncated, genuinely SAT-style reading passage with high factual fidelity, a full 10-question quiz (including vocab-in-context and evidence-pairing pairs), and a standout Tu-4/B-29 Cold War analogy that shares the exact reverse-engineering mechanism of distillation. Opus (c2b157) is nearly equal and also excellent, but Sonnet (fe7a5d) introduces unsourced hard numbers (16M exchanges, 24k accounts) that risk fabrication, and several rivals truncate their passages. 2eb734 has no material factual errors, complete deliverables, and consistently strong dimension scores across the board.
| # | Model | Facts | Analogy | Age-fit | Clarity | Quiz | Glossary | Passage | Overall | Cost | Value | Notes |
| 🥇 |
anthropic:claude-sonnet-4-6
view ↗ |
8 | 9 | 9 | 9 | 9 | 9 | 8 |
9 |
8.57¢ |
6.6 |
Excellent, mechanism-matched analogies (chef ordering every dish, recording a band). Adds unsourced specific figures (16 million exchanges, 24,000 accounts) not in source—minor fabrication risk. Ten well-crafted SAT-style questions with vocab-in-context and evidence pairing. Reading passage truncated but strong. |
| 🥈 |
openrouter:deepseek/deepseek-v4-pro
view ↗ |
9 | 9 | 9 | 9 | 9 | 8 | 9 |
9 |
0.62¢ |
8.7 |
Strong fidelity, complete non-truncated reading passage that is accurate and engaging. Tu-4/B-29 Cold War analogy is precise and shares the same reverse-engineering mechanism. Full 10-question quiz with good vocab and inference items. One of the best all-around. |
| 🥉 |
openrouter:minimax/minimax-m2.7
view ↗ |
9 | 8 | 9 | 9 | 9 | 9 | 9 |
9 |
0.62¢ |
8.7 |
Very faithful, university/answer-key analogy matches the distillation mechanism well. Complete, accurate, engaging reading passage. Ten fair, non-trivial questions with strong distractor analysis. Excellent glossary including safety guardrails. Top-tier. |
| 4 |
openrouter:moonshotai/kimi-k2.6
view ↗ |
8 | 8 | 9 | 9 | 8 | 8 | 8 |
8 |
9.62¢ |
5.9 |
Sophisticated, well-written with conservatory and pharma analogies that fit the harvest-at-scale mechanism. Complete polished passage. Slight vocab-question mismatch ('lightweight' quiz item references a word not clearly in the rewritten passage). Strong overall. |
| 5 |
anthropic:claude-haiku-4-5-20251001
view ↗ |
8 | 7 | 9 | 9 | 9 | 9 | 8 |
8 |
2.57¢ |
6.8 |
Haiku delivers a very readable, well-structured explainer with a full six-term glossary and a strong 10-question quiz with careful trap analysis. Arms-race and doping analogies are decent. Passage truncated at end but accurate. High quality. |
| 6 |
openrouter:mistralai/mistral-large-2512
view ↗ |
8 | 6 | 8 | 8 | 8 | 8 | 8 |
7 |
0.85¢ |
7.3 |
Solid and faithful with complete passage. Analogies (study guide, F1 car, save-scumming) are somewhat generic and don't perfectly capture the query-harvest mechanism. Some added framing (US firms pulling back from collaborations) is speculative. Good quiz coverage. |
| 7 |
openrouter:stepfun/step-3.7-flash
view ↗ |
6 | 8 | 8 | 8 | 8 | 8 | 7 |
7 |
2.09¢ |
6.1 |
Good analogies (valedictorian, burger reverse-engineering). But adds unsupported specifics: says the memo 'names three Chinese firms' (Anthropic named them, not the memo) and asserts they are 'state-linked'—both fabrications. Passage truncated. Quiz strong. |
| 8 |
openrouter:google/gemini-3.1-flash-lite
view ↗ |
8 | 7 | 8 | 8 | 8 | 7 | 8 |
7 |
0.59¢ |
7.3 |
Solid Gemini output with recipe/capture-the-flag analogies (recipe one fits mechanism). Complete accurate passage. Full 10-question quiz. Glossary only 4 terms. Reliable mid-upper tier. |
| 9 |
openrouter:x-ai/grok-4.3
view ↗ |
8 | 6 | 8 | 8 | 8 | 8 | 8 |
7 |
1.64¢ |
6.7 |
Faithful Grok output, complete clean passage, full 10-question quiz with reasonable trap analysis. Analogies (reverse-engineering, scouting playbook) are competent but somewhat generic. Solid, no major errors. |
| 10 |
openrouter:meta-llama/llama-4-maverick
view ↗ |
8 | 6 | 8 | 8 | 8 | 8 | 8 |
7 |
0.26¢ |
7.9 |
Careful GPT-style output that repeatedly signals 'the article says,' keeping it well-grounded. Complete passage. Full quiz. Analogies (lock on lab, film scouting) are okay but not deeply mechanistic. Reliable but slightly hedged/dry. |
| 11 |
openrouter:openai/gpt-5.4-nano
view ↗ |
8 | 6 | 8 | 8 | 8 | 8 | 8 |
7 |
0.63¢ |
7.3 |
GPT-5-nano output is disciplined and well-grounded (hedges appropriately with 'the article says'). Analogies (locked lab equipment, film scouting) are serviceable. Full 10-question quiz, complete passage. Solid mid-upper tier. |
| 12 |
openrouter:qwen/qwen3-max-thinking
view ↗ |
8 | 7 | 8 | 7 | 4 | 8 | 8 |
6 |
1.02¢ |
6.0 |
Clean, faithful explanation and a strong complete passage. Fatal weakness: only ONE quiz question provided despite the rubric expecting a full set. Analogies decent. Underdelivers on quiz volume. |
| 13 |
openrouter:z-ai/glm-5-turbo
view ↗ |
5 | 8 | 8 | 8 | 5 | 8 | 8 |
6 |
3.40¢ |
5.4 |
Rich analogies and good writing, but quiz Q4 asks about 'sharp' and Q9 references material that appears to invent quotes ('central friction point,' McGuire recommendations phrased as passage text) not in the source—fabricated evidence options. Passage truncated. Fact-grounding suffers. |
| 14 |
openrouter:bytedance-seed/seed-2.0-lite
view ↗ |
7 | 7 | 8 | 8 | 7 | 8 | 7 |
6 |
1.73¢ |
6.0 |
Vivid MIT-lab and Twitch analogies. Adds 'April 2026' framing and an unsupported '$1 trillion industry' stat. Q4 asks about 'principal' but references a word ('principal') that isn't in its own rewritten passage as used—slight mismatch. Complete passage. Decent. |
| 15 |
openrouter:amazon/nova-pro-v1
view ↗ |
7 | 5 | 6 | 6 | 6 | 6 | 6 |
5 |
1.34¢ |
5.3 |
Generic chess/homework/doping analogies that don't capture the distillation mechanism. Thin glossary (3 terms). Q10 evidence-pairing cites a fabricated quote ('arms race' line is from a caption, and the inferred claim is weakly grounded). Serviceable but flat. |
| 16 |
openrouter:baidu/ernie-4.5-vl-424b-a47b
view ↗ |
6 | 6 | 7 | 6 | 5 | 6 | 6 |
5 |
0.43¢ |
6.5 |
Brief and functional but underdeveloped: only 3 quiz questions, 3 glossary terms, and no bullets in some sections. Q2 claims 'distil' means 'replicate' via a quote that misquotes the source. Generic counterfeit/plagiarism analogies. Weakest complete entry. |
Scores 0–10 · 🟩 ≥8 · 🟨 4–7 · 🟥 <4 · sorted by Overall. Hover a column header for its full rubric name.
| # | Model | Quality | Cost | Value |
| 🥇 |
openrouter:deepseek/deepseek-v4-pro |
9 |
0.62¢ |
8.7 |
| 🥈 |
openrouter:minimax/minimax-m2.7 |
9 |
0.62¢ |
8.7 |
| 🥉 |
openrouter:meta-llama/llama-4-maverick |
7 |
0.26¢ |
7.9 |
| 4 |
openrouter:mistralai/mistral-large-2512 |
7 |
0.85¢ |
7.3 |
| 5 |
openrouter:google/gemini-3.1-flash-lite |
7 |
0.59¢ |
7.3 |
| 6 |
openrouter:openai/gpt-5.4-nano |
7 |
0.63¢ |
7.3 |
| 7 |
anthropic:claude-haiku-4-5-20251001 |
8 |
2.57¢ |
6.8 |
| 8 |
openrouter:x-ai/grok-4.3 |
7 |
1.64¢ |
6.7 |
| 9 |
anthropic:claude-sonnet-4-6 |
9 |
8.57¢ |
6.6 |
| 10 |
openrouter:baidu/ernie-4.5-vl-424b-a47b |
5 |
0.43¢ |
6.5 |
| 11 |
openrouter:stepfun/step-3.7-flash |
7 |
2.09¢ |
6.1 |
| 12 |
openrouter:qwen/qwen3-max-thinking |
6 |
1.02¢ |
6.0 |
| 13 |
openrouter:bytedance-seed/seed-2.0-lite |
6 |
1.73¢ |
6.0 |
| 14 |
openrouter:moonshotai/kimi-k2.6 |
8 |
9.62¢ |
5.9 |
| 15 |
openrouter:z-ai/glm-5-turbo |
6 |
3.40¢ |
5.4 |
| 16 |
openrouter:amazon/nova-pro-v1 |
5 |
1.34¢ |
5.3 |
| # | Model | Quality | Cost | Value |
| 1 |
openrouter:deepseek/deepseek-v4-pro |
9 |
0.62¢ |
8.7 |
| 2 |
openrouter:minimax/minimax-m2.7 |
9 |
0.62¢ |
8.7 |
| 3 |
anthropic:claude-haiku-4-5-20251001 |
8 |
2.57¢ |
6.8 |
| 4 |
openrouter:meta-llama/llama-4-maverick |
7 |
0.26¢ |
7.9 |
| 5 |
openrouter:mistralai/mistral-large-2512 |
7 |
0.85¢ |
7.3 |
| 6 |
openrouter:google/gemini-3.1-flash-lite |
7 |
0.59¢ |
7.3 |
| 7 |
openrouter:openai/gpt-5.4-nano |
7 |
0.63¢ |
7.3 |
| 8 |
openrouter:x-ai/grok-4.3 |
7 |
1.64¢ |
6.7 |
| 9 |
openrouter:stepfun/step-3.7-flash |
7 |
2.09¢ |
6.1 |
| 10 |
openrouter:qwen/qwen3-max-thinking |
6 |
1.02¢ |
6.0 |
| 11 |
openrouter:bytedance-seed/seed-2.0-lite |
6 |
1.73¢ |
6.0 |
| 12 |
openrouter:baidu/ernie-4.5-vl-424b-a47b |
5 |
0.43¢ |
6.5 |
| 13 |
openrouter:amazon/nova-pro-v1 |
5 |
1.34¢ |
5.3 |
Strong quality without paying flagship prices — the cheap-and-good picks.