🏆 Quality winner: openrouter:deepseek/deepseek-v4-pro
DeepSeek-V4-Pro retains the quality crown: its Soviet Tu-4/B-29 bomber analogy (a reverse-engineered replica that closed a gap yet never matched the original) is still the single best mechanism-match in the field, paired with top-tier execution everywhere. The two newcomers land in the top tier on quality — Claude Opus 4.8 (impeccably clean, faithful, complete; tied at 9) and, just behind, Claude Sonnet 4.6, whose craft is arguably the best of all (the 'reconstruct the sheet music from ten thousand recordings' analogy; an 8-term glossary) but which is capped at 8 for stating as fact unverifiable specifics — '~16 million exchanges via ~24,000 accounts' — that are absent from the source. GLM-5-Turbo and Claude Haiku 4.5 complete the 9-tier. Crucially, on a COST-ADJUSTED basis the picture flips: Opus 4.8 is the most expensive model tested ($0.156) and Sonnet 4.6 the second ($0.086), so despite their quality they rank poorly on value — while DeepSeek delivers the same 9/10 quality at $0.0062 (about 25x cheaper than Opus), making it the value winner too. See the cost-adjusted and under-3-cent tables below. Judge transparency: the judge runs on Opus 4.8, the same family as one contestant; Opus was held to the same bar and did not win.
| # | Model | Facts | Analogy | Age-fit | Clarity | Quiz | Glossary | Passage | Overall | Cost | Value | Notes |
| 🥇 |
openrouter:deepseek/deepseek-v4-pro
view ↗ |
9 | 10 | 9 | 9 | 9 | 9 | 9 |
9 |
0.62¢ |
8.7 |
Best all-round. The B-29/Tu-4 bomber analogy is the standout of the field — a reverse-engineered replica that closed a gap yet never matched the original, mapping distillation precisely. Rich accurate explanation, full 10-question quiz with trap analysis + varied SAT tips, strong 6-term glossary, engaging faithful passage. |
| 🥈 |
openrouter:z-ai/glm-5-turbo
view ↗ |
10 | 8 | 9 | 9 | 10 | 8 | 10 |
9 |
3.40¢ |
7.5 |
Most thorough and faithful output: a six-paragraph passage that captures nuance and even Chris McGuire's specific recommendations. The quiz is the most rigorous, with a consistent Trap A/B/C taxonomy and an SAT tip on every item. Analogies (vocal-stripping, scouting reports) are solid if a touch less vivid than the very best. |
| 🥉 |
anthropic:claude-opus-4-8
view ↗ |
9 | 9 | 9 | 9 | 9 | 9 | 9 |
9 |
15.59¢ |
6.6 |
Excellent and impeccably clean: faithful throughout with no invented specifics, a complete 10-question quiz with Trap labels + SAT tips, a 7-term glossary, and good same-mechanism analogies (reverse-engineer a recipe by tasting; a chess student memorising a grandmaster's recorded games). Tight, well-calibrated five-paragraph passage. Just shy of the top only because its analogies are a touch more familiar than DeepSeek's bomber. NB: Opus 4.8 is the judge's own family — held to the same bar, not favoured. |
| 4 |
anthropic:claude-haiku-4-5-20251001
view ↗ |
9 | 9 | 9 | 9 | 9 | 9 | 9 |
9 |
2.57¢ |
7.5 |
Polished and accurate. The doping analogy — 'distillation isn't cheating, but using someone else's distilled model is' — is sharp and on-mechanism. Strong 6-term glossary, rich quiz with trap analysis, well-structured four-paragraph passage that includes McGuire. A balanced top-tier entry. |
| 5 |
anthropic:claude-sonnet-4-6
view ↗ |
7 | 10 | 9 | 9 | 9 | 10 | 9 |
8 |
8.57¢ |
5.9 |
Superb craft — the 'reconstruct the sheet music from ten thousand recordings' analogy is the best single mechanism-match in the field, plus an 8-term glossary (the most thorough here) and an excellent quiz. Capped at 8 on a fact-grounding violation: it asserts as fact that Anthropic reported '~16 million exchanges via ~24,000 accounts' — specifics absent from the source and unverifiable here, repeated in both the explanation and the passage. Otherwise it would top the field. |
| 6 |
openrouter:moonshotai/kimi-k2.6
view ↗ |
9 | 9 | 9 | 8 | 9 | 8 | 8 |
8 |
9.62¢ |
5.9 |
Sophisticated, on-mechanism analogies (music-conservatory transcription, pharma-abstract scraping). Full quiz with trap analysis and a polished, faithful passage. Minor blemish: 'Chris McGuire' is named twice in consecutive sentences — a small editing lapse. Just below the very top. |
| 7 |
openrouter:stepfun/step-3.7-flash
view ↗ |
9 | 9 | 8 | 8 | 9 | 8 | 8 |
8 |
2.09¢ |
6.8 |
Concrete, same-mechanism analogies (valedictorian's answers; reverse-engineering a competitor's burger) and a full 10-question quiz with explicit trap labels. Long, detailed, faithful passage. Reads slightly long but is high quality throughout. |
| 8 |
openrouter:mistralai/mistral-large-2512
view ↗ |
9 | 8 | 9 | 8 | 9 | 8 | 8 |
8 |
0.85¢ |
8.0 |
Strong 'answer key' hook and a full quiz with trap analysis. Good supply-chain-heist and marathon-sweatband analogies, though 'save-scumming' is a looser fit. Faithful, well-paragraphed passage. Solid top-half entry. |
| 9 |
openrouter:minimax/minimax-m2.7
view ↗ |
9 | 8 | 9 | 8 | 8 | 8 | 8 |
8 |
0.62¢ |
8.0 |
The 'F1 car cloned from race footage, sold at Toyota prices' analogy is vivid and on-mechanism. Full quiz, good 6-term glossary, accurate five-paragraph passage. A well-rounded, faithful entry. |
| 10 |
openrouter:x-ai/grok-4.3
view ↗ |
8 | 7 | 8 | 8 | 7 | 8 | 8 |
7 |
1.64¢ |
6.7 |
Clean and accurate, with a tidy five-paragraph passage. Analogies (reverse-engineer a product, scout a playbook with hidden cameras) are competent but familiar. Quiz is complete but its explanations are thinner — brief trap notes, few SAT tips — than the leaders. |
| 11 |
openrouter:google/gemini-3.1-flash-lite
view ↗ |
8 | 7 | 8 | 8 | 8 | 7 | 7 |
7 |
0.59¢ |
7.3 |
Solid and accurate but lighter: shorter explanation sections and a four-paragraph passage with less depth. Quiz is complete and fair with SAT tips. Analogies (master-chef recipe, capture-the-flag lock-picking) are serviceable. A dependable mid-tier result. |
| 12 |
openrouter:bytedance-seed/seed-2.0-lite
view ↗ |
8 | 8 | 8 | 6 | 8 | 8 | 7 |
7 |
1.73¢ |
6.7 |
Vivid analogies (MIT exam-key theft, Twitch strategy-scraping) and a complete quiz with trap labels. Main weakness: the rewritten passage is a single unbroken wall of text with no paragraph breaks, which hurts scannability for the target reader. |
| 13 |
openrouter:openai/gpt-5.4-nano
view ↗ |
8 | 7 | 8 | 7 | 8 | 8 | 6 |
7 |
0.63¢ |
7.3 |
Competent quiz with trap analysis and a clear glossary. The defining flaw is voice: the explanation and the 'rewritten' passage repeatedly say 'the article explains/describes/frames...', narrating the source instead of producing original SAT-style prose — which undercuts the passage's whole purpose and immersion. |
| 14 |
openrouter:meta-llama/llama-4-maverick
view ↗ |
8 | 6 | 7 | 6 | 7 | 7 | 6 |
6 |
0.26¢ |
7.2 |
Accurate but flat. The hook merely restates the headline rather than sparking curiosity, the analogy is the generic 'copying homework,' and the rewritten passage is one undifferentiated block. Quiz is complete but basic. Functional, not engaging. |
| 15 |
openrouter:qwen/qwen3-max-thinking
view ↗ |
8 | 8 | 8 | 6 | 1 | 8 | 8 |
6 |
1.02¢ |
6.0 |
Strong prose, a good counterfeit-handbag-from-photos analogy, and a faithful passage — but it produced only ONE quiz question instead of ten (the model likely exhausted its token budget on reasoning). The missing quiz is a critical deliverable failure that drags an otherwise good entry down. |
| 16 |
openrouter:amazon/nova-pro-v1
view ↗ |
7 | 5 | 6 | 6 | 7 | 5 | 5 |
5 |
1.34¢ |
5.3 |
Thin throughout: a one-line hook, brief sections, only three glossary terms, and a short single-paragraph passage. Analogies (chess, copying homework, performance-enhancing drugs) are generic and not mechanism-matched. Quiz is complete but basic. Among the weakest. |
| 17 |
openrouter:baidu/ernie-4.5-vl-424b-a47b
view ↗ |
7 | 6 | 6 | 6 | 3 | 5 | 6 |
5 |
0.43¢ |
6.5 |
Incomplete: only three quiz questions (of ten) and three glossary terms, with a short single-block passage. Analogies are listed but generic (secret recipe, counterfeit goods, generic drug). Accurate as far as it goes, but well short of the required deliverables. |
Scores 0–10 · 🟩 ≥8 · 🟨 4–7 · 🟥 <4 · sorted by Overall. Hover a column header for its full rubric name.
| # | Model | Quality | Cost | Value |
| 🥇 |
openrouter:deepseek/deepseek-v4-pro |
9 |
0.62¢ |
8.7 |
| 🥈 |
openrouter:mistralai/mistral-large-2512 |
8 |
0.85¢ |
8.0 |
| 🥉 |
openrouter:minimax/minimax-m2.7 |
8 |
0.62¢ |
8.0 |
| 4 |
openrouter:z-ai/glm-5-turbo |
9 |
3.40¢ |
7.5 |
| 5 |
anthropic:claude-haiku-4-5-20251001 |
9 |
2.57¢ |
7.5 |
| 6 |
openrouter:google/gemini-3.1-flash-lite |
7 |
0.59¢ |
7.3 |
| 7 |
openrouter:openai/gpt-5.4-nano |
7 |
0.63¢ |
7.3 |
| 8 |
openrouter:meta-llama/llama-4-maverick |
6 |
0.26¢ |
7.2 |
| 9 |
openrouter:stepfun/step-3.7-flash |
8 |
2.09¢ |
6.8 |
| 10 |
openrouter:x-ai/grok-4.3 |
7 |
1.64¢ |
6.7 |
| 11 |
openrouter:bytedance-seed/seed-2.0-lite |
7 |
1.73¢ |
6.7 |
| 12 |
anthropic:claude-opus-4-8 |
9 |
15.59¢ |
6.6 |
| 13 |
openrouter:baidu/ernie-4.5-vl-424b-a47b |
5 |
0.43¢ |
6.5 |
| 14 |
openrouter:qwen/qwen3-max-thinking |
6 |
1.02¢ |
6.0 |
| 15 |
anthropic:claude-sonnet-4-6 |
8 |
8.57¢ |
5.9 |
| 16 |
openrouter:moonshotai/kimi-k2.6 |
8 |
9.62¢ |
5.9 |
| 17 |
openrouter:amazon/nova-pro-v1 |
5 |
1.34¢ |
5.3 |
Strong quality without paying flagship prices — the cheap-and-good picks.