🔬 Model Lab

New run Stored runs ⚖️ Judge verdicts 🧮 Math 📊 Math runs 📄 Benchmark paper 📄 3-model paper 📄 Meta: Will Muse Cause a Spark?

⚖️ Judge verdict

The AI Cold War's Newest Weapon: Stealing the Teacher's Answer Key · Judge: anthropic:claude-opus-4-8 · 2026-09-06T13:49:41 · ✅ saved (1f62b5) · 15 models

🔬 View outputs side-by-side → all verdicts →

Note: the judge model is also one of the contestants — scores may carry self-preference bias.
🏆 Quality winner: openrouter:deepseek/deepseek-v4-pro
DeepSeek-v4 wins on the strongest combination of faithful fact-grounding (no invented statistics, unlike Sonnet's fabricated 16M/24,000 figures), a genuinely mechanism-matching Cold War Tu-4/B-29 analogy that captures reverse-engineering-that-closes-a-gap-but-never-matches, a complete and engaging SAT-style passage, and a full, fair 10-question quiz whose items are answerable from the text with strong evidence-pairing. Claude Opus (c2b157) and Sonnet (fe7a5d) are very close, but Opus's Q4/Q5 lean on its own passage vocabulary and Sonnet invents specifics; DeepSeek avoids both pitfalls while matching their polish.
💰 Best value: openrouter:deepseek/deepseek-v4-pro — quality 9/10 at 0.62¢ → value 8.7/10
Value = 0.7·quality + 0.3·cheapness. Quality stays dominant, so a cheap-but-weak output can't win on price alone.

Quality ranking

#ModelFactsAnalogyAge-fitClarityQuizGlossaryPassageOverallCostValueNotes
🥇 anthropic:claude-sonnet-4-6 view ↗ 8999998 9 8.57¢ 6.6 Excellent analogies (chef ordering every dish, band recordings) that match the mechanism precisely. Introduces some invented specifics (16M exchanges, 24,000 accounts) not in the source, a minor grounding risk. Strong 10-question quiz with good SAT tips; passage truncated but high quality.
🥈 openrouter:deepseek/deepseek-v4-pro view ↗ 9999989 9 0.62¢ 8.7 Very faithful; strong Cold War Tu-4/B-29 analogy sharing the reverse-engineering mechanism. Complete, engaging rewritten passage. Full 10-question quiz with tone and evidence-pairing items well grounded.
🥉 openrouter:minimax/minimax-m2.7 view ↗ 9899999 9 0.62¢ 8.7 Highly faithful and clear; university/answer-key and F1-footage analogies fit the mechanism well. Complete, engaging passage; strong quiz with good evidence pairing. One of the best-balanced entries.
4 openrouter:mistralai/mistral-large-2512 view ↗ 8788888 8 0.85¢ 8.0 Solid, well-structured explainer with a complete passage. Marathon/save-scumming analogies are decent but the 'save scumming' one drifts from the actual mechanism. Facts largely accurate; quiz is fair and grounded.
5 openrouter:moonshotai/kimi-k2.6 view ↗ 8999889 8 9.62¢ 5.9 Sophisticated conservatory/pharma analogies matching mechanism; elegant complete passage. Quiz Q4/Q5 ask about words ('sharp', 'lightweight') that appear only in its own rewritten passage rather than the source, a mild fairness issue but internally consistent.
6 openrouter:google/gemini-3.1-flash-lite view ↗ 8788888 8 0.59¢ 8.0 Chef-tasting analogy fits; clean complete passage and a full, fair 10-item quiz with good tone/evidence questions. Solid all-around, slightly less distinctive than the top entries.
7 openrouter:z-ai/glm-5-turbo view ↗ 7899888 8 3.40¢ 6.8 Vivid music-producer analogy and lively 'Why It Matters'. Q4 asks about 'sharp' which appears only in its own passage; adds McGuire details fabricated slightly ('billions', 'from scratch'). Passage truncated but strong.
8 openrouter:openai/gpt-5.4-nano view ↗ 8788888 8 0.63¢ 8.0 Careful, faithful GPT-5-style entry that explicitly attributes claims to 'the article', reducing overclaiming. Complete passage and a well-constructed, grounded 10-item quiz. Analogies (locked lab, film scouting) are good if slightly generic.
9 anthropic:claude-haiku-4-5-20251001 view ↗ 8788888 8 2.57¢ 6.8 Strong nuclear-arms-race and doping analogies (doping is apt for the legal/illegal distinction). Complete passage, full grounded quiz with excellent evidence-pairing. Truncated final paragraph is the only blemish.
10 openrouter:stepfun/step-3.7-flash view ↗ 7888887 7 2.09¢ 6.1 Good SAT/burger analogies and thorough quiz with trap-labeling. Adds unsupported specifics ('state-linked groups', memo 'names three firms', bills 'passed April 2026') that overstate the source. Passage truncated.
11 openrouter:x-ai/grok-4.3 view ↗ 7688888 7 1.64¢ 6.7 Clear and neutral with reasonable scouting/reverse-engineering analogies. Complete passage and full quiz. Some quiz items reference details ('$1 trillion' absent) or added framing; grounding adequate but not exceptional.
12 openrouter:qwen/qwen3-max-thinking view ↗ 8887588 6 1.02¢ 6.0 Concise and accurate with good handbag/answer-key analogies and a clean passage. Major weakness: only ONE quiz question provided, well below the expected set, sharply limiting quiz value.
13 openrouter:amazon/nova-pro-v1 view ↗ 7566666 6 1.34¢ 6.0 Generic chess/doping analogies that don't capture the distillation mechanism well. Q9/Q10 rely on a line ('significant implications for global power dynamics') that exists only in its own added text, weakening grounding. Serviceable but thin.
14 openrouter:bytedance-seed/seed-2.0-lite view ↗ 6788687 6 1.73¢ 6.0 Q4 asks meaning of 'principal' claiming the passage describes 'Chinese entities' as 'principal'—but neither its passage nor source uses that adjective that way, an unfair/fabricated item. Otherwise decent MIT-lab analogy and complete passage.
15 openrouter:baidu/ernie-4.5-vl-424b-a47b view ↗ 7677567 6 0.43¢ 7.2 Compact and readable with a chef-recipe hook, but only THREE quiz questions and three glossary terms, well short of expectations. Analogies are somewhat generic; passage complete but brief.
Scores 0–10 · 🟩 ≥8 · 🟨 4–7 · 🟥 <4 · sorted by Overall. Hover a column header for its full rubric name.

💰 Cost-adjusted ranking (value for money)

#ModelQualityCostValue
🥇 openrouter:deepseek/deepseek-v4-pro 9 0.62¢ 8.7
🥈 openrouter:minimax/minimax-m2.7 9 0.62¢ 8.7
🥉 openrouter:mistralai/mistral-large-2512 8 0.85¢ 8.0
4 openrouter:google/gemini-3.1-flash-lite 8 0.59¢ 8.0
5 openrouter:openai/gpt-5.4-nano 8 0.63¢ 8.0
6 openrouter:baidu/ernie-4.5-vl-424b-a47b 6 0.43¢ 7.2
7 openrouter:z-ai/glm-5-turbo 8 3.40¢ 6.8
8 anthropic:claude-haiku-4-5-20251001 8 2.57¢ 6.8
9 openrouter:x-ai/grok-4.3 7 1.64¢ 6.7
10 anthropic:claude-sonnet-4-6 9 8.57¢ 6.6
11 openrouter:stepfun/step-3.7-flash 7 2.09¢ 6.1
12 openrouter:qwen/qwen3-max-thinking 6 1.02¢ 6.0
13 openrouter:amazon/nova-pro-v1 6 1.34¢ 6.0
14 openrouter:bytedance-seed/seed-2.0-lite 6 1.73¢ 6.0
15 openrouter:moonshotai/kimi-k2.6 8 9.62¢ 5.9

🪙 Best under 3¢ (cost < $0.03), ranked by quality

#ModelQualityCostValue
1 openrouter:deepseek/deepseek-v4-pro 9 0.62¢ 8.7
2 openrouter:minimax/minimax-m2.7 9 0.62¢ 8.7
3 openrouter:mistralai/mistral-large-2512 8 0.85¢ 8.0
4 openrouter:google/gemini-3.1-flash-lite 8 0.59¢ 8.0
5 openrouter:openai/gpt-5.4-nano 8 0.63¢ 8.0
6 anthropic:claude-haiku-4-5-20251001 8 2.57¢ 6.8
7 openrouter:x-ai/grok-4.3 7 1.64¢ 6.7
8 openrouter:stepfun/step-3.7-flash 7 2.09¢ 6.1
9 openrouter:baidu/ernie-4.5-vl-424b-a47b 6 0.43¢ 7.2
10 openrouter:qwen/qwen3-max-thinking 6 1.02¢ 6.0
11 openrouter:amazon/nova-pro-v1 6 1.34¢ 6.0
12 openrouter:bytedance-seed/seed-2.0-lite 6 1.73¢ 6.0
Strong quality without paying flagship prices — the cheap-and-good picks.