🔬 Model Lab

New run Stored runs ⚖️ Judge verdicts 🧮 Math 📊 Math runs 📄 Benchmark paper 📄 3-model paper 📄 Meta: Will Muse Cause a Spark?

⚖️ Judge verdict

The AI Cold War's Newest Weapon: Stealing the Teacher's Answer Key · Judge: anthropic:claude-opus-4-8 · 2026-09-25T04:35:04 · ✅ saved (7ec8b6) · 3 models

🔬 View outputs side-by-side → all verdicts →

🏆 Quality winner: openrouter:deepseek/deepseek-v4-pro
2eb734 edges out 0313fc on the strength of its complete, polished reading passage and its distinctive mechanism-matching analogies (Tu-4 bomber, forking code), plus its clean, non-truncated output. Both are far richer than a0132a, which has only three questions and less depth. 2eb734's quiz is fully self-contained and its vocab questions ('frontier', 'distillation') are grounded in its own passage, whereas 0313fc leans on 'surreptitious', a word it dropped from the rewrite.
💰 Best value: openrouter:deepseek/deepseek-v4-pro — quality 9/10 at 0.62¢ → value 8.7/10
Value = 0.7·quality + 0.3·cheapness. Quality stays dominant, so a cheap-but-weak output can't win on price alone.

Quality ranking

#ModelFactsAnalogyAge-fitClarityQuizGlossaryPassageOverallCostValueNotes
🥇 openrouter:deepseek/deepseek-v4-pro view ↗ 99999910 9 0.62¢ 8.7 Excellent, richly detailed. The Tu-4 bomber and term-paper analogies genuinely share the mechanism of reverse-engineering/extracting knowledge. Ten well-constructed SAT-style questions including proper evidence-pairing, precise glossary, and an outstanding original reading passage. Minor: dates slightly reworked (Feb 2026 vs source's ambiguous 'February') but reasonable.
🥈 anthropic:claude-haiku-4-5-20251001 view ↗ 9999899 8 2.57¢ 6.8 Very strong: doping-scandal and arms-race analogies fit the mechanism well, six glossary terms are precise, and ten quiz questions include good evidence-pairing. One 'surreptitious' vocab question relies on a word in the source not carried into the rewritten passage. Reading passage appears truncated at the end but is otherwise excellent.
🥉 openrouter:baidu/ernie-4.5-vl-424b-a47b view ↗ 8788688 7 0.43¢ 7.9 Solid, faithful summary with good chef/recipe hook and clean glossary. Only 3 quiz questions, and the 'distil' vocab item is debatable (answer C 'replicate' vs D 'steal' is not clearly grounded). Fewer questions and shorter than competitors, but accurate throughout.
Scores 0–10 · 🟩 ≥8 · 🟨 4–7 · 🟥 <4 · sorted by Overall. Hover a column header for its full rubric name.

💰 Cost-adjusted ranking (value for money)

#ModelQualityCostValue
🥇 openrouter:deepseek/deepseek-v4-pro 9 0.62¢ 8.7
🥈 openrouter:baidu/ernie-4.5-vl-424b-a47b 7 0.43¢ 7.9
🥉 anthropic:claude-haiku-4-5-20251001 8 2.57¢ 6.8

🪙 Best under 3¢ (cost < $0.03), ranked by quality

#ModelQualityCostValue
1 openrouter:deepseek/deepseek-v4-pro 9 0.62¢ 8.7
2 anthropic:claude-haiku-4-5-20251001 8 2.57¢ 6.8
3 openrouter:baidu/ernie-4.5-vl-424b-a47b 7 0.43¢ 7.9
Strong quality without paying flagship prices — the cheap-and-good picks.