🔬 Model Lab

New run Stored runs ⚖️ Judge verdicts 🧮 Math 📊 Math runs 📄 Benchmark paper 📄 3-model paper 📄 Meta: Will Muse Cause a Spark?

⚖️ Judge verdict

The AI Cold War's Newest Weapon: Stealing the Teacher's Answer Key · Judge: anthropic:claude-opus-4-8 · 2026-09-19T07:35:08 · ✅ saved (c2d9e3) · 3 models

🔬 View outputs side-by-side → all verdicts →

🏆 Quality winner: openrouter:deepseek/deepseek-v4-pro
Both DeepSeek and Haiku are far richer than Ernie, offering full quizzes, deep glossaries and multi-paragraph passages. DeepSeek edges out Haiku with the most mechanistically precise analogy (Tu-4 reverse-engineering matches distillation's 'copy that doesn't match the original but is cheaper'), consistently self-contained passage prose, and a complete untruncated output, whereas Haiku's passage is cut off mid-sentence and a couple of its vocab questions test words not in its own rewritten passage.
💰 Best value: openrouter:deepseek/deepseek-v4-pro — quality 9/10 at 0.62¢ → value 8.7/10
Value = 0.7·quality + 0.3·cheapness. Quality stays dominant, so a cheap-but-weak output can't win on price alone.

Quality ranking

#ModelFactsAnalogyAge-fitClarityQuizGlossaryPassageOverallCostValueNotes
🥇 openrouter:deepseek/deepseek-v4-pro view ↗ 9999999 9 0.62¢ 8.7 Excellent depth and accuracy; the Tu-4/B-29 Cold War analogy shares the reverse-engineering mechanism precisely. Ten well-constructed SAT-style questions including evidence-pairing and vocab-in-context, all answerable from the passage. Rich, original multi-paragraph passage and precise glossary. Minor: some questions reference the original FT text rather than only the rewritten passage.
🥈 anthropic:claude-haiku-4-5-20251001 view ↗ 9999899 8 2.57¢ 6.8 Very strong: engaging, accurate, with sharp analogies (doping scandal, arms race) that fit the 'legal-vs-illicit' distinction well. Ten good questions, though Q4 ('surreptitious') and some evidence questions test the source text rather than the rewritten passage, slightly unfair. Passage is excellent and comprehensive; the truncated final word is a small flaw.
🥉 openrouter:baidu/ernie-4.5-vl-424b-a47b view ↗ 8788688 7 0.43¢ 7.9 Solid, faithful summary with clean structure and good analogies (recipe, plagiarism). But only 3 quiz questions, and Q2 answer 'replicate' over 'steal' is defensible but debatable; some quiz explanations reference 'traps' oddly. Passage is accurate but compresses heavily.
Scores 0–10 · 🟩 ≥8 · 🟨 4–7 · 🟥 <4 · sorted by Overall. Hover a column header for its full rubric name.

💰 Cost-adjusted ranking (value for money)

#ModelQualityCostValue
🥇 openrouter:deepseek/deepseek-v4-pro 9 0.62¢ 8.7
🥈 openrouter:baidu/ernie-4.5-vl-424b-a47b 7 0.43¢ 7.9
🥉 anthropic:claude-haiku-4-5-20251001 8 2.57¢ 6.8

🪙 Best under 3¢ (cost < $0.03), ranked by quality

#ModelQualityCostValue
1 openrouter:deepseek/deepseek-v4-pro 9 0.62¢ 8.7
2 anthropic:claude-haiku-4-5-20251001 8 2.57¢ 6.8
3 openrouter:baidu/ernie-4.5-vl-424b-a47b 7 0.43¢ 7.9
Strong quality without paying flagship prices — the cheap-and-good picks.