🔬 Model Lab

New run Stored runs ⚖️ Judge verdicts 🧮 Math 📊 Math runs 📄 Benchmark paper 📄 3-model paper 📄 Meta: Will Muse Cause a Spark?

⚖️ Judge verdict

The AI Cold War's Newest Weapon: Stealing the Teacher's Answer Key · Judge: anthropic:claude-opus-4-8 · 2026-09-05T00:17:39 · ✅ saved (7e6779) · 3 models

🔬 View outputs side-by-side → all verdicts →

🏆 Quality winner: openrouter:deepseek/deepseek-v4-pro
2eb734 combines the most precise, mechanism-matched analogies (Cold War Tu-4 reverse-engineering directly parallels distillation that copies but doesn't fully match), a complete and polished four-paragraph reading passage, a full 10-question set with correct evidence-pairing, and the richest accurate glossary. 0313fc is comparably strong but its rewritten passage is cut off mid-sentence, and a0132a offers only 3 quiz questions and a shorter passage.
💰 Best value: openrouter:deepseek/deepseek-v4-pro — quality 9/10 at 0.62¢ → value 8.7/10
Value = 0.7·quality + 0.3·cheapness. Quality stays dominant, so a cheap-but-weak output can't win on price alone.

Quality ranking

#ModelFactsAnalogyAge-fitClarityQuizGlossaryPassageOverallCostValueNotes
🥇 openrouter:deepseek/deepseek-v4-pro view ↗ 9999999 9 0.62¢ 8.7 Excellent mechanism-matched analogies (Tu-4/B-29 reverse-engineering, term-paper paraphrasing, open-source forking). Ten well-constructed SAT-style questions including proper evidence-pairing. Rich glossary and a polished, accurate, multi-paragraph passage. Facts carefully grounded with dates.
🥈 anthropic:claude-haiku-4-5-20251001 view ↗ 9899998 8 2.57¢ 6.8 Very strong, faithful, and well-organized with good analogies (arms race, doping scandal) and 10 fair questions with evidence-pairing. Passage appears truncated at the end ('respon'). Slightly behind the winner on analogy depth and passage completeness.
🥉 openrouter:baidu/ernie-4.5-vl-424b-a47b view ↗ 8788688 7 0.43¢ 7.9 Solid and faithful with good analogies (chef's recipe, generic drug). Only 3 quiz questions where competitors offer 10 with evidence-pairing; Q2 answer 'replicate' vs 'condense' is a bit arguable. Passage is accurate and engaging but compresses heavily.
Scores 0–10 · 🟩 ≥8 · 🟨 4–7 · 🟥 <4 · sorted by Overall. Hover a column header for its full rubric name.

💰 Cost-adjusted ranking (value for money)

#ModelQualityCostValue
🥇 openrouter:deepseek/deepseek-v4-pro 9 0.62¢ 8.7
🥈 openrouter:baidu/ernie-4.5-vl-424b-a47b 7 0.43¢ 7.9
🥉 anthropic:claude-haiku-4-5-20251001 8 2.57¢ 6.8

🪙 Best under 3¢ (cost < $0.03), ranked by quality

#ModelQualityCostValue
1 openrouter:deepseek/deepseek-v4-pro 9 0.62¢ 8.7
2 anthropic:claude-haiku-4-5-20251001 8 2.57¢ 6.8
3 openrouter:baidu/ernie-4.5-vl-424b-a47b 7 0.43¢ 7.9
Strong quality without paying flagship prices — the cheap-and-good picks.