🔬 Model Lab

New run Stored runs ⚖️ Judge verdicts 🧮 Math 📊 Math runs 📄 Benchmark paper 📄 3-model paper 📄 Meta: Will Muse Cause a Spark?

⚖️ Judge verdict

The AI Cold War's Newest Weapon: Stealing the Teacher's Answer Key · Judge: anthropic:claude-opus-4-8 · 2026-09-04T20:55:58 · ✅ saved (e4bd4c) · 3 models

🔬 View outputs side-by-side → all verdicts →

🏆 Quality winner: openrouter:deepseek/deepseek-v4-pro
Run 2eb734 wins on analogy quality (the Tu-4/B-29 Cold War reverse-engineering analogy precisely mirrors distillation's mechanism—copying that's cheaper but doesn't fully match the original), the richest and most self-contained rewritten passage, and a complete 10-question quiz whose questions are all answerable from its own passage. Run 0313fc is very close but its 'surreptitious' vocab question references a word not in its rewritten passage. Run a0132a is solid but weaker with only 3 quiz questions.
💰 Best value: openrouter:deepseek/deepseek-v4-pro — quality 9/10 at 0.62¢ → value 8.7/10
Value = 0.7·quality + 0.3·cheapness. Quality stays dominant, so a cheap-but-weak output can't win on price alone.

Quality ranking

#ModelFactsAnalogyAge-fitClarityQuizGlossaryPassageOverallCostValueNotes
🥇 openrouter:deepseek/deepseek-v4-pro view ↗ 9999999 9 0.62¢ 8.7 Excellent mechanistic analogies (Tu-4/B-29 reverse-engineering, filming opponent's practice, forking code) that share the actual distillation mechanism. Ten well-constructed SAT-style questions including evidence-pairing, all answerable from the rich passage. Glossary is precise and comprehensive; dates and names all accurate.
🥈 anthropic:claude-haiku-4-5-20251001 view ↗ 9899898 8 2.57¢ 6.8 Strong faithful reporting with good structure and ten quiz questions. Q4 asks about 'surreptitious' which appears in source but not in the rewritten passage, slightly unfair for a passage-based quiz. Analogies (arms race, doping scandal) are good; passage is accurate and engaging.
🥉 openrouter:baidu/ernie-4.5-vl-424b-a47b view ↗ 8788688 7 0.43¢ 7.9 Solid factual grounding and clean structure. Only 3 quiz questions (vs 10 for others), and Q2's 'distil' answer choosing 'replicate' over 'condense' is debatable. Analogies (recipe, plagiarism) are decent but generic; passage is accurate and engaging.
Scores 0–10 · 🟩 ≥8 · 🟨 4–7 · 🟥 <4 · sorted by Overall. Hover a column header for its full rubric name.

💰 Cost-adjusted ranking (value for money)

#ModelQualityCostValue
🥇 openrouter:deepseek/deepseek-v4-pro 9 0.62¢ 8.7
🥈 openrouter:baidu/ernie-4.5-vl-424b-a47b 7 0.43¢ 7.9
🥉 anthropic:claude-haiku-4-5-20251001 8 2.57¢ 6.8

🪙 Best under 3¢ (cost < $0.03), ranked by quality

#ModelQualityCostValue
1 openrouter:deepseek/deepseek-v4-pro 9 0.62¢ 8.7
2 anthropic:claude-haiku-4-5-20251001 8 2.57¢ 6.8
3 openrouter:baidu/ernie-4.5-vl-424b-a47b 7 0.43¢ 7.9
Strong quality without paying flagship prices — the cheap-and-good picks.