🔬 Model Lab

New run Stored runs ⚖️ Judge verdicts 🧮 Math 📊 Math runs 📄 Benchmark paper 📄 3-model paper 📄 Meta: Will Muse Cause a Spark?

⚖️ Judge verdict

The AI Cold War's Newest Weapon: Stealing the Teacher's Answer Key · Judge: anthropic:claude-opus-4-8 · 2026-09-25T06:33:32 · ✅ saved (d16317) · 22 models

🔬 View outputs side-by-side → all verdicts →

🏆 Quality winner: openrouter:deepseek/deepseek-v4-pro
Run 2eb734 combines flawless fact-grounding with the best analogies in the field—the Tu-4/B-29 Cold War reverse-engineering example precisely mirrors distillation's mechanism (a cheaper copy that doesn't match the original but closes a strategic gap), which is exactly the article's point. It has a complete 10-question quiz with well-constructed evidence-pairing and vocab-in-context items all answerable from its original, complete, engaging passage, plus a rich six-term glossary. c2b157 and e2507b are close runners-up but slightly less distinctive in analogy; several otherwise-strong entries lost points for truncated rewritten passages.
💰 Best value: openrouter:deepseek/deepseek-v4-pro — quality 9/10 at 0.62¢ → value 8.7/10
Value = 0.7·quality + 0.3·cheapness. Quality stays dominant, so a cheap-but-weak output can't win on price alone.

Quality ranking

#ModelFactsAnalogyAge-fitClarityQuizGlossaryPassageOverallCostValueNotes
🥇 openrouter:deepseek/deepseek-v4-pro view ↗ 10999999 9 0.62¢ 8.7 Facts are accurate and richly detailed (Tu-4 analogy is apt and shares mechanism of reverse-engineering with cost gap). Full 10-question quiz with strong SAT-style evidence pairing and vocab-in-context items answerable from the rewritten passage. Original, engaging passage. Very strong all around.
🥈 openrouter:minimax/minimax-m2.7 view ↗ 10899999 9 0.62¢ 8.7 Highly accurate, well-structured, with a clear F1-car and university analogy that captures the cost-vs-effort mechanism. Full 10-question quiz well-anchored to the passage; glossary precise. Passage is original and readable.
🥉 openrouter:moonshotai/kimi-k2.6 view ↗ 9999999 9 9.62¢ 6.6 Excellent conservatory/pharma analogies matching the same mechanism. Rewritten passage is complete, original and sophisticated. Quiz is well-crafted with fair evidence-pairing. Faithful to source facts including McGuire and Liu Pengyu.
4 openrouter:mistralai/mistral-large-2512 view ↗ 9899889 8 0.85¢ 8.0 Solid, faithful, engaging. 'Surreptitious' vocab question (Q5) is grounded in the source's word but the rewritten passage itself doesn't use it, a minor mismatch. Good analogies (cheat code, supply chain heist) and a strong original passage.
5 openrouter:stepfun/step-3.7-flash view ↗ 9998898 8 2.09¢ 6.8 Strong content and excellent SAT analogies (valedictorian, burger joint). Q5 tests 'surreptitious' which appears in source but not clearly in the truncated rewritten passage. The rewritten_passage is cut off mid-sentence, hurting completeness.
6 openrouter:qwen/qwen3-max-thinking view ↗ 9898887 8 1.02¢ 7.4 Good quiz and glossary; answer-key analogy is apt. But rewritten_passage is truncated mid-sentence, reducing quality. Q5 'surreptitious' vocab relies on a word that may not appear in the truncated passage.
7 openrouter:x-ai/grok-4.3 view ↗ 9788888 8 1.64¢ 7.4 Faithful and clean; analogies (reverse-engineering, scouting playbook) are decent but somewhat generic. Complete passage, solid full quiz. Slightly less vivid than top entries but accurate throughout.
8 openrouter:google/gemini-3.1-flash-lite view ↗ 8899988 8 0.59¢ 8.0 Clean structure, chef-recipe and capture-the-flag analogies fit reasonably. Full 10-question quiz with good SAT scaffolding. Complete, engaging passage. Solid and well-rounded.
9 openrouter:z-ai/glm-5-turbo view ↗ 8899898 8 3.40¢ 6.8 Strong music-producer/scouting analogies and clear structure. Q4 tests 'sharp' ('drew a sharp distinction') which the writer's own rewritten passage uses—clever and self-consistent. Rewritten passage truncated at end, a small deduction.
10 openrouter:openai/gpt-5.4-nano view ↗ 9799999 8 0.63¢ 8.0 Careful to attribute claims to 'the article,' which is scrupulously faithful though slightly repetitive. Strong 6-term glossary, full solid quiz, complete engaging passage. Analogies (textbook/study-guide, film scouting) are decent.
11 anthropic:claude-haiku-4-5-20251001 view ↗ 9899998 8 2.57¢ 6.8 Very faithful with named companies and dates. Arms-race and doping analogies work; doping is slightly generic. Full quiz with excellent evidence pairing and vocab items. Rewritten passage is complete and strong but ends slightly truncated.
12 openrouter:amazon/nova-pro-v1 view ↗ 8677777 7 1.34¢ 6.7 Accurate but thinner. Chess/doping analogies are generic and doping doesn't share the distillation mechanism well. Passage is a compressed rewrite that's serviceable. Only 3 glossary terms; quiz is fine but less rigorous.
13 openrouter:meta-llama/llama-4-maverick view ↗ 8687887 7 0.26¢ 7.9 Accurate but analogies are brief and generic (student copying, reverse-engineering). Explanation sections are short. Quiz is complete and fair. Passage is competent but less vivid than leaders.
14 openrouter:bytedance-seed/seed-2.0-lite view ↗ 7688787 7 1.73¢ 6.7 Serviceable and mostly accurate but adds an unsourced '$1 trillion industry' figure. Analogies are ok. Q4 tests 'principal'/'distil' meaning that may not be perfectly grounded in the rewritten passage. Reasonable overall.
15 openrouter:baidu/ernie-4.5-vl-424b-a47b view ↗ 7687677 6 0.43¢ 7.2 Accurate but thin: only 3 quiz questions and 3 glossary terms. Analogies (chef, counterfeit) are generic. Passage is competent but compressed. Least developed of the fuller entries.
16 view ↗ 0000000 0 – – Duplicate id placeholder; ignore.
17 view ↗ 0000000 0 – – placeholder
18 view ↗ 0000000 0 – – placeholder
19 view ↗ 0000000 0 – – placeholder
20 view ↗ 0000000 0 – – placeholder
21 view ↗ 0000000 0 – – placeholder
22 view ↗ 0000000 0 – – placeholder
Scores 0–10 · 🟩 ≥8 · 🟨 4–7 · 🟥 <4 · sorted by Overall. Hover a column header for its full rubric name.

💰 Cost-adjusted ranking (value for money)

#ModelQualityCostValue
🥇 openrouter:deepseek/deepseek-v4-pro 9 0.62¢ 8.7
🥈 openrouter:minimax/minimax-m2.7 9 0.62¢ 8.7
🥉 openrouter:mistralai/mistral-large-2512 8 0.85¢ 8.0
4 openrouter:google/gemini-3.1-flash-lite 8 0.59¢ 8.0
5 openrouter:openai/gpt-5.4-nano 8 0.63¢ 8.0
6 openrouter:meta-llama/llama-4-maverick 7 0.26¢ 7.9
7 openrouter:qwen/qwen3-max-thinking 8 1.02¢ 7.4
8 openrouter:x-ai/grok-4.3 8 1.64¢ 7.4
9 openrouter:baidu/ernie-4.5-vl-424b-a47b 6 0.43¢ 7.2
10 openrouter:stepfun/step-3.7-flash 8 2.09¢ 6.8
11 openrouter:z-ai/glm-5-turbo 8 3.40¢ 6.8
12 anthropic:claude-haiku-4-5-20251001 8 2.57¢ 6.8
13 openrouter:amazon/nova-pro-v1 7 1.34¢ 6.7
14 openrouter:bytedance-seed/seed-2.0-lite 7 1.73¢ 6.7
15 openrouter:moonshotai/kimi-k2.6 9 9.62¢ 6.6

🪙 Best under 3¢ (cost < $0.03), ranked by quality

#ModelQualityCostValue
1 openrouter:deepseek/deepseek-v4-pro 9 0.62¢ 8.7
2 openrouter:minimax/minimax-m2.7 9 0.62¢ 8.7
3 openrouter:mistralai/mistral-large-2512 8 0.85¢ 8.0
4 openrouter:google/gemini-3.1-flash-lite 8 0.59¢ 8.0
5 openrouter:openai/gpt-5.4-nano 8 0.63¢ 8.0
6 openrouter:qwen/qwen3-max-thinking 8 1.02¢ 7.4
7 openrouter:x-ai/grok-4.3 8 1.64¢ 7.4
8 openrouter:stepfun/step-3.7-flash 8 2.09¢ 6.8
9 anthropic:claude-haiku-4-5-20251001 8 2.57¢ 6.8
10 openrouter:meta-llama/llama-4-maverick 7 0.26¢ 7.9
11 openrouter:amazon/nova-pro-v1 7 1.34¢ 6.7
12 openrouter:bytedance-seed/seed-2.0-lite 7 1.73¢ 6.7
13 openrouter:baidu/ernie-4.5-vl-424b-a47b 6 0.43¢ 7.2
Strong quality without paying flagship prices — the cheap-and-good picks.