🔬 Model Lab

New run Stored runs ⚖️ Judge verdicts 🧮 Math 📊 Math runs 📄 Benchmark paper 📄 3-model paper 📄 Meta: Will Muse Cause a Spark?

⚖️ Judge verdict

The AI Cold War's Newest Weapon: Stealing the Teacher's Answer Key · Judge: anthropic:claude-opus-4-8 · 2026-09-05T00:17:38 · ✅ saved (f65374) · 15 models

🔬 View outputs side-by-side → all verdicts →

🏆 Quality winner: openrouter:deepseek/deepseek-v4-pro
The DeepSeek entry combines the most distinctive, mechanism-matching analogy (Soviet Tu-4 reverse-engineered from the B-29, which precisely mirrors how distilled models approximate but don't match originals) with a complete, fully answerable 10-question SAT-style quiz, a precise six-term glossary, and an original, engaging, non-truncated rewritten passage. It edges out c2b157 and e2507b, which are equally faithful and well-structured but have slightly more generic analogies, and beats the strong Claude/GPT entries because those passages are truncated. Weaker candidates either supplied only 1-3 quiz questions (f95e70, a0132a) or relied on generic chess/homework analogies (fa777b).
💰 Best value: openrouter:deepseek/deepseek-v4-pro — quality 9/10 at 0.62¢ → value 8.7/10
Value = 0.7·quality + 0.3·cheapness. Quality stays dominant, so a cheap-but-weak output can't win on price alone.

Quality ranking

#ModelFactsAnalogyAge-fitClarityQuizGlossaryPassageOverallCostValueNotes
🥇 openrouter:deepseek/deepseek-v4-pro view ↗ 9999999 9 0.62¢ 8.7 Excellent fact fidelity with strong Cold War Tu-4/B-29 analogy sharing the reverse-engineering mechanism. Full 10-question quiz is fair and answerable from the rewritten passage. Glossary precise; passage is engaging and accurate.
🥈 openrouter:minimax/minimax-m2.7 view ↗ 9899999 9 0.62¢ 8.7 Very faithful, well-organized with a sharp university/answer-key analogy and F1 comparison. Quiz answerable from the passage with good distractors and clean evidence-pairing. Glossary strong including AI safety guardrails.
🥉 openrouter:moonshotai/kimi-k2.6 view ↗ 9899989 9 9.62¢ 6.6 Highly faithful, incorporates McGuire and the entity-list detail into an original, sophisticated passage. Conservatory and pharma analogies share the core mechanism well. Strong, trap-aware quiz with accurate evidence pairing.
4 anthropic:claude-haiku-4-5-20251001 view ↗ 9899998 9 2.57¢ 7.5 Very faithful with a rich, well-structured passage (though truncated at the very end). Strong nuclear-arms and doping-scandal analogies capturing the mechanism. Excellent, trap-aware quiz and thorough glossary.
5 openrouter:mistralai/mistral-large-2512 view ↗ 8899888 8 0.85¢ 8.0 Strong analogies (marathon sweatband, supply chain heist) and solid structure. Q5 asks about 'surreptitious' which appears in source but not clearly in the rewritten passage, a minor calibration issue. Otherwise faithful and readable.
6 openrouter:google/gemini-3.1-flash-lite view ↗ 8888888 8 0.59¢ 8.0 Solid chef-recipe and blueprint analogies that fit the mechanism. Faithful passage and well-constructed 10-question quiz with reasonable evidence pairing. Clear, scannable structure.
7 openrouter:z-ai/glm-5-turbo view ↗ 8898897 8 3.40¢ 6.8 Strong analogies and detailed trap-labeled quiz, but includes a fabricated vocab question ('sharp' in 'drew a sharp distinction') that relies on the rewritten passage's own added phrasing, and the passage is truncated. Otherwise rich and well-grounded.
8 openrouter:x-ai/grok-4.3 view ↗ 9788888 8 1.64¢ 7.4 Faithful and complete with a clean, original passage incorporating McGuire. Analogies (playbook cameras, reverse-engineering) are adequate but somewhat generic. Solid trap-aware quiz answerable from passage.
9 openrouter:bytedance-seed/seed-2.0-lite view ↗ 8788888 8 1.73¢ 7.4 Faithful, complete passage and strong SAT-style quiz with good traps. MIT-lab and gaming analogies are decent but a bit forced. Slightly overstates figures ('$1 trillion industry') but broadly accurate.
10 openrouter:stepfun/step-3.7-flash view ↗ 8888787 7 2.09¢ 6.1 Good SAT-tutor-style questions and analogies, but the rewritten passage is truncated mid-sentence, hurting completeness. Q5 vocab 'surreptitious' not present in the truncated passage. Otherwise solid content.
11 openrouter:qwen/qwen3-max-thinking view ↗ 8788588 7 1.02¢ 6.7 Clean, accurate passage and decent analogies (answer key, counterfeit handbags). Major weakness: only ONE quiz question provided despite the 10-question expectation, sharply reducing quiz value. Glossary and hook strong.
12 openrouter:meta-llama/llama-4-maverick view ↗ 7687776 7 0.26¢ 7.9 Accurate but leans heavily on 'the article says' framing rather than crisp explanation. Analogies serviceable (locked lab, film scouting). Quiz is fair but Q5 'unacceptable' meaning-question is odd. Passage is competent, less engaging.
13 openrouter:openai/gpt-5.4-nano view ↗ 8688888 7 0.63¢ 7.3 Accurate and readable; passage stays close to source with appropriate hedging. Analogies (textbook/locked equipment) are okay but generic. Quiz is fair and answerable. Solid but unremarkable.
14 openrouter:amazon/nova-pro-v1 view ↗ 7566666 6 1.34¢ 6.0 Generic chess/doping/homework analogies that don't capture the distillation mechanism precisely. Q1 evidence-pairing references a caption line ('arms race') as though authored. Content is accurate but thin and less engaging; glossary only three terms.
15 openrouter:baidu/ernie-4.5-vl-424b-a47b view ↗ 7777577 6 0.43¢ 7.2 Chef-recipe and counterfeit analogies are fine and grounded. But only 3 quiz questions provided and glossary only 3 terms, reducing depth. Q2 vocab answer 'replicate' over 'condense' is arguable. Passage accurate but compressed.
Scores 0–10 · 🟩 ≥8 · 🟨 4–7 · 🟥 <4 · sorted by Overall. Hover a column header for its full rubric name.

💰 Cost-adjusted ranking (value for money)

#ModelQualityCostValue
🥇 openrouter:deepseek/deepseek-v4-pro 9 0.62¢ 8.7
🥈 openrouter:minimax/minimax-m2.7 9 0.62¢ 8.7
🥉 openrouter:mistralai/mistral-large-2512 8 0.85¢ 8.0
4 openrouter:google/gemini-3.1-flash-lite 8 0.59¢ 8.0
5 openrouter:meta-llama/llama-4-maverick 7 0.26¢ 7.9
6 anthropic:claude-haiku-4-5-20251001 9 2.57¢ 7.5
7 openrouter:x-ai/grok-4.3 8 1.64¢ 7.4
8 openrouter:bytedance-seed/seed-2.0-lite 8 1.73¢ 7.4
9 openrouter:openai/gpt-5.4-nano 7 0.63¢ 7.3
10 openrouter:baidu/ernie-4.5-vl-424b-a47b 6 0.43¢ 7.2
11 openrouter:z-ai/glm-5-turbo 8 3.40¢ 6.8
12 openrouter:qwen/qwen3-max-thinking 7 1.02¢ 6.7
13 openrouter:moonshotai/kimi-k2.6 9 9.62¢ 6.6
14 openrouter:stepfun/step-3.7-flash 7 2.09¢ 6.1
15 openrouter:amazon/nova-pro-v1 6 1.34¢ 6.0

🪙 Best under 3¢ (cost < $0.03), ranked by quality

#ModelQualityCostValue
1 openrouter:deepseek/deepseek-v4-pro 9 0.62¢ 8.7
2 openrouter:minimax/minimax-m2.7 9 0.62¢ 8.7
3 anthropic:claude-haiku-4-5-20251001 9 2.57¢ 7.5
4 openrouter:mistralai/mistral-large-2512 8 0.85¢ 8.0
5 openrouter:google/gemini-3.1-flash-lite 8 0.59¢ 8.0
6 openrouter:x-ai/grok-4.3 8 1.64¢ 7.4
7 openrouter:bytedance-seed/seed-2.0-lite 8 1.73¢ 7.4
8 openrouter:meta-llama/llama-4-maverick 7 0.26¢ 7.9
9 openrouter:openai/gpt-5.4-nano 7 0.63¢ 7.3
10 openrouter:qwen/qwen3-max-thinking 7 1.02¢ 6.7
11 openrouter:stepfun/step-3.7-flash 7 2.09¢ 6.1
12 openrouter:baidu/ernie-4.5-vl-424b-a47b 6 0.43¢ 7.2
13 openrouter:amazon/nova-pro-v1 6 1.34¢ 6.0
Strong quality without paying flagship prices — the cheap-and-good picks.