🔬 Model Lab

New run Stored runs ⚖️ Judge verdicts 🧮 Math 📊 Math runs 📄 Benchmark paper 📄 3-model paper 📄 Meta: Will Muse Cause a Spark?

⚖️ Judge verdict

The AI Cold War's Newest Weapon: Stealing the Teacher's Answer Key · Judge: anthropic:claude-opus-4-8 · 2026-09-11T15:01:25 · ✅ saved (50ae34) · 17 models

🔬 View outputs side-by-side → all verdicts →

Note: the judge model is also one of the contestants — scores may carry self-preference bias.
🏆 Quality winner: anthropic:claude-opus-4-8
Run 91e0a8 (Opus) wins on the combination of flawless fact-grounding (correctly stresses distilled models are weaker yet still cheap), mechanism-true analogies (tasting a dish thousands of times; memorizing a grandmaster's recorded games), a complete and genuinely SAT-caliber rewritten passage that closes on a real intellectual open question, and a ten-item quiz with precise trap taxonomy and clean evidence-pairing tied to lines actually in its passage. Runners-up e2507b and 2eb734 are nearly as strong—e2507b has superb prose and precise conservatory/pharma analogies, 2eb734 has the standout Tu-4/B-29 reverse-engineering analogy—but 91e0a8 edges them on passage quality and consistency. Candidates that fabricated figures (fe7a5d's 16M/24k numbers, 0313fc's $1T, a95998's 'state-linked') or under-delivered quizzes (f95e70 with one question, a0132a and a-vl variant with three) rank lower.
💰 Best value: openrouter:deepseek/deepseek-v4-pro — quality 9/10 at 0.62¢ → value 8.7/10
Value = 0.7·quality + 0.3·cheapness. Quality stays dominant, so a cheap-but-weak output can't win on price alone.

Quality ranking

#ModelFactsAnalogyAge-fitClarityQuizGlossaryPassageOverallCostValueNotes
🥇 anthropic:claude-sonnet-4-6 view ↗ 8999998 9 8.57¢ 6.6 Excellent depth with strong same-mechanism analogies (chef ordering every dish, recording performances). Invents specific figures (16 million exchanges, 24,000 accounts) not in source—a fact-grounding risk, though plausible. Ten well-crafted SAT-style questions with evidence pairs; passage truncated but engaging.
🥈 openrouter:deepseek/deepseek-v4-pro view ↗ 9899989 9 0.62¢ 8.7 Faithful, well-organized, with a complete polished rewritten passage. Tu-4/B-29 Cold War analogy is precise and shares the reverse-engineering mechanism. Strong 10-question quiz with good vocab-in-context and evidence pairing. Minor: quiz asks about 'How To Think About It' tone, relying on the rewritten sections.
🥉 openrouter:minimax/minimax-m2.7 view ↗ 9899889 9 0.62¢ 8.7 Very faithful, careful to note distilled models don't match originals. University/answer-key analogy is apt. Complete strong passage. Quiz Q4 asks about 'leverage' which appears in source—good. Well-balanced and accurate throughout with clear structure.
4 openrouter:moonshotai/kimi-k2.6 view ↗ 9999989 9 9.62¢ 6.6 Highly faithful, sophisticated prose pitched perfectly for older teens. Conservatory/pharmaceutical analogies precisely share the incentive-structure mechanism. Complete, richly written passage that integrates McGuire quotes accurately. Excellent 10-question quiz with careful evidence pairing. One of the strongest.
5 anthropic:claude-opus-4-8 view ↗ 99999910 9 15.59¢ 6.6 Outstanding across the board: faithful, precise about distilled models being weaker, complete and beautifully written SAT-style passage ending on a genuine open question. Recipe-tasting and chess-game-memorizing analogies match the mechanism. Ten sharp questions with excellent trap explanations and evidence pairing. Best overall.
6 openrouter:mistralai/mistral-large-2512 view ↗ 8788888 8 0.85¢ 8.0 Solid, faithful coverage with the marathon-sweatband analogy (decent mechanism match). Adds some speculative claims (US firms pulling back from collaborations) but flags them as bigger-picture. Complete passage, competent 10-question quiz. Some analogies (cheat code save-scumming) don't share the distillation mechanism well.
7 openrouter:z-ai/glm-5-turbo view ↗ 9899889 8 3.40¢ 6.8 Very faithful with strong music-producer/scouting analogies and detailed trap taxonomy in quiz explanations. However, quiz Q4 asks about 'sharp' and Q9/Q10 reference lines ('close the competitive gap...partly because of export controls') that appear in its own rewritten passage, not verbatim source—slight self-referential risk. Passage truncated but excellent.
8 openrouter:openai/gpt-5.4-nano view ↗ 9788888 8 0.63¢ 8.0 Notably careful and faithful, repeatedly attributing claims to 'the article' and preserving nuance (distilled models lag but still shift power). Export-controls-as-lock analogy is apt. Complete passage, strong 10-question quiz with good evidence pairing. Analogies solid though the sports one is a touch generic.
9 openrouter:stepfun/step-3.7-flash view ↗ 6888887 7 2.09¢ 6.1 Good analogies and structure but overstates facts: calls Chinese groups 'state-linked' and says memo 'specifically names three firms' (Anthropic named them, not the memo). Invents specific dates. Quiz is decent with SAT-style traps but the passage is truncated. Fact drift lowers overall.
10 openrouter:google/gemini-3.1-flash-lite view ↗ 8788778 7 0.59¢ 7.3 Solid, accurate, complete passage with chef/recipe and capture-the-flag analogies. Quiz Q5/Q6 on author tone rely on the rewritten passage rather than source, slightly weakening fairness. Only 4 glossary terms. Competent but not standout.
11 openrouter:x-ai/grok-4.3 view ↗ 8688788 7 1.64¢ 6.7 Faithful and clean but analogies (reverse-engineering, hidden-camera scouting) are somewhat generic. Sections are brief. Complete passage. Quiz Q5 asks about 'industrial-scale' meaning—fair—but overall quiz is competent rather than distinguished. Solid mid-tier entry.
12 openrouter:meta-llama/llama-4-maverick view ↗ 8688788 7 0.26¢ 7.9 Careful, faithful reporting that appropriately hedges ('the article states'). Analogies (textbook/study-guide, film scouting) are okay but not especially mechanism-tight. Good glossary. Quiz is fair with vocab-in-context and evidence pairing, but slightly less punchy than the top tier. Complete passage.
13 openrouter:qwen/qwen3-max-thinking view ↗ 8788588 6 1.02¢ 6.0 Clean, accurate explainer with good analogies and a complete polished passage. Fatal weakness: only ONE quiz question provided despite requesting multiple MCQs. That single question is well-built, but the quiz dimension is severely underdelivered.
14 openrouter:amazon/nova-pro-v1 view ↗ 8567666 6 1.34¢ 6.0 Accurate but thin and somewhat dumbed-down. Analogies (copying homework, chess, doping) are generic and don't capture the query-API distillation mechanism. Only 3 glossary terms. Quiz Q9/Q10 reference claims from the rewritten passage but feel circular. Serviceable but bland.
15 anthropic:claude-haiku-4-5-20251001 view ↗ 6788677 6 2.57¢ 5.4 Vivid MIT-lab and Twitch-scraping analogies but introduces unsupported specifics: fabricates DeepSeek release timing details, asserts '$1 trillion global industry' figure not in source, and claims the memo named three firms (Anthropic did). Quiz Q4 asks about 'principal' but source says 'principally'—minor slip. Fact drift hurts it.
16 openrouter:baidu/ernie-4.5-vl-424b-a47b view ↗ 7677567 5 0.43¢ 6.5 Concise and accurate but thin: only 3 glossary terms and only 3 quiz questions (rubric expects a fuller set). Analogies (recipe heist, counterfeit goods, generic drug) are decent but not deeply mechanism-matched. Complete passage. Underdelivers on quiz breadth.
17 view ↗ 0000000 0 – – Placeholder—ignore.
Scores 0–10 · 🟩 ≥8 · 🟨 4–7 · 🟥 <4 · sorted by Overall. Hover a column header for its full rubric name.

💰 Cost-adjusted ranking (value for money)

#ModelQualityCostValue
🥇 openrouter:deepseek/deepseek-v4-pro 9 0.62¢ 8.7
🥈 openrouter:minimax/minimax-m2.7 9 0.62¢ 8.7
🥉 openrouter:mistralai/mistral-large-2512 8 0.85¢ 8.0
4 openrouter:openai/gpt-5.4-nano 8 0.63¢ 8.0
5 openrouter:meta-llama/llama-4-maverick 7 0.26¢ 7.9
6 openrouter:google/gemini-3.1-flash-lite 7 0.59¢ 7.3
7 openrouter:z-ai/glm-5-turbo 8 3.40¢ 6.8
8 openrouter:x-ai/grok-4.3 7 1.64¢ 6.7
9 anthropic:claude-sonnet-4-6 9 8.57¢ 6.6
10 openrouter:moonshotai/kimi-k2.6 9 9.62¢ 6.6
11 anthropic:claude-opus-4-8 9 15.59¢ 6.6
12 openrouter:baidu/ernie-4.5-vl-424b-a47b 5 0.43¢ 6.5
13 openrouter:stepfun/step-3.7-flash 7 2.09¢ 6.1
14 openrouter:qwen/qwen3-max-thinking 6 1.02¢ 6.0
15 openrouter:amazon/nova-pro-v1 6 1.34¢ 6.0
16 anthropic:claude-haiku-4-5-20251001 6 2.57¢ 5.4

🪙 Best under 3¢ (cost < $0.03), ranked by quality

#ModelQualityCostValue
1 openrouter:deepseek/deepseek-v4-pro 9 0.62¢ 8.7
2 openrouter:minimax/minimax-m2.7 9 0.62¢ 8.7
3 openrouter:mistralai/mistral-large-2512 8 0.85¢ 8.0
4 openrouter:openai/gpt-5.4-nano 8 0.63¢ 8.0
5 openrouter:meta-llama/llama-4-maverick 7 0.26¢ 7.9
6 openrouter:google/gemini-3.1-flash-lite 7 0.59¢ 7.3
7 openrouter:x-ai/grok-4.3 7 1.64¢ 6.7
8 openrouter:stepfun/step-3.7-flash 7 2.09¢ 6.1
9 openrouter:qwen/qwen3-max-thinking 6 1.02¢ 6.0
10 openrouter:amazon/nova-pro-v1 6 1.34¢ 6.0
11 anthropic:claude-haiku-4-5-20251001 6 2.57¢ 5.4
12 openrouter:baidu/ernie-4.5-vl-424b-a47b 5 0.43¢ 6.5
Strong quality without paying flagship prices — the cheap-and-good picks.