🔬 Model Lab

New run Stored runs ⚖️ Judge verdicts 🧮 Math 📊 Math runs 📄 Benchmark paper 📄 3-model paper 📄 Meta: Will Muse Cause a Spark?

⚖️ Judge verdict

The AI Cold War's Newest Weapon: Stealing the Teacher's Answer Key · Judge: anthropic:claude-opus-4-8 · 2026-09-12T02:56:41 · ✅ saved (520a1e) · 17 models

🔬 View outputs side-by-side → all verdicts →

Note: the judge model is also one of the contestants — scores may carry self-preference bias.
🏆 Quality winner: anthropic:claude-opus-4-8
Claude Opus (91e0a8) wins on the strength of its reading passage—the most polished, original, and genuinely SAT-style prose of any entry, with a memorable closing rhetorical question. Its analogies (tasting a dish thousands of times; a chess student memorizing recorded games) precisely capture distillation's mechanism of copying behavior via outputs rather than stealing source. Its full 10-question quiz uses vocab words that actually appear in its own passage ('distill', 'frontier') and its evidence-pairing questions are internally consistent—a discipline several rivals (48b0f2, fe7a5d) failed by writing questions about words absent from their passages or by inventing figures. It is uniformly excellent across all seven dimensions with no fact-grounding fabrications, edging out the very strong c2b157 and e2507b.
💰 Best value: openrouter:minimax/minimax-m2.7 — quality 9/10 at 0.62¢ → value 8.7/10
Value = 0.7·quality + 0.3·cheapness. Quality stays dominant, so a cheap-but-weak output can't win on price alone.

Quality ranking

#ModelFactsAnalogyAge-fitClarityQuizGlossaryPassageOverallCostValueNotes
🥇 anthropic:claude-sonnet-4-6 view ↗ 89999108 9 8.57¢ 6.6 Excellent explainer with strong analogies (chef ordering every dish, recording a band). Fabricates specific numbers (16 million exchanges, 24,000 accounts) not in source, a minor fact-grounding penalty. Full 10-question quiz with strong SAT-style vocab and evidence-pairing questions; passage is truncated.
🥈 view ↗ 9899999 9 – – Very faithful to source; Tu-4/B-29 Cold War analogy is apt and mechanism-matched. Complete, engaging rewritten passage. Quiz well-constructed with fair tone and inference questions. One of the strongest overall.
🥉 openrouter:minimax/minimax-m2.7 view ↗ 9899999 9 0.62¢ 8.7 Highly faithful and well-organized. University/answer-key and F1 analogies match mechanism well. Complete, original, engaging passage; quiz vocab words ('leverage', 'surreptitious') genuinely appear in the source. Excellent all-around.
4 openrouter:moonshotai/kimi-k2.6 view ↗ 9999899 9 9.62¢ 6.6 Excellent conservatory/pharma analogies with matched incentive structure. Faithful, richly written passage that incorporates McGuire and embassy quotes. Quiz solid; vocab 'frontier'/'lightweight' appear in passage. One of the best.
5 anthropic:claude-opus-4-8 view ↗ 99999910 9 15.59¢ 6.6 Outstanding entry. Recipe-tasting and chess-memorization analogies are precise and mechanism-matched. Beautifully written, complete, original passage with a strong closing question. Full 10-question quiz with excellent vocab-in-context and evidence-pairing; vocab words genuinely appear. Top tier.
6 openrouter:mistralai/mistral-large-2512 view ↗ 8788888 8 0.85¢ 8.0 Solid and accurate; 'save scumming' cheat-code analogy is a bit forced but engaging. Good structure and passage. Quiz asks about a 'surreptitious' vocab word but the rewritten passage doesn't actually use it, a minor internal-consistency slip.
7 openrouter:google/gemini-3.1-flash-lite view ↗ 8788888 8 0.59¢ 8.0 Faithful, clear, well-structured with a complete engaging passage. Chef/blueprint analogies are decent. Full quiz with good vocab and evidence-pairing questions. Solid mid-upper tier entry.
8 openrouter:meta-llama/llama-4-maverick view ↗ 8788878 8 0.26¢ 8.6 Clean, accurate, complete passage. Recipe/capture-the-flag analogies work. Full quiz with fair questions and good distractor logic. Smaller glossary (4 terms) but solid. Reliable upper-mid entry.
9 openrouter:bytedance-seed/seed-2.0-lite view ↗ 8788888 8 1.73¢ 7.4 Faithful and complete. MIT/Twitch analogies are vivid though slightly hyperbolic. Full quiz; Q4 asks about 'principal' which the source does use ('principally based in China'). Good glossary and passage. Solid.
10 openrouter:openai/gpt-5.4-nano view ↗ 8788888 8 0.63¢ 8.0 Careful, source-hedged prose ('the article says'), which is accurate but slightly clunky for an explainer. Lock/scouting analogies are decent. Complete passage and full quiz with fair questions. Reliable.
11 anthropic:claude-haiku-4-5-20251001 view ↗ 8899998 8 2.57¢ 6.8 Strong, engaging, well-structured. Nuclear-arms-race and doping analogies fit well. Complete quiz with good vocab and evidence-pairing. Passage truncated at end. Adds a bit of interpretive framing but stays grounded.
12 openrouter:stepfun/step-3.7-flash view ↗ 6888888 7 2.09¢ 6.1 Strong analogies and clear structure. Fabricates 'state-linked' and claims memo 'specifically names' three firms (source attributes to Anthropic, not memo); overstates certainty. Passage truncated. Quiz decent but rests on some invented details.
13 openrouter:z-ai/glm-5-turbo view ↗ 6899698 7 3.40¢ 6.1 Very well-written with strong analogies and glossary. However, quiz Q4 asks about 'sharp' and Q6 references phrasing ('close the competitive gap...partly because') that don't match the actual rewritten passage—questions were written for a different passage version. Fact 'billions to train from scratch' invented. Passage truncated.
14 openrouter:x-ai/grok-4.3 view ↗ 8688788 7 1.64¢ 6.7 Concise and accurate; complete passage. Analogies (reverse-engineering, playbook scouting) are generic but serviceable. Quiz Q5 vocab 'industrial-scale' is reasonable. Solid but not standout.
15 openrouter:qwen/qwen3-max-thinking view ↗ 8787488 6 1.02¢ 6.0 Good hook, accurate content, complete passage. Analogies (answer key, handbag counterfeit) are reasonable. Major weakness: only ONE quiz question provided instead of a full set, sharply limiting quiz utility.
16 openrouter:amazon/nova-pro-v1 view ↗ 8466666 5 1.34¢ 5.3 Accurate but bland. Generic analogies (chess, doping) don't capture the distillation mechanism. Some quiz questions (tone, central idea) reference the explainer rather than the passage. Thinner glossary and shorter passage than peers.
17 openrouter:baidu/ernie-4.5-vl-424b-a47b view ↗ 7677467 5 0.43¢ 6.5 Accurate enough with a decent recipe-heist hook, but thin: only 3 quiz questions, 3 glossary terms, and a short passage. Analogies are generic. Missing empty 'bullets' arrays and shorter overall than competitors.
Scores 0–10 · 🟩 ≥8 · 🟨 4–7 · 🟥 <4 · sorted by Overall. Hover a column header for its full rubric name.

💰 Cost-adjusted ranking (value for money)

#ModelQualityCostValue
🥇 openrouter:minimax/minimax-m2.7 9 0.62¢ 8.7
🥈 openrouter:meta-llama/llama-4-maverick 8 0.26¢ 8.6
🥉 openrouter:mistralai/mistral-large-2512 8 0.85¢ 8.0
4 openrouter:google/gemini-3.1-flash-lite 8 0.59¢ 8.0
5 openrouter:openai/gpt-5.4-nano 8 0.63¢ 8.0
6 openrouter:bytedance-seed/seed-2.0-lite 8 1.73¢ 7.4
7 anthropic:claude-haiku-4-5-20251001 8 2.57¢ 6.8
8 openrouter:x-ai/grok-4.3 7 1.64¢ 6.7
9 anthropic:claude-sonnet-4-6 9 8.57¢ 6.6
10 openrouter:moonshotai/kimi-k2.6 9 9.62¢ 6.6
11 anthropic:claude-opus-4-8 9 15.59¢ 6.6
12 openrouter:baidu/ernie-4.5-vl-424b-a47b 5 0.43¢ 6.5
13 openrouter:stepfun/step-3.7-flash 7 2.09¢ 6.1
14 openrouter:z-ai/glm-5-turbo 7 3.40¢ 6.1
15 openrouter:qwen/qwen3-max-thinking 6 1.02¢ 6.0
16 openrouter:amazon/nova-pro-v1 5 1.34¢ 5.3

🪙 Best under 3¢ (cost < $0.03), ranked by quality

#ModelQualityCostValue
1 openrouter:minimax/minimax-m2.7 9 0.62¢ 8.7
2 openrouter:meta-llama/llama-4-maverick 8 0.26¢ 8.6
3 openrouter:mistralai/mistral-large-2512 8 0.85¢ 8.0
4 openrouter:google/gemini-3.1-flash-lite 8 0.59¢ 8.0
5 openrouter:openai/gpt-5.4-nano 8 0.63¢ 8.0
6 openrouter:bytedance-seed/seed-2.0-lite 8 1.73¢ 7.4
7 anthropic:claude-haiku-4-5-20251001 8 2.57¢ 6.8
8 openrouter:x-ai/grok-4.3 7 1.64¢ 6.7
9 openrouter:stepfun/step-3.7-flash 7 2.09¢ 6.1
10 openrouter:qwen/qwen3-max-thinking 6 1.02¢ 6.0
11 openrouter:baidu/ernie-4.5-vl-424b-a47b 5 0.43¢ 6.5
12 openrouter:amazon/nova-pro-v1 5 1.34¢ 5.3
Strong quality without paying flagship prices — the cheap-and-good picks.