🔬 Model Lab

New run Stored runs ⚖️ Judge verdicts 🧮 Math 📊 Math runs 📄 Benchmark paper 📄 3-model paper 📄 Meta: Will Muse Cause a Spark?

⚖️ Judge verdict

The AI Cold War's Newest Weapon: Stealing the Teacher's Answer Key · Judge: anthropic:claude-opus-4-8 · 2026-09-06T13:48:56 · ✅ saved (d03298) · 15 models

🔬 View outputs side-by-side → all verdicts →

🏆 Quality winner: openrouter:minimax/minimax-m2.7
c2b157 (minimax) combines a fully complete, accurate rewritten passage, a rich six-term glossary, and a fair 10-question quiz whose vocab and evidence-pairing items are all answerable from its own passage, with analogies (world-class university, F1 car) that genuinely share the copy-the-output mechanism. Its closest rivals e2507b and 2eb734 are equally strong, but 2eb734's passage and e2507b are also excellent; c2b157 edges ahead on the balance of completeness (no truncation, unlike 0313fc, a95998, and haiku d7fde8-style truncations) and precise, safety-aware glossary. Several otherwise-strong runs (f95e70 one question, a0132a three questions, fa777b generic analogies) fall behind on quiz depth or analogy fidelity.
💰 Best value: openrouter:deepseek/deepseek-v4-pro — quality 9/10 at 0.62¢ → value 8.7/10
Value = 0.7·quality + 0.3·cheapness. Quality stays dominant, so a cheap-but-weak output can't win on price alone.

Quality ranking

#ModelFactsAnalogyAge-fitClarityQuizGlossaryPassageOverallCostValueNotes
🥇 openrouter:deepseek/deepseek-v4-pro view ↗ 10999999 9 0.62¢ 8.7 Excellent fact grounding with the Tu-4/B-29 Cold War analogy sharing the same reverse-engineering mechanism. Full 10-question quiz with strong SAT-style vocab and evidence-pairing items. Rewritten passage is accurate and engaging.
🥈 openrouter:minimax/minimax-m2.7 view ↗ 10899999 9 0.62¢ 8.7 Very accurate, well-structured, with a good glossary including safety guardrails. Quiz is fair and answerable from the passage; vocab items ('leverage','surreptitious') tie back to text. Analogies (university, F1 car) share the copying mechanism.
🥉 openrouter:moonshotai/kimi-k2.6 view ↗ 9999999 9 9.62¢ 6.6 Conservatory/pharma analogies share the incentive-structure mechanism well. Quiz vocab ('frontier','lightweight') and evidence pairing solid. Rewritten passage is complete, accurate, and elegantly written—one of the strongest.
4 anthropic:claude-haiku-4-5-20251001 view ↗ 9899998 9 2.57¢ 7.5 Nuclear arms race and doping analogies fit; explains the 'legal vs illegal' distinction cleanly. Strong 10-question quiz with good vocab and evidence pairing. Rich glossary. Rewritten passage accurate but truncated at the very end.
5 openrouter:mistralai/mistral-large-2512 view ↗ 9899888 8 0.85¢ 8.0 Solid and faithful; 'save scumming' and supply-chain heist analogies are decent. Q5 uses 'surreptitious' as a vocab word but the term doesn't appear in the rewritten passage, a minor answerability issue. Strong prose.
6 openrouter:stepfun/step-3.7-flash view ↗ 9998887 8 2.09¢ 6.8 Strong analogies (SAT tutor, burger recipe) with matching mechanism. Full quiz with trap taxonomy. However the rewritten_passage is truncated mid-sentence, hurting reading_passage_quality.
7 openrouter:google/gemini-3.1-flash-lite view ↗ 9889888 8 0.59¢ 8.0 Chef recipe and capture-the-flag analogies fit reasonably. Full 10-question quiz, clean structure. Accurate rewritten passage. Solid all-around but glossary is 4 terms and slightly generic.
8 openrouter:x-ai/grok-4.3 view ↗ 9799889 8 1.64¢ 7.4 Clean, accurate, complete passage. Playbook-scouting analogy is fine but student-copying is generic. Full quiz answerable from the passage; vocab items grounded. Reliable but analogies less distinctive than top runs.
9 openrouter:openai/gpt-5.4-nano view ↗ 9888888 8 0.63¢ 8.0 Careful, hedged ('the article says') framing keeps facts tight. Textbook/scouting analogies fit. Quiz well-constructed with grounded vocab. Slightly repetitive meta-referencing ('the article') in the passage but accurate.
10 openrouter:z-ai/glm-5-turbo view ↗ 8899697 7 3.40¢ 6.1 Music-producer and scouting analogies are strong. However Q4 tests the word 'sharp' via a phrase ('drew a sharp distinction') that does not appear in the truncated passage—an answerability failure. Rewritten passage is also truncated.
11 openrouter:bytedance-seed/seed-2.0-lite view ↗ 8778787 7 1.73¢ 6.7 MIT-lab and Twitch-scraping analogies are decent. Q4 tests 'principal' but the rewritten passage uses 'principally' rather than 'principal' as adjective—minor mismatch. Otherwise complete and reasonable.
12 openrouter:qwen/qwen3-max-thinking view ↗ 8687589 6 1.02¢ 6.0 Good rewritten passage and glossary, but only ONE quiz question provided, severely undercutting quiz_quality. Analogies (answer key, handbag counterfeit) are okay but the handbag one shares less mechanism.
13 openrouter:amazon/nova-pro-v1 view ↗ 8566666 6 1.34¢ 6.0 Faithful but thin. Chess and doping analogies are generic and don't capture the distillation mechanism. Short glossary (3 terms), brief passage. Q10 quotes a caption-style line as evidence somewhat weakly.
14 openrouter:meta-llama/llama-4-maverick view ↗ 8577777 6 0.26¢ 7.2 Faithful but generic analogies (copying test answers, stealing lesson plans). Quiz functional but somewhat shallow. Serviceable rewritten passage. Nothing stands out.
15 openrouter:baidu/ernie-4.5-vl-424b-a47b view ↗ 8777567 6 0.43¢ 7.2 Chef-recipe and counterfeit analogies fine, but only 3 quiz questions and a 3-term glossary make it feel abbreviated. Q2 defines 'distil' as 'replicate' over 'steal' somewhat loosely. Passage accurate but compressed.
Scores 0–10 · 🟩 ≥8 · 🟨 4–7 · 🟥 <4 · sorted by Overall. Hover a column header for its full rubric name.

💰 Cost-adjusted ranking (value for money)

#ModelQualityCostValue
🥇 openrouter:deepseek/deepseek-v4-pro 9 0.62¢ 8.7
🥈 openrouter:minimax/minimax-m2.7 9 0.62¢ 8.7
🥉 openrouter:mistralai/mistral-large-2512 8 0.85¢ 8.0
4 openrouter:google/gemini-3.1-flash-lite 8 0.59¢ 8.0
5 openrouter:openai/gpt-5.4-nano 8 0.63¢ 8.0
6 anthropic:claude-haiku-4-5-20251001 9 2.57¢ 7.5
7 openrouter:x-ai/grok-4.3 8 1.64¢ 7.4
8 openrouter:meta-llama/llama-4-maverick 6 0.26¢ 7.2
9 openrouter:baidu/ernie-4.5-vl-424b-a47b 6 0.43¢ 7.2
10 openrouter:stepfun/step-3.7-flash 8 2.09¢ 6.8
11 openrouter:bytedance-seed/seed-2.0-lite 7 1.73¢ 6.7
12 openrouter:moonshotai/kimi-k2.6 9 9.62¢ 6.6
13 openrouter:z-ai/glm-5-turbo 7 3.40¢ 6.1
14 openrouter:qwen/qwen3-max-thinking 6 1.02¢ 6.0
15 openrouter:amazon/nova-pro-v1 6 1.34¢ 6.0

🪙 Best under 3¢ (cost < $0.03), ranked by quality

#ModelQualityCostValue
1 openrouter:deepseek/deepseek-v4-pro 9 0.62¢ 8.7
2 openrouter:minimax/minimax-m2.7 9 0.62¢ 8.7
3 anthropic:claude-haiku-4-5-20251001 9 2.57¢ 7.5
4 openrouter:mistralai/mistral-large-2512 8 0.85¢ 8.0
5 openrouter:google/gemini-3.1-flash-lite 8 0.59¢ 8.0
6 openrouter:openai/gpt-5.4-nano 8 0.63¢ 8.0
7 openrouter:x-ai/grok-4.3 8 1.64¢ 7.4
8 openrouter:stepfun/step-3.7-flash 8 2.09¢ 6.8
9 openrouter:bytedance-seed/seed-2.0-lite 7 1.73¢ 6.7
10 openrouter:meta-llama/llama-4-maverick 6 0.26¢ 7.2
11 openrouter:baidu/ernie-4.5-vl-424b-a47b 6 0.43¢ 7.2
12 openrouter:qwen/qwen3-max-thinking 6 1.02¢ 6.0
13 openrouter:amazon/nova-pro-v1 6 1.34¢ 6.0
Strong quality without paying flagship prices — the cheap-and-good picks.