๐Ÿ”ฌ Model Lab

New run Stored runs โš–๏ธ Judge verdicts ๐Ÿงฎ Math ๐Ÿ“Š Math runs ๐Ÿ“„ Benchmark paper ๐Ÿ“„ 3-model paper ๐Ÿ“„ Meta: Will Muse Cause a Spark?

โš–๏ธ Judge verdict

The AI Cold War's Newest Weapon: Stealing the Teacher's Answer Key ยท Judge: claude-code (Opus 4.8, subscription) ยท 2026-05-30T14:07:24 ยท โœ… saved (86f9ad) ยท 15 models

๐Ÿ”ฌ View outputs side-by-side โ†’ all verdicts โ†’

๐Ÿ† Quality winner: openrouter:deepseek/deepseek-v4-pro
DeepSeek-V4-Pro wins on the strength of the single best analogy in the field โ€” the Soviet Tu-4 bomber reverse-engineered from captured B-29s: a replica that closed a strategic gap yet never matched the original. That maps the distillation mechanism AND its 'good-enough, far-cheaper' payoff more precisely than any rival, and it pairs the analogy with a rich, accurate explanation, a full 10-question quiz with trap analysis and varied SAT tips, a strong glossary, and an engaging, faithful passage. GLM-5-Turbo is a very close second (the most thorough and faithful passage, capturing even Chris McGuire's specific recommendations, plus the most rigorous quiz with a consistent Trap A/B/C taxonomy) and Claude Haiku 4.5 third (sharp doping analogy, very polished, complete). The field's clear failures: Qwen3-Max-Thinking produced only 1 of 10 quiz questions (it spent its budget on reasoning), while Nova-Pro and Ernie-4.5-VL were thin and, in Ernie's case, returned only 3 quiz questions and 3 glossary terms.

Quality ranking

#ModelFactsAnalogyAge-fitClarityQuizGlossaryPassageOverallCostValueNotes
๐Ÿฅ‡ openrouter:deepseek/deepseek-v4-pro view โ†— 91099999 9 โ€“ โ€“ Best all-round. The B-29/Tu-4 bomber analogy is the standout of the field โ€” a reverse-engineered replica that closed a gap yet never matched the original, mapping distillation precisely. Rich accurate explanation, full 10-question quiz with trap analysis + varied SAT tips, strong 6-term glossary, engaging faithful passage.
๐Ÿฅˆ openrouter:z-ai/glm-5-turbo view โ†— 1089910810 9 โ€“ โ€“ Most thorough and faithful output: a six-paragraph passage that captures nuance and even Chris McGuire's specific recommendations. The quiz is the most rigorous, with a consistent Trap A/B/C taxonomy and an SAT tip on every item. Analogies (vocal-stripping, scouting reports) are solid if a touch less vivid than the very best.
๐Ÿฅ‰ anthropic:claude-haiku-4-5-20251001 view โ†— 9999999 9 โ€“ โ€“ Polished and accurate. The doping analogy โ€” 'distillation isn't cheating, but using someone else's distilled model is' โ€” is sharp and on-mechanism. Strong 6-term glossary, rich quiz with trap analysis, well-structured four-paragraph passage that includes McGuire. A balanced top-tier entry.
4 openrouter:moonshotai/kimi-k2.6 view โ†— 9998988 8 โ€“ โ€“ Sophisticated, on-mechanism analogies (music-conservatory transcription, pharma-abstract scraping). Full quiz with trap analysis and a polished, faithful passage. Minor blemish: 'Chris McGuire' is named twice in consecutive sentences โ€” a small editing lapse. Just below the very top.
5 openrouter:stepfun/step-3.7-flash view โ†— 9988988 8 โ€“ โ€“ Concrete, same-mechanism analogies (valedictorian's answers; reverse-engineering a competitor's burger) and a full 10-question quiz with explicit trap labels. Long, detailed, faithful passage. Reads slightly long but is high quality throughout.
6 openrouter:mistralai/mistral-large-2512 view โ†— 9898988 8 โ€“ โ€“ Strong 'answer key' hook and a full quiz with trap analysis. Good supply-chain-heist and marathon-sweatband analogies, though 'save-scumming' is a looser fit. Faithful, well-paragraphed passage. Solid top-half entry.
7 openrouter:minimax/minimax-m2.7 view โ†— 9898888 8 โ€“ โ€“ The 'F1 car cloned from race footage, sold at Toyota prices' analogy is vivid and on-mechanism. Full quiz, good 6-term glossary, accurate five-paragraph passage. A well-rounded, faithful entry.
8 openrouter:x-ai/grok-4.3 view โ†— 8788788 7 โ€“ โ€“ Clean and accurate, with a tidy five-paragraph passage. Analogies (reverse-engineer a product, scout a playbook with hidden cameras) are competent but familiar. Quiz is complete but its explanations are thinner โ€” brief trap notes, few SAT tips โ€” than the leaders.
9 openrouter:google/gemini-3.1-flash-lite view โ†— 8788877 7 โ€“ โ€“ Solid and accurate but lighter: shorter explanation sections and a four-paragraph passage with less depth. Quiz is complete and fair with SAT tips. Analogies (master-chef recipe, capture-the-flag lock-picking) are serviceable. A dependable mid-tier result.
10 openrouter:bytedance-seed/seed-2.0-lite view โ†— 8886887 7 โ€“ โ€“ Vivid analogies (MIT exam-key theft, Twitch strategy-scraping) and a complete quiz with trap labels. Main weakness: the rewritten passage is a single unbroken wall of text with no paragraph breaks, which hurts scannability for the target reader.
11 openrouter:openai/gpt-5.4-nano view โ†— 8787886 7 โ€“ โ€“ Competent quiz with trap analysis and a clear glossary. The defining flaw is voice: the explanation and the 'rewritten' passage repeatedly say 'the article explains/describes/framesโ€ฆ', narrating the source instead of producing original SAT-style prose โ€” which undercuts the passage's whole purpose and immersion.
12 openrouter:meta-llama/llama-4-maverick view โ†— 8676776 6 โ€“ โ€“ Accurate but flat. The hook merely restates the headline rather than sparking curiosity, the analogy is the generic 'copying homework,' and the rewritten passage is one undifferentiated block. Quiz is complete but basic. Functional, not engaging.
13 openrouter:qwen/qwen3-max-thinking view โ†— 8886188 6 โ€“ โ€“ Strong prose, a good counterfeit-handbag-from-photos analogy, and a faithful passage โ€” but it produced only ONE quiz question instead of ten (the model likely exhausted its token budget on reasoning). The missing quiz is a critical deliverable failure that drags an otherwise good entry down.
14 openrouter:amazon/nova-pro-v1 view โ†— 7566755 5 โ€“ โ€“ Thin throughout: a one-line hook, brief sections, only three glossary terms, and a short single-paragraph passage. Analogies (chess, copying homework, performance-enhancing drugs) are generic and not mechanism-matched. Quiz is complete but basic. Among the weakest.
15 openrouter:baidu/ernie-4.5-vl-424b-a47b view โ†— 7666356 5 โ€“ โ€“ Incomplete: only three quiz questions (of ten) and three glossary terms, with a short single-block passage. Analogies are listed but generic (secret recipe, counterfeit goods, generic drug). Accurate as far as it goes, but well short of the required deliverables.
Scores 0โ€“10 ยท ๐ŸŸฉ โ‰ฅ8 ยท ๐ŸŸจ 4โ€“7 ยท ๐ŸŸฅ <4 ยท sorted by Overall. Hover a column header for its full rubric name.