๐Ÿ”ฌ Model Lab

New run Stored runs โš–๏ธ Judge verdicts ๐Ÿงฎ Math ๐Ÿ“Š Math runs ๐Ÿ“„ Benchmark paper ๐Ÿ“„ 3-model paper ๐Ÿ“„ Meta: Will Muse Cause a Spark?

โš–๏ธ Judge verdict

The AI Cold War's Newest Weapon: Stealing the Teacher's Answer Key ยท Judge: anthropic:claude-opus-4-8 ยท 2026-05-30T13:50:31 ยท โœ… saved (919ca9) ยท 3 models

๐Ÿ”ฌ View outputs side-by-side โ†’ all verdicts โ†’

๐Ÿ† Quality winner: openrouter:deepseek/deepseek-v4-pro
DeepSeek-v4 edges out Haiku with the most mechanically precise analogy (Tu-4 reverse-engineering that explicitly mirrors 'cheaper copy that never fully matches original'), the richest accurate detail (dates, named firms, cause-effect of export controls), and a longer, more vivid rewritten passage. Both it and Haiku have strong 10-question quizzes, but DeepSeek's vocab items are better anchored in the passage and its tone/inference questions are slightly fairer. Ernie's run is good but offers only 3 questions and a thinner treatment.

Quality ranking

#ModelFactsAnalogyAge-fitClarityQuizGlossaryPassageOverallCostValueNotes
๐Ÿฅ‡ openrouter:deepseek/deepseek-v4-pro view โ†— 9999999 9 โ€“ โ€“ Excellent depth and fidelity; the Tu-4/B-29 Cold War analogy genuinely shares the reverse-engineering-falls-short mechanism. Ten varied, fair SAT-style questions with evidence-pairing, rich glossary, and a multi-paragraph engaging passage. Strongest across every dimension.
๐Ÿฅˆ anthropic:claude-haiku-4-5-20251001 view โ†— 8899899 8 โ€“ โ€“ Strong, engaging, ten well-built questions with evidence-pairing and good distractors. One vocab item ('surreptitious') is harder to ground from the rewritten passage alone, and the doping/answer-key analogies are slightly generic. Excellent glossary and rewritten passage.
๐Ÿฅ‰ openrouter:baidu/ernie-4.5-vl-424b-a47b view โ†— 8788688 7 โ€“ โ€“ Solid and accurate with good chef-recipe and academic-plagiarism analogies. Only 3 quiz questions and one ('distil' means replicate vs steal) is debatable. Glossary is concise and useful, passage well-written but compact.
Scores 0โ€“10 ยท ๐ŸŸฉ โ‰ฅ8 ยท ๐ŸŸจ 4โ€“7 ยท ๐ŸŸฅ <4 ยท sorted by Overall. Hover a column header for its full rubric name.