๐Ÿ”ฌ Model Lab

New run Stored runs โš–๏ธ Judge verdicts ๐Ÿงฎ Math ๐Ÿ“Š Math runs ๐Ÿ“„ Benchmark paper ๐Ÿ“„ 3-model paper ๐Ÿ“„ Meta: Will Muse Cause a Spark?

โš–๏ธ Judge verdict

The AI Cold War's Newest Weapon: Stealing the Teacher's Answer Key ยท Judge: anthropic:claude-opus-4-8 ยท 2026-05-30T15:43:59 ยท โœ… saved (e5c57c) ยท 14 models

๐Ÿ”ฌ View outputs side-by-side โ†’ all verdicts โ†’

๐Ÿ† Quality winner: openrouter:moonshotai/kimi-k2.6
Kimi-k2.6 wins on the strength of its analogies, which uniquely match the actual mechanism (a conservatory allowing transcription until a rival sends ghost students at scale; a pharma firm whose published abstracts are scraped and reverse-engineered)โ€”both capture the 'legitimate-process-weaponized-at-scale' core better than the generic chess/doping analogies elsewhere. Its rewritten passage is complete, accurate, and incorporates source details like McGuire and the embassy quote, while its quiz features tight evidence-pairing and instructive trap explanations. DeepSeek (2eb734) and Minimax (c2b157) are nearly as strong with full ten-question quizzes and rich glossaries, but their analogies (Tu-4/B-29, answer key) are slightly less mechanism-precise. Several otherwise-strong entries (f95e70 with one quiz question, a0132a with three, plus several truncated passages) are penalized for incompleteness.

Quality ranking

#ModelFactsAnalogyAge-fitClarityQuizGlossaryPassageOverallCostValueNotes
๐Ÿฅ‡ openrouter:deepseek/deepseek-v4-pro view โ†— 10999999 9 โ€“ โ€“ Excellent fact fidelity with the Tu-4/B-29 Cold War analogy sharing genuine mechanism (reverse-engineering that doesn't fully match original). Ten well-constructed quiz questions with proper evidence-pairing. Rich glossary and a polished, original rewritten passage.
๐Ÿฅˆ openrouter:minimax/minimax-m2.7 view โ†— 9899999 9 โ€“ โ€“ Accurate and well-organized with strong 'splinternet' framing. Quiz questions tie tightly to the rewritten passage and evidence-pairing works. 'Surreptitious' question references a word not in the rewritten passage but the quotes are real.
๐Ÿฅ‰ openrouter:moonshotai/kimi-k2.6 view โ†— 9999999 9 โ€“ โ€“ Outstanding music-conservatory and pharma-abstract analogies that precisely match the mechanism. Faithful, complete rewritten passage incorporating McGuire and embassy details. Strong evidence-pairing quiz with clear trap explanations.
4 openrouter:mistralai/mistral-large-2512 view โ†— 9899889 8 โ€“ โ€“ Strong, faithful coverage with vivid analogies (answer key, marathon sweatband). The 'surreptitious' vocab question asks about a word not in the rewritten passage, a minor flaw. Otherwise clear and well-pitched.
5 openrouter:google/gemini-3.1-flash-lite view โ†— 9889889 8 โ€“ โ€“ Clean, faithful, with apt chef and capture-the-flag analogies. Good ten-question quiz with evidence pairing. Solid rewritten passage. Slightly less distinctive than the top finishers.
6 openrouter:z-ai/glm-5-turbo view โ†— 9899797 8 โ€“ โ€“ Vivid music-producer and scouting-report analogies, strong glossary. However the 'sharp' vocab question (Q4) tests a phrase the model fabricated as if from the source, and the rewritten passage is truncated. Otherwise excellent.
7 openrouter:x-ai/grok-4.3 view โ†— 9788888 8 โ€“ โ€“ Accurate and complete with reasonable reverse-engineering and scouting analogies. Good quiz with evidence pairing keyed to the rewritten passage. Slightly more compact and less inventive than top entries.
8 openrouter:openai/gpt-5.4-nano view โ†— 9888898 8 โ€“ โ€“ Careful, honest framing (consistently flags 'the article says'), strong export-controls/lock and scouting analogies. Good glossary and complete passage. Slightly hedged tone but very faithful to source.
9 openrouter:stepfun/step-3.7-flash view โ†— 8887887 7 โ€“ โ€“ Solid analogies and good SAT-style traps, but the rewritten passage is truncated mid-sentence, hurting completeness. Adds 'state-linked' framing not strictly in source. 'Surreptitious' question references content not in the truncated passage.
10 openrouter:qwen/qwen3-max-thinking view โ†— 9888589 7 โ€“ โ€“ Strong rewritten passage and good analogies, but only ONE quiz question is provided despite the format expecting a full set, a major shortfall. Otherwise accurate and well-pitched.
11 openrouter:meta-llama/llama-4-maverick view โ†— 8788877 7 โ€“ โ€“ Generic but functional analogies (arms race, doping). Quiz is solid. The rewritten passage adds 'urgent respon[se]' and is truncated; states the memo focuses on three companies somewhat overconfidently. Reasonable but not standout.
12 openrouter:bytedance-seed/seed-2.0-lite view โ†— 8788788 7 โ€“ โ€“ Good analogies (MIT exam keys, Twitch scouting). Q4 tests 'principal'โ€”a word that appears in source as 'principally,' a slight stretch. Otherwise accurate, complete passage and decent quiz with evidence pairing.
13 openrouter:amazon/nova-pro-v1 view โ†— 7566666 6 โ€“ โ€“ Competent but generic. The chess and doping analogies are weak/mechanism-mismatched. Quiz questions sometimes cite phrases ('global power dynamics') that exist in the explainer rather than the source. Thinner glossary and shorter passage.
14 openrouter:baidu/ernie-4.5-vl-424b-a47b view โ†— 7576566 5 โ€“ โ€“ Only three quiz questions and a thin three-term glossary. Q2 marks 'replicate' over the source's literal definition awkwardly. Analogies are generic. Shorter and less rigorous than peers.
Scores 0โ€“10 ยท ๐ŸŸฉ โ‰ฅ8 ยท ๐ŸŸจ 4โ€“7 ยท ๐ŸŸฅ <4 ยท sorted by Overall. Hover a column header for its full rubric name.