The AI Cold War's Newest Weapon: Stealing the Teacher's Answer Key ยท Judge: claude-code (Opus 4.8, subscription) ยท 2026-05-30T14:07:24
ยท โ
saved (86f9ad) ยท 15 models
๐ฌ View outputs side-by-side โ all verdicts โ
| # | Model | Facts | Analogy | Age-fit | Clarity | Quiz | Glossary | Passage | Overall | Cost | Value | Notes |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ๐ฅ | openrouter:deepseek/deepseek-v4-pro
view โ |
9 | 10 | 9 | 9 | 9 | 9 | 9 | 9 | โ | โ | Best all-round. The B-29/Tu-4 bomber analogy is the standout of the field โ a reverse-engineered replica that closed a gap yet never matched the original, mapping distillation precisely. Rich accurate explanation, full 10-question quiz with trap analysis + varied SAT tips, strong 6-term glossary, engaging faithful passage. |
| ๐ฅ | openrouter:z-ai/glm-5-turbo
view โ |
10 | 8 | 9 | 9 | 10 | 8 | 10 | 9 | โ | โ | Most thorough and faithful output: a six-paragraph passage that captures nuance and even Chris McGuire's specific recommendations. The quiz is the most rigorous, with a consistent Trap A/B/C taxonomy and an SAT tip on every item. Analogies (vocal-stripping, scouting reports) are solid if a touch less vivid than the very best. |
| ๐ฅ | anthropic:claude-haiku-4-5-20251001
view โ |
9 | 9 | 9 | 9 | 9 | 9 | 9 | 9 | โ | โ | Polished and accurate. The doping analogy โ 'distillation isn't cheating, but using someone else's distilled model is' โ is sharp and on-mechanism. Strong 6-term glossary, rich quiz with trap analysis, well-structured four-paragraph passage that includes McGuire. A balanced top-tier entry. |
| 4 | openrouter:moonshotai/kimi-k2.6
view โ |
9 | 9 | 9 | 8 | 9 | 8 | 8 | 8 | โ | โ | Sophisticated, on-mechanism analogies (music-conservatory transcription, pharma-abstract scraping). Full quiz with trap analysis and a polished, faithful passage. Minor blemish: 'Chris McGuire' is named twice in consecutive sentences โ a small editing lapse. Just below the very top. |
| 5 | openrouter:stepfun/step-3.7-flash
view โ |
9 | 9 | 8 | 8 | 9 | 8 | 8 | 8 | โ | โ | Concrete, same-mechanism analogies (valedictorian's answers; reverse-engineering a competitor's burger) and a full 10-question quiz with explicit trap labels. Long, detailed, faithful passage. Reads slightly long but is high quality throughout. |
| 6 | openrouter:mistralai/mistral-large-2512
view โ |
9 | 8 | 9 | 8 | 9 | 8 | 8 | 8 | โ | โ | Strong 'answer key' hook and a full quiz with trap analysis. Good supply-chain-heist and marathon-sweatband analogies, though 'save-scumming' is a looser fit. Faithful, well-paragraphed passage. Solid top-half entry. |
| 7 | openrouter:minimax/minimax-m2.7
view โ |
9 | 8 | 9 | 8 | 8 | 8 | 8 | 8 | โ | โ | The 'F1 car cloned from race footage, sold at Toyota prices' analogy is vivid and on-mechanism. Full quiz, good 6-term glossary, accurate five-paragraph passage. A well-rounded, faithful entry. |
| 8 | openrouter:x-ai/grok-4.3
view โ |
8 | 7 | 8 | 8 | 7 | 8 | 8 | 7 | โ | โ | Clean and accurate, with a tidy five-paragraph passage. Analogies (reverse-engineer a product, scout a playbook with hidden cameras) are competent but familiar. Quiz is complete but its explanations are thinner โ brief trap notes, few SAT tips โ than the leaders. |
| 9 | openrouter:google/gemini-3.1-flash-lite
view โ |
8 | 7 | 8 | 8 | 8 | 7 | 7 | 7 | โ | โ | Solid and accurate but lighter: shorter explanation sections and a four-paragraph passage with less depth. Quiz is complete and fair with SAT tips. Analogies (master-chef recipe, capture-the-flag lock-picking) are serviceable. A dependable mid-tier result. |
| 10 | openrouter:bytedance-seed/seed-2.0-lite
view โ |
8 | 8 | 8 | 6 | 8 | 8 | 7 | 7 | โ | โ | Vivid analogies (MIT exam-key theft, Twitch strategy-scraping) and a complete quiz with trap labels. Main weakness: the rewritten passage is a single unbroken wall of text with no paragraph breaks, which hurts scannability for the target reader. |
| 11 | openrouter:openai/gpt-5.4-nano
view โ |
8 | 7 | 8 | 7 | 8 | 8 | 6 | 7 | โ | โ | Competent quiz with trap analysis and a clear glossary. The defining flaw is voice: the explanation and the 'rewritten' passage repeatedly say 'the article explains/describes/framesโฆ', narrating the source instead of producing original SAT-style prose โ which undercuts the passage's whole purpose and immersion. |
| 12 | openrouter:meta-llama/llama-4-maverick
view โ |
8 | 6 | 7 | 6 | 7 | 7 | 6 | 6 | โ | โ | Accurate but flat. The hook merely restates the headline rather than sparking curiosity, the analogy is the generic 'copying homework,' and the rewritten passage is one undifferentiated block. Quiz is complete but basic. Functional, not engaging. |
| 13 | openrouter:qwen/qwen3-max-thinking
view โ |
8 | 8 | 8 | 6 | 1 | 8 | 8 | 6 | โ | โ | Strong prose, a good counterfeit-handbag-from-photos analogy, and a faithful passage โ but it produced only ONE quiz question instead of ten (the model likely exhausted its token budget on reasoning). The missing quiz is a critical deliverable failure that drags an otherwise good entry down. |
| 14 | openrouter:amazon/nova-pro-v1
view โ |
7 | 5 | 6 | 6 | 7 | 5 | 5 | 5 | โ | โ | Thin throughout: a one-line hook, brief sections, only three glossary terms, and a short single-paragraph passage. Analogies (chess, copying homework, performance-enhancing drugs) are generic and not mechanism-matched. Quiz is complete but basic. Among the weakest. |
| 15 | openrouter:baidu/ernie-4.5-vl-424b-a47b
view โ |
7 | 6 | 6 | 6 | 3 | 5 | 6 | 5 | โ | โ | Incomplete: only three quiz questions (of ten) and three glossary terms, with a short single-block passage. Analogies are listed but generic (secret recipe, counterfeit goods, generic drug). Accurate as far as it goes, but well short of the required deliverables. |