🔬 Model Lab

New run Stored runs ⚖️ Judge verdicts 🧮 Math 📊 Math runs 📄 Benchmark paper 📄 3-model paper 📄 Meta: Will Muse Cause a Spark?

⚖️ Judge verdict

India's Economic Paradox: Winning Votes, Losing Investors · Judge: anthropic:claude-opus-4-8 · 2026-09-05T00:16:55 · ✅ saved (571a8c) · 2 models

🔬 View outputs side-by-side → all verdicts →

Note: the judge model is also one of the contestants — scores may carry self-preference bias.
🏆 Quality winner: anthropic:claude-opus-4-8
Run 83ff29 sticks tightly to verifiable source facts and dates, while 705590 introduces likely-fabricated specifics (West Bengal win, '8th globally' per-capita rank, precise IMF role) that risk misleading students. 83ff29's analogies capture the actual mechanism—political dominance eroding reform pressure—rather than generic 'different muscles' framing, and its quiz items are answerable from its passage, whereas several of 705590's questions cite text absent from its severely truncated passage.
💰 Best value: anthropic:claude-opus-4-8 — quality 9/10 at 17.06¢ → value 6.6/10
Value = 0.7·quality + 0.3·cheapness. Quality stays dominant, so a cheap-but-weak output can't win on price alone.

Quality ranking

#ModelFactsAnalogyAge-fitClarityQuizGlossaryPassageOverallCostValueNotes
🥇 anthropic:claude-opus-4-8 view ↗ 9999898 9 17.06¢ 6.6 Strong grounding in specific source details (QCOs 14→765, 12% rupee fall, BIT review Feb 2025). The restaurant/kitchen and health-inspector analogies share the actual mechanism (dominance removes accountability pressure). Passage is engaging though truncated mid-sentence with an odd 'agents' field. Quiz is fair and answerable, glossary precise.
🥈 anthropic:claude-haiku-4-5-20251001 view ↗ 5888583 6 3.02¢ 5.4 Invents several facts not clearly in source: a 'West Bengal' electoral victory, per-capita rank '8th globally', 2015 BIT '2015 revision' framing, and specific IMF appointment details—these look fabricated or embellished. Quiz questions reference text (e.g. 'seeking long-term certainty') not present in the truncated rewritten_passage, so some MCQs aren't answerable from the passage. Analogies are decent but multiple and generic; reading passage is cut off almost immediately.
Scores 0–10 · 🟩 ≥8 · 🟨 4–7 · 🟥 <4 · sorted by Overall. Hover a column header for its full rubric name.

💰 Cost-adjusted ranking (value for money)

#ModelQualityCostValue
🥇 anthropic:claude-opus-4-8 9 17.06¢ 6.6
🥈 anthropic:claude-haiku-4-5-20251001 6 3.02¢ 5.4