🔬 Model Lab

New run Stored runs ⚖️ Judge verdicts 🧮 Math 📊 Math runs 📄 Benchmark paper 📄 3-model paper 📄 Meta: Will Muse Cause a Spark?

⚖️ Judge verdict

India's Economic Paradox: Winning Votes, Losing Investors · Judge: anthropic:claude-opus-4-8 · 2026-09-06T13:48:14 · ✅ saved (711c81) · 2 models

🔬 View outputs side-by-side → all verdicts →

Note: the judge model is also one of the contestants — scores may carry self-preference bias.
🏆 Quality winner: anthropic:claude-opus-4-8
The Opus run is more tightly grounded in verifiable-sounding source facts (QCO 14→765, Feb 2025 BIT review, ~12% rupee fall) and its analogies share the actual mechanism of the topic. The Haiku run invents specifics (West Bengal win, exact 8th rank, 2015 BIT exit clauses) that read as hallucinations, and its quiz explanations cite passage text that isn't present, plus its rewritten_passage is essentially missing. Opus wins on fact_grounding, quiz fairness, and completeness despite its own awkwardly split passage fields.
💰 Best value: anthropic:claude-opus-4-8 — quality 8/10 at 17.06¢ → value 5.9/10
Value = 0.7·quality + 0.3·cheapness. Quality stays dominant, so a cheap-but-weak output can't win on price alone.

Quality ranking

#ModelFactsAnalogyAge-fitClarityQuizGlossaryPassageOverallCostValueNotes
🥇 anthropic:claude-opus-4-8 view ↗ 8998997 8 17.06¢ 5.9 Strong, faithful treatment: FDI, QCO surge (14→765), BIT review Feb 2025, rupee ~12% fall are all specific and consistent. The restaurant/kitchen and 'complain to chef's family before a health inspector' analogies genuinely match the mechanism (complacency and forced-local-courts). Quiz is fair and mostly answerable from the passage. The rewritten_passage is oddly split into a truncated main field plus an 'agents' continuation field, hurting reading-passage quality, but content is engaging and original.
🥈 anthropic:claude-haiku-4-5-20251001 view ↗ 4787582 5 3.02¢ 4.7 Introduces likely fabricated specifics not supported by the source: a 'West Bengal' electoral victory, a precise '8th globally' per-capita rank, and a '2015 BIT revision with exit clauses' framing that conflicts with the other candidate's Feb 2025 review detail. The rewritten_passage is truncated to a single fragment, and several quiz explanations quote 'passage' text that doesn't appear, undermining fairness. Analogies (student president, football team) are decent but more generic.
Scores 0–10 · 🟩 ≥8 · 🟨 4–7 · 🟥 <4 · sorted by Overall. Hover a column header for its full rubric name.

💰 Cost-adjusted ranking (value for money)

#ModelQualityCostValue
🥇 anthropic:claude-opus-4-8 8 17.06¢ 5.9
🥈 anthropic:claude-haiku-4-5-20251001 5 3.02¢ 4.7