🔬 Model Lab

New run Stored runs ⚖️ Judge verdicts 🧮 Math 📊 Math runs 📄 Benchmark paper 📄 3-model paper 📄 Meta: Will Muse Cause a Spark?

⚖️ Judge verdict

India's Economic Paradox: Winning Votes, Losing Investors · Judge: anthropic:claude-opus-4-8 · 2026-09-04T06:34:55 · ✅ saved (d6dc1e) · 2 models

🔬 View outputs side-by-side → all verdicts →

Note: the judge model is also one of the contestants — scores may carry self-preference bias.
🏆 Quality winner: anthropic:claude-opus-4-8
The Opus run stays tightly grounded in verifiable article details and its quiz answers are all supported by its own passage/explanation, whereas the Haiku run invents specifics (West Bengal, 8th rank, 2015 exit clauses) and pairs quiz questions with quoted evidence that never appears in its truncated passage, breaking the answerable-from-passage requirement. Both passages are truncated, but Opus's overall fact fidelity, analogy precision, and internally consistent quiz make it the clear winner.
💰 Best value: anthropic:claude-opus-4-8 — quality 8.6/10 at 17.06¢ → value 6.3/10
Value = 0.7·quality + 0.3·cheapness. Quality stays dominant, so a cheap-but-weak output can't win on price alone.

Quality ranking

#ModelFactsAnalogyAge-fitClarityQuizGlossaryPassageOverallCostValueNotes
🥇 anthropic:claude-opus-4-8 view ↗ 9999996 8.6 17.06¢ 6.3 Facts are specific and consistent (12% rupee fall, QCOs 14→765, Sitharaman Feb 2025, BIT courts-before-arbitration). Restaurant/kitchen and 'call the health inspector' analogies genuinely share the mechanism. Quiz is fair and SAT-styled. The rewritten_passage is truncated mid-sentence, and the odd 'agents' field looks like a passage fragment escaped its container, hurting passage quality.
🥈 anthropic:claude-haiku-4-5-20251001 view ↗ 5898683 6.1 3.02¢ 5.5 Introduces several likely-invented specifics not clearly in source: 'West Bengal victory,' 'per capita rank 8th,' '2015 BIT exit clauses,' Bhalla's exact IMF/PMEAC roles — these read as fabricated or over-precise. Analogies (student president, restaurant, football) are good. But several quiz items quote 'evidence' passages that don't appear in the truncated rewritten_passage, making them unanswerable, and the passage is cut off almost immediately.
Scores 0–10 · 🟩 ≥8 · 🟨 4–7 · 🟥 <4 · sorted by Overall. Hover a column header for its full rubric name.

💰 Cost-adjusted ranking (value for money)

#ModelQualityCostValue
🥇 anthropic:claude-opus-4-8 8.6 17.06¢ 6.3
🥈 anthropic:claude-haiku-4-5-20251001 6.1 3.02¢ 5.5