🔬 Model Lab

New run Stored runs ⚖️ Judge verdicts 🧮 Math 📊 Math runs 📄 Benchmark paper 📄 3-model paper 📄 Meta: Will Muse Cause a Spark?

⚖️ Judge verdict

India's Economic Paradox: Winning Votes, Losing Investors · Judge: anthropic:claude-opus-4-8 · 2026-09-24T14:38:18 · ✅ saved (f7e7ae) · 2 models

🔬 View outputs side-by-side → all verdicts →

Note: the judge model is also one of the contestants — scores may carry self-preference bias.
🏆 Quality winner: anthropic:claude-opus-4-8
The Opus run is more faithful (grounded specific figures, careful hedging like 'reform still awaited'), delivers mechanism-matched analogies, and provides a nearly complete engaging passage with well-paired evidence questions. The Haiku run fabricates specifics (West Bengal win, 8th-place rank, 2015 exit clauses) and its rewritten_passage is truncated to a single unfinished sentence, undermining a core deliverable.
💰 Best value: anthropic:claude-opus-4-8 — quality 9/10 at 17.06¢ → value 6.6/10
Value = 0.7·quality + 0.3·cheapness. Quality stays dominant, so a cheap-but-weak output can't win on price alone.

Quality ranking

#ModelFactsAnalogyAge-fitClarityQuizGlossaryPassageOverallCostValueNotes
🥇 anthropic:claude-opus-4-8 view ↗ 8999998 9 17.06¢ 6.6 Strong grounding with specific figures (12% rupee fall, QCOs 14→765, Feb 2025 BIT review) that align with a coherent source narrative. The five-star kitchen and restaurant/health-inspector analogies genuinely share the reform-complacency and legal-friction mechanisms. Quiz is fair and answerable; the split rewritten_passage/'agents' field is odd but content is engaging.
🥈 anthropic:claude-haiku-4-5-20251001 view ↗ 5888683 6 3.02¢ 5.4 Introduces several claims that appear invented or uncertain relative to the other candidate (a 'West Bengal' victory, per-capita rank '8th globally', a '2015 BIT revision with exit clauses')—these look fabricated with false precision. The rewritten_passage is truncated after one clause, badly hurting reading quality, and some quiz explanations quote lines not present in the passage.
Scores 0–10 · 🟩 ≥8 · 🟨 4–7 · 🟥 <4 · sorted by Overall. Hover a column header for its full rubric name.

💰 Cost-adjusted ranking (value for money)

#ModelQualityCostValue
🥇 anthropic:claude-opus-4-8 9 17.06¢ 6.6
🥈 anthropic:claude-haiku-4-5-20251001 6 3.02¢ 5.4