🔬 Model Lab

New run Stored runs ⚖️ Judge verdicts 🧮 Math 📊 Math runs 📄 Benchmark paper 📄 3-model paper 📄 Meta: Will Muse Cause a Spark?

⚖️ Judge verdict

India's Economic Paradox: Winning Votes, Losing Investors · Judge: anthropic:claude-opus-4-8 · 2026-09-25T02:48:22 · ✅ saved (f0c6d7) · 2 models

🔬 View outputs side-by-side → all verdicts →

Note: the judge model is also one of the contestants — scores may carry self-preference bias.
🏆 Quality winner: anthropic:claude-opus-4-8
Run 83ff29 stays tighter to source facts (correctly identifies Bhalla as ex-IMF/PM adviser, gives verifiable numbers, dates the BIT review to Feb 2025 rather than inventing a 2015 revision), and its quiz items are answerable and fair. Run 705590 adds unverifiable specifics (West Bengal win, per-capita '8th', 2015 BIT clause) and its rewritten_passage is essentially empty, making it weaker on grounding and the core SAT-passage deliverable despite good analogies.
💰 Best value: anthropic:claude-opus-4-8 — quality 9/10 at 17.06¢ → value 6.6/10
Value = 0.7·quality + 0.3·cheapness. Quality stays dominant, so a cheap-but-weak output can't win on price alone.

Quality ranking

#ModelFactsAnalogyAge-fitClarityQuizGlossaryPassageOverallCostValueNotes
🥇 anthropic:claude-opus-4-8 view ↗ 9999998 9 17.06¢ 6.6 Strong fidelity to the source with specific figures (12% rupee fall, QCOs 14→765, Feb 2025 BIT review). Analogies (restaurant kitchen, health inspector delay) share the same mechanism as complacency and arbitration barriers. Rewritten passage is truncated/split awkwardly into 'agents' field but overall prose is engaging; quiz is fair and grounded.
🥈 anthropic:claude-haiku-4-5-20251001 view ↗ 5888683 6 3.02¢ 5.4 Introduces several likely-invented specifics not clearly in source: a West Bengal electoral win, 'per capita ranks 8th', a '2015 BIT revision', and QCOs framed narrowly at foreign firms—these raise fact-grounding concerns. Analogies are vivid and mechanism-matching, but several quiz answer explanations cite quotes ('the deeper danger...') that don't cleanly match a coherent passage, and the rewritten_passage is cut off after one line.
Scores 0–10 · 🟩 ≥8 · 🟨 4–7 · 🟥 <4 · sorted by Overall. Hover a column header for its full rubric name.

💰 Cost-adjusted ranking (value for money)

#ModelQualityCostValue
🥇 anthropic:claude-opus-4-8 9 17.06¢ 6.6
🥈 anthropic:claude-haiku-4-5-20251001 6 3.02¢ 5.4