🔬 Model Lab

New run Stored runs ⚖️ Judge verdicts 🧮 Math 📊 Math runs 📄 Benchmark paper 📄 3-model paper 📄 Meta: Will Muse Cause a Spark?

⚖️ Judge verdict

India's Economic Paradox: Winning Votes, Losing Investors · Judge: anthropic:claude-opus-4-8 · 2026-09-11T10:47:35 · ✅ saved (614b53) · 2 models

🔬 View outputs side-by-side → all verdicts →

Note: the judge model is also one of the contestants — scores may carry self-preference bias.
🏆 Quality winner: anthropic:claude-opus-4-8
Run 83ff29 is far more faithful to the source—it hedges attribution to Bhalla, uses verifiable figures, and its analogies map tightly onto the reform-complacency mechanism. Run 705590 invents concrete facts (West Bengal win, 2015 exit clauses, 8th-place ranking, quiz explanations citing nonexistent lines) and its reading passage is barely more than a fragment, whereas 83ff29's is fuller and more engaging despite its own truncation.
💰 Best value: anthropic:claude-opus-4-8 — quality 8/10 at 17.06¢ → value 5.9/10
Value = 0.7·quality + 0.3·cheapness. Quality stays dominant, so a cheap-but-weak output can't win on price alone.

Quality ranking

#ModelFactsAnalogyAge-fitClarityQuizGlossaryPassageOverallCostValueNotes
🥇 anthropic:claude-opus-4-8 view ↗ 8899897 8 17.06¢ 5.9 Strong grounding with specific figures (12% rupee fall, QCOs 14→765, Feb 2025 BIT review) that align well and are used consistently. The restaurant/kitchen and 'complain to the chef's own family' analogies genuinely share the topic's mechanism. Reading passage is engaging but appears cut off mid-sentence ('inflation is contained') with the 'agents' field oddly split, hurting passage coherence.
🥈 anthropic:claude-haiku-4-5-20251001 view ↗ 4788583 5 3.02¢ 4.7 Introduces several likely-invented specifics: a 'West Bengal victory,' a '2015 BIT revision with exit clauses,' and a precise 'per capita GDP ranks 8th' figure that don't clearly derive from the other candidate's shared facts and read as fabricated. Analogies (student president, football team) are decent. The rewritten_passage is truncated to a single fragment, and several quiz explanations quote passage text that doesn't appear, undermining fairness.
Scores 0–10 · 🟩 ≥8 · 🟨 4–7 · 🟥 <4 · sorted by Overall. Hover a column header for its full rubric name.

💰 Cost-adjusted ranking (value for money)

#ModelQualityCostValue
🥇 anthropic:claude-opus-4-8 8 17.06¢ 5.9
🥈 anthropic:claude-haiku-4-5-20251001 5 3.02¢ 4.7