๐Ÿ”ฌ Model Lab

New run Stored runs โš–๏ธ Judge verdicts ๐Ÿงฎ Math ๐Ÿ“Š Math runs ๐Ÿ“„ Benchmark paper ๐Ÿ“„ 3-model paper ๐Ÿ“„ Meta: Will Muse Cause a Spark?

โš–๏ธ Judge verdict

Judge: anthropic:claude-opus-4-8 ยท 2026-05-30T11:08:49 ยท โœ… saved (2c4318) ยท 2 models

๐Ÿ”ฌ View outputs side-by-side โ†’ all verdicts โ†’

๐Ÿ† Quality winner: 20260530-110537-83ff29
Run 1 is more faithful to the source's specific data and Bhalla's framing (four agents, band-aids vs surgery, exhaust-Indian-courts arbitration), with tighter, SAT-grade quiz items and a precise glossary. Run 2 invents details (West Bengal win, 8th-place per-capita rank, a 2015 BIT exit-clause revision) and anchors quiz answers to text not present, undermining fact_grounding and quiz fairness. Both passages are truncated, but Run 1's substantive content and accuracy clearly outweigh Run 2.

Quality ranking

#ModelFactsAnalogyAge-fitClarityQuizGlossaryPassageOverallCostValueNotes
๐Ÿฅ‡ view โ†— 8899997 8 โ€“ โ€“ Strong, faithful treatment with specific figures (12% rupee fall, QCOs 14โ†’765, Feb 2025 BIT review) that track Bhalla's argument carefully. Analogies (restaurant/health inspector for arbitration) share genuine mechanism. The rewritten_passage appears truncated mid-sentence ('inflation is contained') with the rest spilling into a separate 'agents' field, hurting passage quality, but quiz and glossary are excellent and SAT-calibrated.
๐Ÿฅˆ view โ†— 5788683 6 โ€“ โ€“ Engaging analogies (student president, restaurant, football) but introduces likely-invented specifics: 'West Bengal victory,' per-capita rank '8th,' a '2015 BIT revision with exit clauses' โ€” these conflict with run 1's grounding and read as fabricated. Several quiz answers cite 'the passage states' for facts not clearly in the source, and the rewritten_passage is severely truncated ('presents a political and economic').
Scores 0โ€“10 ยท ๐ŸŸฉ โ‰ฅ8 ยท ๐ŸŸจ 4โ€“7 ยท ๐ŸŸฅ <4 ยท sorted by Overall. Hover a column header for its full rubric name.