๐Ÿ”ฌ Model Lab

New run Stored runs โš–๏ธ Judge verdicts ๐Ÿงฎ Math ๐Ÿ“Š Math runs ๐Ÿ“„ Benchmark paper ๐Ÿ“„ 3-model paper ๐Ÿ“„ Meta: Will Muse Cause a Spark?
โš ๏ธ Runs call real paid APIs. All sampled questions are sent to each model in one call (so its "Total time" is the real round-trip). Web search is OFF โ€” pure reasoning, graded objectively against the known answer. Each model shows full step-by-step working for every question (in the drill-down). If a model can't fit all the questions and drops some answers, those score wrong โ€” a fair capability strike.
Leave all ticked to draw from every difficulty.
(1โ€“50; sampled at random โ€” every model gets the SAME set)
anthropic
openai
google
x-ai
meta-llama
deepseek
qwen
moonshotai
z-ai
minimax
baidu
bytedance-seed
stepfun
amazon
mistralai
meta
~deepseek
~z-ai
Selected (0): none yet โ€” tick models above
โ€” they'll be pre-ticked next time

A large matrix (> 60 model-questions) requires the box above. Runs launch in the background โ†’ results page (auto-refreshes as answers land).

Past sessions