๐Ÿ”ฌ Model Lab

New run Stored runs โš–๏ธ Judge verdicts ๐Ÿงฎ Math ๐Ÿ“Š Math runs ๐Ÿ“„ Benchmark paper ๐Ÿ“„ 3-model paper ๐Ÿ“„ Meta: Will Muse Cause a Spark?
โš ๏ธ Runs call real paid APIs (not the Claude subscription). Estimated cost for the whole matrix is shown before you run. Anthropic runs use the production web_search tool; other providers don't โ€” cross-provider runs aren't apples-to-apples on fact-checking.
Ctrl/Cmd-click to pick several (110 runnable). The raw original input is fed to the model (not the processed text): ๐Ÿ–ผ = original image โ†’ runs on vision models ยท ๐Ÿ“„ PDF (text extracted) ยท ๐Ÿ“ text. Hover an article for its summary.
๐Ÿ“„ โ€ฆor upload a new PDF to test on a fresh article:
Text is extracted from the PDF and run through the selected models (same as a ๐Ÿ“„ library article).
anthropic
openai
google
x-ai
meta-llama
deepseek
qwen
moonshotai
z-ai
minimax
baidu
bytedance-seed
stepfun
amazon
mistralai
meta
~deepseek
~z-ai
Selected (0): none yet โ€” tick models above; remove with ร—
๐Ÿ‘ = vision-capable โ€” REQUIRED for ๐Ÿ–ผ image articles (text-only models 404 on images). out = output-token price (usually the dominant cost). ๐ŸŒ web search is on for every model (Anthropic native; others via OpenRouter plugin, ~$0.005/run) โ€” apples-to-apples. Only each lab's latest generation is listed. refresh
Leave as-is to use the original (production) prompt, or pick a saved variant / edit it to A/B-test.
Picking one fills the box below; edit freely afterwards.
Saving makes it selectable here next time. original is always available and is never overwritten.

Runs launch in the background and you go straight to the status page โ€” no waiting on a spinner. A large matrix (> 12 cells) requires the checkbox above.