GenAware Model Lab

LLM benchmarks on AMC 8 / AJHSME competition math: accuracy, speed, tokens and cost by difficulty tier.

Current study (3 Sep 2026)

Meta: Will Muse Cause a Spark?Muse Spark 1.3 vs Claude Opus 5, GPT-5.6 Terra Pro, Gemini 3.8 Flash, GLM and Qwen3.8 on AMC 8 math by difficulty. 31 rounds, 1,650 graded answers. Investor note, 38 charts, PDF/XLSX/CSV. Live math runsEvery batch as it lands: per-model grid, cost, tokens, timing and full step-by-step working.

Earlier papers

Brute Force vs. Finesse (May 2026)16 models, 4,896 graded attempts. Chinese flagships vs Western budget and frontier models. 3-model paper (Jun 2026)GPT-5.5 vs Gemini Pro vs Llama-4 Maverick.

Lab

Run a math benchmarkPick models, tier and batch size. Calls paid APIs. Model Lab homeArticle-analysis runs, judge verdicts, stored runs.