GenAware Model Lab
LLM benchmarks on AMC 8 / AJHSME competition math: accuracy, speed, tokens and cost by difficulty tier.
Current study (3 Sep 2026)
Meta: Will Muse Cause a Spark?Muse Spark 1.3 vs Claude Opus 5, GPT-5.6 Terra Pro, Gemini 3.8 Flash, GLM and Qwen3.8 on AMC 8 math by difficulty. 31 rounds, 1,650 graded answers. Investor note, 38 charts, PDF/XLSX/CSV.
Live math runsEvery batch as it lands: per-model grid, cost, tokens, timing and full step-by-step working.
Earlier papers
Brute Force vs. Finesse (May 2026)16 models, 4,896 graded attempts. Chinese flagships vs Western budget and frontier models.
3-model paper (Jun 2026)GPT-5.5 vs Gemini Pro vs Llama-4 Maverick.