GenAware Model Lab
Model research · First look
03 Sep 2026 · Meta Platforms (META)

Meta: Will Muse Cause a Spark?

Muse Spark 1.3 puts Meta back on the frontier watchlist. It is not yet a revenue story.

Constructive on capability. Neutral on monetisation.

Four months ago we put Meta's Llama-4 in the dead lane of our model map: slow and unreliable when problems got hard, with no reason to deploy it over cheaper or better alternatives. Muse Spark 1.3, released this month, changes that verdict. Across 31 rounds and 1650 graded answers on AMC 8 competition math, Muse scored 99.0% on the questions it answered, level with Claude Opus 5 and GPT-5.6 Terra Pro and inside a point of Gemini 3.8 Flash, and it was perfect on the two hardest tiers. It was the second-fastest model in the field and its metered bill, 0.27 cents a question, was a third of Opus 5's and a fifth of Terra Pro's. What it has not yet shown is consistency: its time per batch swung from 35 to 170 seconds on the hardest problems and one hard batch blew a three-minute deadline, which is exactly the trait that keeps a model out of the interactive products where the margin sits. We read Muse as proof that Meta's AI capex is producing a frontier-grade model and as cost-of-goods relief for Meta's own surfaces, not as a new revenue line. That is a start, and it is worth taking seriously.

1

Meta now builds frontier-grade models.

In May, on the same problem pool, Llama-4 Maverick scored 84.2% overall and 65.5% on the hardest tier. Muse scores 99.0% and 100.0%. On the like-for-like view the entire field now sits within about three points and Muse is inside that band at every difficulty tier, with every hard and stretch question answered correctly. The "Meta cannot build at the frontier" pillar of the bear case does not survive this data.

2

The gap to Anthropic is efficiency, not intelligence.

Think of a smart school math kid and an olympiad winner. Both get the answer. The olympiad winner has muscle memory: Opus 5 used 326 output tokens a question to Muse's 612, took 37 seconds a batch to Muse's 52, barely slowed as problems got harder and never ran past about a minute. Muse works everything from first principles, with 66% of its output spent thinking, and its time nearly triples from easy to stretch. Speed and consistency on hard work remain Anthropic's moat.

3

This is a cost story for Meta, not a revenue story.

A model that is cheap, accurate and variable in latency is the profile you run inside your own free products at scale, not one you sell by the token to enterprises paying for predictability. Muse lowers what Meta AI, WhatsApp and Instagram cost to serve and removes Meta's dependence on rivals' models. It does not, on this evidence, earn Meta a seat in the API market. The next release, and whether the latency tail narrows, decides which of those two stories dominates.

⬇ PDF⬇ Data workbook (XLSX)⬇ Raw CSVFull data edition (38 charts, all tables)Live session pages

01Four months on: Meta in May vs Meta in September

Same problem pool and tiers. May figures are Llama-4 Maverick from our 31 May 2026 study (12-question batches, 15-minute deadline), so accuracy is directly comparable and timing and cost are indicative. Light purple is May, dark purple is now. The hardest-tier jump, 65.5% to 100.0%, is the single most important number in this report; the cost chart is the honest counterweight: Meta traded a cheap, weak model for a frontier-priced, frontier-grade one.

Accuracy % by tier: Meta's model in May vs September (excl. failed attempts) higher is better
May, easy95.2%
Sep, easy97.5%
May, medium89.3%
Sep, medium98.8%
May, hard86.9%
Sep, hard100%
May, stretch65.5%
Sep, stretch100%
May, all tiers84.2%
Sep, all tiers99%
■ Meta, May 2026 (Llama-4 Maverick)   ■ Meta, Sep 2026
Stretch-tier accuracy %: Meta vs Anthropic, then and now higher is better
Claude Opus 4.8, May100%
Muse Spark 1.3, Sep100%
Claude Opus 5, Sep100%
Llama-4 Maverick, May65.5%
■ Meta, May 2026 (Llama-4 Maverick)   ■ Meta, Sep 2026   ■ Western frontier
Seconds per question by tier: May vs September lower is better
May, easy4.5 s
Sep, easy2.8 s
May, medium8.6 s
Sep, medium5.6 s
May, hard9.9 s
Sep, hard4.7 s
May, stretch12.5 s
Sep, stretch7.9 s
May, all tiers8.9 s
Sep, all tiers5.2 s
■ Meta, May 2026 (Llama-4 Maverick)   ■ Meta, Sep 2026
Cents per question: May vs September (Muse is priced as a flagship; Llama-4 was not) lower is better
May, stretch0.03¢
Sep, stretch0.39¢
May, all tiers0.02¢
Sep, all tiers0.27¢
■ Meta, May 2026 (Llama-4 Maverick)   ■ Meta, Sep 2026

02The evidence

Muse is shown in purple throughout. The left-hand accuracy chart drops failed attempts (the like-for-like view); the right-hand one scores a timed-out batch as zero (what a buyer under a time budget experiences).

Accuracy % EXCLUDING failed attempts, all tiers pooled (like-for-like) higher is better
Gemini 3.8 Flash99.4%
Claude Opus 599.3%
Muse Spark 1.399%
GPT-5.6 Terra Pro98.7%
Qwen3.8-27B97.9%
GLM (latest)96%
■ Meta, Sep 2026   ■ Western frontier   ■ Western mid-tier   ■ Chinese flagship / open
Accuracy % OVERALL, all tiers pooled (a timed-out batch scores 0/10) higher is better
Gemini 3.8 Flash99.4%
Claude Opus 599.3%
GPT-5.6 Terra Pro98.7%
Muse Spark 1.395.8%
GLM (latest)77.4%
Qwen3.8-27B60%
■ Meta, Sep 2026   ■ Western frontier   ■ Western mid-tier   ■ Chinese flagship / open
STRETCH tier - accuracy % excluding failed attempts higher is better
Gemini 3.8 Flash100%
Claude Opus 5100%
Muse Spark 1.3100%
GPT-5.6 Terra Pro100%
GLM (latest)95.7%
Qwen3.8-27B93.3%
■ Meta, Sep 2026   ■ Western frontier   ■ Western mid-tier   ■ Chinese flagship / open
Seconds per 10-question batch, all tiers pooled lower is better
Claude Opus 537.3 s
Muse Spark 1.352.1 s
GPT-5.6 Terra Pro52.9 s
Gemini 3.8 Flash55.0 s
GLM (latest)68.8 s
Qwen3.8-27B102 s
■ Meta, Sep 2026   ■ Western frontier   ■ Western mid-tier   ■ Chinese flagship / open
Seconds per batch by tier - Muse vs the fast models (how steeply time climbs with difficulty) lower is better
Muse Spark 1.3 - easy28.5 s
Muse Spark 1.3 - medium56.2 s
Muse Spark 1.3 - hard47.2 s
Muse Spark 1.3 - stretch79.3 s
Claude Opus 5 - easy25.2 s
Claude Opus 5 - medium30.2 s
Claude Opus 5 - hard47.2 s
Claude Opus 5 - stretch51.4 s
Gemini 3.8 Flash - easy34.7 s
Gemini 3.8 Flash - medium48.4 s
Gemini 3.8 Flash - hard58.3 s
Gemini 3.8 Flash - stretch81.9 s
GPT-5.6 Terra Pro - easy30.5 s
GPT-5.6 Terra Pro - medium47.9 s
GPT-5.6 Terra Pro - hard60.5 s
GPT-5.6 Terra Pro - stretch75.6 s
■ Meta, Sep 2026   ■ Western frontier   ■ Western mid-tier
STRETCH tier - seconds per 10-question batch lower is better
Claude Opus 551.4 s
GPT-5.6 Terra Pro75.6 s
Muse Spark 1.379.3 s
Gemini 3.8 Flash81.9 s
GLM (latest)100 s
Qwen3.8-27B123 s
■ Meta, Sep 2026   ■ Western frontier   ■ Western mid-tier   ■ Chinese flagship / open
Output tokens per question, all tiers pooled (incl. hidden reasoning) lower is better
Claude Opus 5326
Muse Spark 1.3612
GLM (latest)683
Gemini 3.8 Flash693
Qwen3.8-27B810
GPT-5.6 Terra Pro999
■ Meta, Sep 2026   ■ Western frontier   ■ Western mid-tier   ■ Chinese flagship / open
Reasoning share of output %, all tiers pooled lower is better
GPT-5.6 Terra Pro38.6%
Claude Opus 538.7%
Gemini 3.8 Flash52.8%
GLM (latest)55.3%
Qwen3.8-27B56.1%
Muse Spark 1.366%
■ Meta, Sep 2026   ■ Western frontier   ■ Western mid-tier   ■ Chinese flagship / open
Cents per question, all tiers pooled (metered) lower is better
Qwen3.8-27B0.23¢
Gemini 3.8 Flash0.27¢
Muse Spark 1.30.27¢
GLM (latest)0.31¢
Claude Opus 50.88¢
GPT-5.6 Terra Pro1.48¢
■ Meta, Sep 2026   ■ Western frontier   ■ Western mid-tier   ■ Chinese flagship / open
Cents per CORRECT answer, all tiers pooled lower is better
Qwen3.8-27B0.23¢
Gemini 3.8 Flash0.27¢
Muse Spark 1.30.28¢
GLM (latest)0.33¢
Claude Opus 50.89¢
GPT-5.6 Terra Pro1.50¢
■ Meta, Sep 2026   ■ Western frontier   ■ Western mid-tier   ■ Chinese flagship / open

Scorecard

Muse Spark 1.3vs. fieldRead
Accuracy, excl. failed99.0%#3 of 6; leader 99.4%Parity. Perfect on hard and stretch tiers.
Accuracy, overall (timeouts = 0)95.8%#4; Opus 99.3%, GPT 98.7%One blown deadline in 31 batches costs it 3 points.
Speed, s per 10-question batch52#2; Opus 37, GPT 53, Gemini 55Fast on average...
Speed range on hardest tier35 to 170 sOpus never exceeded about a minute...but wildly variable. This is the weakness.
Cost, cents per question0.27¢#3; Opus 0.88¢ (3.3x), GPT 1.48¢ (5.5x)Budget-tier bill at flagship list price.
Reasoning share of output66%Highest in field; Opus 39%Buys accuracy with compute. Works; caps price cuts.

Attempt accounting

ModelBatches attemptedAnsweredTimed out (180 s)Credit/policy errorsAnswered % of valid attemptsAcc % excl. failedAcc % overall
Gemini 3.8 Flash31310010099.499.4
Claude Opus 531290210099.399.3
Muse Spark 1.33130109799.095.8
GPT-5.6 Terra Pro31310010098.798.7
Qwen3.8-27B31191206197.960.0
GLM (latest)3125608196.077.4

Timeouts are recorded when a single 10-question call exceeds 180 s. They may reflect provider/OpenRouter routing conditions rather than the model, so the excluding-failed view is the like-for-like comparison; the overall view is the buyer's experience under a time budget.

03Views, risks and method

What would change our view

Risks to the view

One domain (competition math, text only), a 99% ceiling that cannot separate the leaders, small per-tier samples (one wrong answer moves a tier 10 points), provider routing inside every timing. This is a first look, not a coverage initiation.

How the study was run

WhoOne person directing an AI coding agent (Claude Code). No engineering team, no data-science team, no vendor access.
TimeStart to published report: about two hours. The model runs themselves took 78 minutes.
Cost$10.02 of metered OpenRouter spend across all six models, plus a $20 top-up mid-run when the account hit zero.
QuestionsOur own curated store of about 700 text-only AMC 8 / AJHSME problems with verified answer keys, built earlier for a kids' math site.
InfrastructureA $20-a-month Vultr box we already had. The GenAware Model Lab that ran our May and June papers was switched back on, its lab app patched (metered cost, reasoning tokens, a 3-minute deadline, provider defaults), and pointed at the new model list.
Runs31 rounds, 1650 graded model-questions, every model seeing the identical ten questions per round, all metrics captured live and published to this page while the run was still going.
ReproducibilityEvery batch, every model's full working, and every metered cost sits in the raw CSV and the live session pages linked at the top. Change one line to test a different model.