Muse Spark 1.3: Meta's Frontier Coding and Agentic Model, Explained
Meta's fourth Muse Spark release in five months reaches the frontier — with the top agentic-work benchmark, unchanged pricing, and a max-reasoning variant in limited preview.

Muse Spark 1.3 is Meta's new frontier model for coding and agentic work, released September 2, 2026 — Meta's fourth Muse Spark model in five months. The headline: its top variant lands just behind Claude on the leading intelligence index, it posts the #1 score on a key agentic-work benchmark, and it does it at the lowest cost per task of any model in its class. Here's what changed versus Muse Spark 1.2, the two variants, the benchmarks, the pricing, and where it fits for people building agents.
What Is Muse Spark 1.3?
Muse Spark 1.3 is a reasoning model from Meta, tuned for long-horizon coding and agentic workflows. It accepts text, image, and video input, has a 1M-token context window, and is available in Muse Code and Meta's Model API from launch day.
Meta's own framing is that the model is built to sustain longer work by collaborating with the user and juggling multiple workflows in a single long thread. In practice that means asking clarifying questions when a prompt is ambiguous, invoking help when it's stuck, and confirming before consequential actions — behavior aimed at agents that run for a long time rather than one-shot chat.
Meta AI chief Alexandr Wang told Axios the update is "very competitive with frontier models" and helps pave the way for products like personal agents that work on your behalf. The launch landed the same week as Google's Gemini 3.8 Flash and Anthropic's Claude Fable 5.1 — a crowded few days for model releases.
The Two Variants: xhigh vs max
Muse Spark 1.3 ships as two models with different reasoning budgets:
- Muse Spark 1.3 (xhigh) — available now in Muse Code and the API.
- Muse Spark 1.3 (max) — a higher-effort variant in limited preview for Meta's partners, rolling out more broadly after additional safety testing.
The difference is how hard each variant thinks. On the Artificial Analysis Intelligence Index, xhigh scores 61 and max scores 62. The max variant reaches its higher numbers by spending more — roughly 62% more reasoning on one agentic eval and 28% more on another compared to xhigh. If you want the best result and can absorb the token cost, max is the target; if you want the price-performance sweet spot available today, xhigh is it.
Benchmarks: Where Muse Spark 1.3 Reaches the Frontier
On the Artificial Analysis Intelligence Index, Muse Spark 1.3 (xhigh) enters at 61 — up 4 points from Muse Spark 1.2 (57) and 8 points from Muse Spark 1.1 (53). It ties GPT-5.6 Sol (max) and Grok 4.6 (high). The max variant at 62 sits behind only Claude Fable 5.1 and Claude Opus 5 at the top of the index.
Most of the gain comes from agentic and scientific tasks, not raw knowledge.
| Benchmark | Muse Spark 1.2 | 1.3 (xhigh) | 1.3 (max) |
|---|---|---|---|
| Tau3-Bench Banking | 35% | 47% | 52% |
| Terminal-Bench 2.1 | 80% | 85% | 86% |
| GDPval-AA v2 (Elo) | 1,615 | 1,709 | 1,754 |
| CritPt (scientific) | 18% | 26% | ~25% |
| GPQA Diamond | 90% | 94% | 94% |
Source: Artificial Analysis, September 2, 2026.
The standout is Tau3-Bench Banking: the max variant's 52% is the #1 score among all models, and the 12-to-17-point jump over Muse Spark 1.2 is the single biggest driver of the index improvement. Scientific reasoning rose across the board, led by an 8-point CritPt gain for xhigh.
The Regressions Meta Didn't Headline
Two benchmarks went the other way versus Muse Spark 1.2. Both variants dropped 4 points on AA-LCR (83% to 79%), and AA-Omniscience (Accuracy) fell 3 points for xhigh and 1 point for max. Per Artificial Analysis, the omniscience drop comes from a higher abstention rate — the model declining to answer when unsure — which also lowered its hallucination rate. That's arguably a feature for agent work: a model that says "I don't know" instead of confidently guessing is safer to hand a long, consequential task.
Meta says it trained Muse Spark 1.3 to have a better sense of what it does and doesn't know, and to flag when it hits a wall instead of hallucinating an outcome — consistent with that abstention behavior.
Pricing: The Cheapest Frontier-Tier Model per Task
Pricing is unchanged from Muse Spark 1.2: $1.25 per 1M input tokens and $4.25 per 1M output tokens, with cache hits discounted to $0.15 per 1M. Wang characterized the pricing as "aggressive."
The comparison that matters is cost per task. Artificial Analysis measures Muse Spark 1.3 (xhigh) at $0.55 per Intelligence Index task — the lowest cost of any model scoring 59 or above. Its direct peers at 61 cost far more: Grok 4.6 (high) at $0.94 and GPT-5.6 Sol (max) at $0.95, a premium of over 70%. Meta hasn't published pricing for the max variant yet.
One caveat: cost per task rose versus Muse Spark 1.2 ($0.40), because the new model consumes ~57% more input tokens per agentic task. It's cheaper than its peers, but not cheaper than its predecessor.
What Changed for Coding
Meta says Muse Spark 1.3 was trained on more long-horizon coding tasks and is more usable in real engineering work. Compared to Muse Spark 1.2, it takes fewer turns where they aren't needed, is less verbose, and has a cleaner coding style. In Meta's own engineer comparisons it used roughly 20% fewer tool calls and 25% fewer tokens — efficiency that adds up across a long agent run.
Meta also reports safety improvements most relevant to agents: stronger resistance to prompt injection and adversarial inputs, and better calibration on which actions are irreversible before proceeding.
How It Compares to the Rest of the Field
Muse Spark 1.3 (xhigh) is best understood as the value pick at the frontier's edge: index-tied with GPT-5.6 Sol (max) and Grok 4.6 (high), leading them on cost per task, and topping the field on agentic banking tasks. It trails Claude Fable 5.1 (66) and Claude Opus 5 (63) on the overall index — so if you want the single strongest model regardless of price, Claude still leads; if you want frontier-tier agentic performance at the lowest cost, Muse Spark 1.3 is the standout.
For a head-to-head of this week's releases, see our Muse Spark 1.3 vs Gemini 3.8 Flash vs Claude Fable 5.1 comparison.
What's Next
Meta says its roadmap includes bigger models and a Muse Spark open-weights release. That follows Muse Glimmer, the 30B open-weight model distilled from Muse Spark that Meta shipped in August — so the open-weight family looks set to keep growing. Wang also flagged safety as one of the most important internal topics right now, noting Meta has more powerful models in development.
Bottom Line
Muse Spark 1.3 is Meta reaching the frontier on agentic and coding work at a genuinely competitive price. The xhigh variant is available today and delivers the lowest cost per task of any model in its intelligence tier; the max variant edges just behind Claude at the top of the index. The honest asterisks: benchmarks are early, cost per task rose versus 1.2, and two evals regressed slightly. But if you build agents and care about long-horizon reliability per dollar, it belongs on your shortlist.
Put Muse Spark 1.3 to Work in an AI Workforce
A frontier model like Muse Spark 1.3 only earns its keep once it's wired into real workflows — tools, files, code review, and multi-step plans — not a chat window. Eigent is an open-source Cowork desktop app that runs a multi-agent AI workforce locally and is model-agnostic, so you can route the right model to each task and keep sensitive work on your own machine. See how agents can review GitHub PRs end to end, then download Eigent to try it.
Frequently Asked Questions
What is Muse Spark 1.3?
Muse Spark 1.3 is Meta's frontier reasoning model for coding and agentic work, released September 2, 2026. It supports text, image, and video input with a 1M-token context window and is available in Muse Code and Meta's Model API. It's Meta's fourth Muse Spark release in five months.
What is the difference between Muse Spark 1.3 xhigh and max?
They differ by reasoning budget. Muse Spark 1.3 (xhigh) is available now and scores 61 on the Artificial Analysis Intelligence Index. Muse Spark 1.3 (max) is a higher-effort variant in limited preview that scores 62 by spending significantly more reasoning tokens, rolling out more widely after extra safety testing.
How much does Muse Spark 1.3 cost?
Pricing for the xhigh variant is unchanged from Muse Spark 1.2: $1.25 per 1M input tokens and $4.25 per 1M output tokens, with cache hits at $0.15 per 1M. Artificial Analysis measures it at $0.55 per Intelligence Index task — the lowest of any model scoring 59 or above. Pricing for the max variant hasn't been announced.
Is Muse Spark 1.3 better than Muse Spark 1.2?
Yes on most tasks. The xhigh variant gained 4 index points and posted large jumps on agentic benchmarks — Tau3-Bench Banking rose from 35% to 47%, Terminal-Bench 2.1 from 80% to 85%. Two evals regressed slightly (AA-LCR and AA-Omniscience), and cost per task rose because it uses more input tokens.
How does Muse Spark 1.3 compare to Claude and GPT?
Muse Spark 1.3 (xhigh) ties GPT-5.6 Sol (max) and Grok 4.6 (high) at 61 on the intelligence index while costing far less per task. The max variant at 62 trails only Claude Fable 5.1 (66) and Claude Opus 5 (63). Claude still leads on the overall index; Muse Spark leads on cost per task and on the Tau3-Bench Banking agentic eval.
Will Muse Spark 1.3 be open-weight?
Meta has said an open-weights Muse Spark release is on its roadmap, following the 30B open-weight Muse Glimmer model it shipped in August. A specific date for Muse Spark 1.3 open weights hasn't been announced.
Recent Posts

Claude Fable 5.1 and Mythos 5.1: What's New, Explained
Claude Fable 5.1 and Mythos 5.1 explained: the same model in two safeguard tiers, with new benchmarks, roughly 25 to 45 percent lower cost, and access details.

Gemini 3.8 Flash: What's New for Coding and AI Agents
Gemini 3.8 Flash brings big coding and agentic-reasoning gains at the same low price, plus a new 3.8 Flash Cyber variant. Benchmarks, pricing, and how to use it.

GPT-6 Astra: What OpenAI's New Model Actually Does
GPT-6 Astra is OpenAI's new flagship model. Here's what it does, how it benchmarks, what it costs, and the caveats behind the AGI headlines.