|
Vendors do not usually publish the chart where they come second. On August 5, Meta Superintelligence Labs published three of them.
Meta released Muse Code, a terminal coding agent in beta, alongside Muse Spark 1.2, the coding-focused model that runs it. The launch post carries three benchmark bar charts and states no placements in the prose at all. Read the charts and you see why: Claude Opus 5 is ahead on every board Meta chose to show, including the internal one Meta built and controls.
Which makes the launch more interesting, not less. The score is not the product here. The product is a runtime designed to survive a 24 hour job, and a second price tier that undercuts the entire coding-model market by more than an order of magnitude, as long as you let Meta keep what your agent reads.
|
THE SHAPE OF ITMeta lists two model IDs against the same checkpoint. One is priced like a frontier model. The other is priced like nothing, and the difference is settled in your source code.
|
$ 01
What actually shipped
Two things on the same day, deliberately coupled. Muse Code is a terminal agent installed with one line on macOS or Linux. No Windows build, no GUI, no IDE extension. Meta describes its job as planning changes, writing code, and validating results across large repositories.
Muse Spark 1.2 is the model underneath it: a coding-focused revision of Muse Spark 1.1 with a 1,048,576 token context window, closed weights, served through the Meta Model API. Meta says it scaled training compute on coding tasks and widened the diversity of training environments, targeting long-horizon work rather than autocomplete.
The coupling is the claim. Meta states the two were co-trained, so the model was tuned inside the runtime it ships with. Three bundled skills describe the intended working style better than the marketing copy does: /plan produces an approval-gated plan, /grill stress-tests that plan until it holds, and /goal drives toward completion. Plan, attack the plan, then execute. That is the pattern experienced Claude Code and Codex users converged on by hand, shipped as first-class commands.
$ 02
The rate card, and what the gap is worth
There is no subscription. Billing is per token, on one of two model IDs. Meta's developer post is direct about the trade: muse-spark-1.2 runs at standard pay-as-you-go rates and Meta states those prompts are not used to improve its products. muse-spark-1.2-contributor is rate-limited by tokens in a rolling five hour window instead of by request count, is available in select countries, and carries permission for Meta to train on what you send.
| RATE CARD · USD PER 1M TOKENS |
STANDARD |
CONTRIBUTOR |
GAP |
| Input |
1.25 |
0.10 |
1.15 12.5x |
| Output |
4.25 |
0.20 |
4.05 21.3x |
| Cached input |
0.15 |
0.002 |
0.148 75x |
| Meta trains on your data |
no |
yes |
the whole trade |
Rates from Meta's developer post and Meta Model API listings, August 5, 2026. Gap column is arithmetic on those two published prices.
That gap column is the number worth sitting with. Meta lists both IDs against the same checkpoint, so nothing about the model changes between the two rows. The 1.15 per million input tokens is what Meta is willing to forgo in exchange for the right to train on what you sent. Meta set that figure, not a market. It is a valuation of your prompts, published on a price page.
A third path exists and got almost no coverage: Meta's developer post says it is beginning to accept requests for zero data retention through its sales team. If you are in a regulated environment, that line matters more than either price.
$ 03
Why the price is shaped like that
The obvious read is that Meta is buying market share. The launch post suggests something more specific. Under a heading called Self-Improvement, Meta describes how 1.2 was trained: Muse Spark 1.1 generated challenging coding environments and instruction-following templates, then graded candidate solutions against those requirements, producing what Meta calls a scalable training dataset for 1.2.
Meta also states the co-training used rejection-sampled harness trajectories. Put plainly: the input to this model was agent runs. Not code, not documentation. Traces of an agent working inside a repository, with a grade attached.
| THE TRAINING INPUT, AS META DESCRIBES IT |
| SOURCE |
WHAT IT PRODUCED |
STATUS |
| Muse Spark 1.1 |
Coding environments and instruction templates, self-graded |
stated by meta |
| Muse Code harness |
Rejection-sampled trajectories, plus recipes for goals, compaction, subagents |
stated by meta |
| Long-horizon runs |
Whole-repository generation, large end-to-end projects, auto-research |
stated by meta |
| Contributor tier |
Real agent runs, real repositories, permission attached at the model ID |
reading, not stated |
Rows one to three paraphrase Meta's launch post. The bottom row is inference from the pricing structure. Meta has not said the contributor tier feeds a specific future model.
Hold that pipeline next to the price sheet and the shape resolves. The thing Meta spent 1.1's inference budget manufacturing synthetically is the exact thing the contributor tier acquires organically, from repositories that actually exist, with human steering already in the trace. Meta priced it at roughly a twelfth of list because that is what the data appears to be worth on the other side of the ledger.
$ 04
The scores are harness scores
Meta's evaluation methodology report is unusually candid, and it undercuts its own charts. Every model was run inside its own vendor agent: Muse Code for 1.2, mini-swe-agent for 1.1, Claude Code for Opus, Codex for GPT, Grok Build for Grok, Antigravity for Gemini, Kimi Code for Kimi. Meta adds that its setup may not be tuned for third-party models, so those results may not reflect their best performance.
The cleanest way to see what that does is to take one fixed thing, the 1.1 to 1.2 upgrade, and measure it three ways.
| ONE UPGRADE, THREE HARNESSES · TERMINAL-BENCH 2.1 |
| MEASURED BY |
1.1 |
1.2 |
DELTA |
Meta's deck mini-swe-agent → muse code |
76.2 |
82.9 |
+6.7 |
Artificial Analysis one fixed harness, both models |
78 |
80 |
+2.0 |
Vals AI terminus 2, same for every model |
1.2 ranks 14 of 50 on Terminal-Bench 2.1, and 5 of 45 on the overall Vals Index at 71.88 percent |
| The tell: Meta's own model page lists Muse Spark 1.1 at 80.0 on this benchmark. Meta's 1.2 comparison deck implies 76.2 for the same model. Same vendor, same model, same benchmark, 3.8 points apart. The variable is the harness. |
Meta figures from its launch charts and methodology report. Independent figures from Artificial Analysis and Vals AI. No verified entry for either 1.2 or Opus 5 exists on the official Terminal-Bench leaderboard as of writing.
A Terminal-Bench row is not a model score. It is a score for a harness plus a model plus an effort setting, and Meta swapped the harness underneath the two ends of its own generational comparison. Roughly two thirds of the headline gain sits in a variable Meta changed.
One improvement survives independent measurement cleanly, and it is worth naming because it is genuine. Artificial Analysis reports Muse Spark 1.2's GDPval-AA v2 Elo rose 260 points to 1631, fifth among all models it has benchmarked and ahead of Claude Opus 4.8 at max effort. Agentic knowledge work was the clearest gap in 1.1. That gap closed measurably, in someone else's harness.
The trade shows up in latency. Artificial Analysis clocked time to first token at 26.12 seconds against 1.1's 2.90 seconds at maximum reasoning effort, buying three points of Intelligence Index, 51 to 54. Nine times the wait is absurd for a chat completion and invisible on a twelve-file refactor. That tells you what this model was built for.
$ 05
The runtime is the actual pitch
Strip out the charts and what remains is an engineering argument about not dying. Two mechanisms carry it.
The first is the event log. Muse Code appends every model call, tool run, approval, and edit to a local append-only log. Meta calls the runtime replay-exact and restart-safe: crash it, lose the machine, and the agent resumes precisely where it stopped instead of re-deriving context from scratch. Competitors have crash recovery of varying quality. None of them made it the headline. It is also an audit artifact, a complete local record of what an autonomous process did to your working tree, which is the exact document a security review asks for and which most agent CLIs cannot produce.
The second is async background agents. The distinction from ordinary subagents is that they are not spawned per subtask and torn down. They stay alive across the session, carry out next steps on their own, and choose when to report back to the main loop. Meta's claim is lower latency and less steering, because the orchestrator is not paying spin-up cost every time it delegates.
The supporting evidence is one case study: Muse Code running up to 24 hours across more than 1,000 tool calls, iteratively optimizing GPU kernels on NVIDIA Hopper hardware. KDA against the FLA Triton baseline, with third-party kernel libraries prohibited, and MLA against a PyTorch reference at batch size 1, 64 heads, sequence length 8192, latent dimension 512. Meta published no speedup figure, so the interesting quantity is not the result. It is the duration. A thousand tool calls without human rescue is a different failure surface than a fifty-call refactor, and it is a workload where a 3.8 point benchmark deficit matters far less than whether the process is alive at hour nine.
$ 06
Run the arithmetic on your own loop
Take a realistic agent step: 60,000 tokens of repository context in, 3,000 tokens of plan and patch out. Two independent choices (which tier, and whether your context hits the cache) produce four very different bills.
| COST PER 1,000 AGENT STEPS · 60K IN / 3K OUT |
| CONFIGURATION |
PER STEP |
PER 1,000 STEPS |
| Standard, cold context |
8.8 cents |
87.75 |
| Standard, cache hit |
2.2 cents |
21.75 |
| Contributor, cold context |
0.66 cents |
6.60 |
| Contributor, cache hit |
0.07 cents |
0.72 |
| Same work, top row to bottom row, is a factor of 122. Meta's own 24 hour, 1,000 call kernel run costs on the order of 88 dollars in tokens on the tier that protects your data, and under a dollar on the tier that does not. |
Computed from Meta's published per-token rates. Reasoning tokens bill at the output rate, so real spend on hard tasks runs above these figures.
|
THE DECISION THIS PUTS IN FRONT OF YOUA terminal coding agent reads your codebase. That is its entire function. So the prompts on the contributor tier contain your source, your internal APIs, your comments explaining why the workaround exists, and whatever your test fixtures happen to hold. For a side project or an open tree, the arithmetic is trivially in favor. For a proprietary repository, or any repository within reach of customer data, this is a licensing decision wearing a price tag, and it belongs in front of whoever signs off on data handling rather than whoever watches the cloud bill.
|
$ 07
What nobody has checked yet
Five things worth holding loosely.
01 Nothing about the event log has been verified by anyone outside Meta. Replay-exact and restart-safe are strong assertions about a beta runtime.
02 The co-training premium is untested. Nobody has run Muse Spark 1.2 inside a rival harness to see whether the pairing does anything, or whether the model simply behaves like a competent coding model wherever you put it.
03 The 82.9 belongs to the model inside Muse Code. Call the API from your own framework and you have bought the model half and none of the harness half. Two separate evaluations, easily conflated.
04 Meta benchmarked against the second model at two of the three major labs: Opus 5 rather than Fable 5, GPT-5.6 Terra rather than Sol. The model holding the top verified Terminal-Bench slot does not appear on Meta's chart at all.
05 No open weights. From the company that spent three years as the standard-bearer for open models, the launch post does not mention weights at all. Treat this as a hosted dependency with exactly one provider.
Practical shape of the week: if you run long-horizon autonomous jobs, migrations, monorepo dependency upgrades, anything you currently babysit because the agent loses its place, Muse Code earns an afternoon on a scratch clone. The event log and persistent background agents target exactly that failure mode, and the benchmark deficit is smallest where task duration is longest.
If you are running a team on Claude Code or Codex against a proprietary repository, there is nothing here that forces a move. Muse Code sits second or third on every board its own vendor selected, the standard tier prices at parity with the field rather than below it, and the differentiator is a reliability property nobody outside Meta has stress-tested.
The thing worth watching is not the next benchmark. Every lab can move a benchmark by moving a harness, and Meta just demonstrated that inside a single generational comparison. What has not happened yet is somebody outside Meta running a thousand tool calls through that event log and reporting what came out.
|
Meta published the charts it loses on because the charts were never the argument. The argument is that a coding agent is a data collection surface, and Meta is the first lab to price that honestly on a public page.
|
|