| researchaudio.io · issue |
diff card-4p6.pdf → 4p7card.pdf |
|
2 files model cards |
3 corrected in place, aug 17 |
1 deleted hallucination row |
5 peers off the coding charts |
|
Grok 4.7 Beats a Grok 4.6 That Was Corrected in August
Three regressions shrank by revision. Opus 5 left the coding charts. Biology fell by design.
september 21, 2026 · source: the Grok 4.7 model card, the revised Grok 4.6 model card, xAI's launch post, Artificial Analysis, Datacurve
|
|
On August 13 this newsletter read the Grok 4.6 model card and counted five rows that moved the wrong way: self-harm compliance, dishonesty under pressure, sycophancy, dual-use cyber compliance, and hallucination rate. Four days later the PDF at the same URL was replaced. A single changelog line says results on HackerBench v0.2, Self-harm, MASK and LAB were corrected. Three of the five regressions got smaller by revision. The Grok 4.7 card, published today, measures itself against the corrected numbers.
That is the frame for the new card. A model card is a set of comparisons, and every comparator in this one moved between August and September: the 4.6 baseline was corrected in place, three benchmarks changed version, and the peer set on the coding charts no longer includes the models Datacurve scores highest. Below: what moved, what was kept, what was dropped, and what to do with a number whose PDF can change under you.
|
|
Grok 4.7 shipped September 21 in Cursor, Grok Build and the API, with gateways and Office add-ins listed in the card and consumer surfaces promised later. The launch post says it uses a new, larger base model; the card gives no size, and the 2.1 trillion figure circulating on X comes from Musk rather than either document, so it stays out of this issue. Pretraining cutoff moved from January 2026 to June 2026, with generated data through August. The model page keeps everything else: 500k context, 2.00 and 6.00 per million tokens, 0.50 cached, doubled rates above 200k prompt tokens, four reasoning efforts with high as default. A fast variant costs twice as much.
Two sentences carried over from the 4.6 card word for word. The footnote disclosing supplemental training on anonymized Cursor workflow data, and the overview line saying the model reaches results with fewer steps and fewer output tokens than other frontier models. Section 04 puts a meter on the second one.
The launch table pairs Grok 4.7 at xhigh with Grok 4.6 at high in every row. The card itself has 4.6 at xhigh where the gap is narrower: on EEBench the launch table shows 53.0 to 64.0, the card shows 60.0 to 66.0. Read the card, and 4.7 has two different EEBench scores at the same effort published on the same day, 64.0 and 66.0. The card does not reconcile them.
|
| @@ 02 · the baseline moved @@ |
The five regressions from the August 13 issue met three fates. Three were corrected in place on August 17. One was carried unchanged. One row no longer exists in the new card. The Grok 4.5 column did not move at all across the two revisions, which says the correction touched the 4.6 runs, not the suites.
| row (lower is better) | 4.5 | 4.6 aug 12 | 4.6 aug 17 | 4.6 in 4.7 card | 4.7 | fate |
| self-harm compliance | 0.50 | 3.7 | 0.84 | 0.84 | 1.05 | corrected; 4.7 worse than the corrected 4.6 |
| MASK-Rectified dishonesty | 0.67 | 3.8 | 1.90 | 1.90 | 0.00 | corrected |
| HackerBench dual-use compliance | 7.8 | 16.7 v0.2 | 6.9 v0.2 | 5.93 v0.3 | 3.31 high · 4.02 xhigh | corrected, then re-versioned |
| sycophancy | 0.01 | 0.04 | 0.04 | 0.04 | 0.03 | carried |
| hallucination rate | 0.98 | 1.7 | 1.7 | no row | no row | deleted with its section |
aug 12 figures preserved by NOPE Insights (self-harm, MASK) and the BenchLM mirror (HackerBench), both matching what this newsletter read on aug 13. aug 17 and sept 21 figures from the two cards.
A correction can mean a re-run, a regrade, or a broken pipeline fixed. xAI gives one line and no method. What it changes is the story, in both directions. Against the August 12 numbers, Grok 4.7 would improve on all three rows by wide margins: 1.05 against 3.7, 0.00 against 3.8, 3.31 against 16.7. Against the corrected numbers it improves on two and regresses on self-harm, which the new card states in plain words in section 9. So the revision cut 4.6's regressions in August and cut 4.7's headline improvement in September. That is a point in xAI's favour: the correction was not shaped for the next launch.
HackerBench moved a second time without a correction. The 4.7 card runs version 0.3 and lists Grok 4.6 at 5.93, down from 6.9 on version 0.2. In the same chart Grok 4.5 reads 7.8 and GPT-5.6 Sol reads 35.7, the exact decimals from the v0.2 chart. Either those two were not re-run on v0.3, or they were and landed on identical values. The card does not say which.
The hallucination row did not get corrected. Its whole section went. The 4.6 card had a Factuality evaluation (0.98 for 4.5, 1.7 for 4.6) and DeepSearchQA; neither appears in the 4.7 card. By my count the 4.6 card carried 39 evaluation subsections and the 4.7 card carries 29. The R&D enablement section (MTS eval, InferenceEval, KernelBench) is gone, as are APEX-SWE, FrontierCode, APEX-Agents and OfficeQA; GDPval and AA-Briefcase moved from the card to the launch post. Artificial Analysis's independent hallucination measure improved from 34 to 29 percent, so the deleted row would probably have read fine. The reader cannot check that from the card.
|
| @@ 03 · the peers moved @@ |
DeepSWE v1.1 is the cleanest case because every model runs in the same harness (mini-SWE-agent, graded by Datacurve). In August the 4.6 card's DeepSWE chart had Opus 5 on top at 74.0 and ten entries. In September the 4.7 card's chart has five, Opus 5 is not one of them, and Grok 4.7's 71.0 sits second behind GPT-5.6 Sol. Datacurve's own leaderboard, updated September 3, has three models at 74.
| DeepSWE v1.1, pass@1 | 4.6 card, aug | 4.7 card, sept | Datacurve board, sept 3 | cost per task |
| GPT-6 Astra (xhigh) | − | − | 74 | 6.52 |
| Gemini 3.8 Flash (high) | − | − | 74 | 2.36 |
| Opus 5 (max) | 74.0 | − | 74 | 11.84 |
| GPT-5.6 Sol (max) | 73.0 | 72.7 | 73 | 6.46 |
| Grok 4.7 (high) | − | 71.0 | not listed yet | − |
| Fable 5 (max / xhigh) | 70.0 | 69.7 | 70 | 13.41 |
| Kimi K3 (max) | 69.0 | − | 69 | 4.65 |
| GLM-5.3 (max) | − | − | 69 | 3.99 |
| Grok 4.6 | 67.0 xhigh · 65.9 high | 65.2 high | 67 medium | 3.45 |
| Sonnet 5 (max) | 54.0 | 54 | 54 | 26.40 |
cost per task in dollars, from Datacurve's cost column; confidence intervals on the board are plus or minus 1 to 6 points, so 71.0 overlaps the 74s.
Two things are true at once. The 4.7 card's peer numbers come from Datacurve runs instead of provider cards, which would explain why Sol reads 72.7 rather than 73.0 and Fable 5 reads 69.7; that is a better method than the 4.6 card's. And the roster shrank at the same time, dropping the three models that would sit above Grok 4.7. Gemini 3.8 Flash reaches 74 at 2.36 per task on Datacurve's meter, against 11.84 for Opus 5 and 6.52 for Astra.
GPT-6 Astra follows a pattern across the card. It appears on EEBench, HealthBench Professional, the LatchBio suite and BioSecBench, and leads three of the four. It appears on none of the five coding charts: CursorBench, DeepSWE, Terminal-Bench, FrontierSWE, SWE-Marathon. On Datacurve it is tied first on DeepSWE; on Artificial Analysis's Coding Agent Index, Grok 4.7 with Grok Build ranks fourth behind Fable 5.1, Astra and Opus 5. The card does not explain how its rosters are chosen.
Grok 4.6 itself has three DeepSWE numbers before 4.7 is even added: 65.9 at high in its own card, 65.2 at high in the 4.7 card, 67 at medium as Datacurve's listed configuration. Same model, same harness, same grader, three values. Any one of them is fine. Comparing across them is not.
|
| @@ 04 · one benchmark, three numbers @@ |
The card's largest gain is Terminal-Bench 4.0, 20.3 to 38.0. The card also says absolute scores remain sensitive to the agent harness. Here is the sensitivity, measured on the same model on the same day.
| Terminal-Bench 4.0, Grok 4.7 xhigh | who ran it | harness | score |
| model card | Harbor | Grok Build | 38.0 |
| AA Coding Agent Index | Artificial Analysis | Grok Build | 33 |
| AA Intelligence Index | Artificial Analysis | standardized | 26 |
| tokens and cost per AA index task | index | output tokens | cost |
| Grok 4.7 (xhigh), 2.00 / 6.00 | 46 | 81k (59k reasoning) | 3.74 |
| Grok 4.7 (high) | 46 | 66k | 2.73 |
| Grok 4.6 (high) | 44 | 36k | − |
| GPT-5.6 Sol (max), 4.00 / 20.00 | 47 | 29k | 1.99 |
| GPT-6 Astra (max) | 53 | 27k | − |
scores and tokens from Artificial Analysis's launch analysis and comparison page; the standardized-harness 26 as reported by The Decoder. Fable 5.1 reads 57.9 in the card and 55 in AA's harness; Astra, absent from the card, reads about 60 in AA's harness.
Twelve points on one model between the card and the standardized run, and the card is the top of the three. None is wrong; each is a different system. Terminal-Bench 4.0 gives agents up to eight hours, and the harness decides how those hours are spent.
The token line is where the carried-over sentence fails. The overview still says fewer output tokens than other frontier models. Artificial Analysis meters 81,000 output tokens per task at xhigh, 125 percent more than Grok 4.6 at high and 196 percent more than Astra at max. At 2.00 and 6.00 that comes to 3.74 per task; Sol at max, priced 4.00 and 20.00, comes to 1.99 on the same index for one more point. The card's claim about cheap tokens is true and the claim about fewer tokens is not, and the second one is what decides the bill. Same lesson as the September 6 issue on Astra and Fable 5.1: the rate card is not the invoice.
|
| @@ 05 · the drops the card keeps @@ |
Not every decline was corrected or deleted. Section 7 of the card reports Grok 4.7 scoring below Grok 4.6 on eight biology and chemistry suites, and explains it as intended: the result of safer RL environments and better selectivity of training data, as opposed to enhanced safeguards or refusal training. In the same section it says general biological knowledge not useful for dual-use capability holds the same or sees some diminishment. The table sorts the eight by how the card itself describes each suite.
| suite | card's description | 4.5 | 4.6 | 4.7 | 4.6 to 4.7 |
| ProtocolQA open-ended | everyday PCR, cloning, cell culture | 87.0 | 79.6 | 70.4 | −9.2 |
| BixBench zero-shot | no weapons-enabling content | 93.8 | 93.8 | 88.4 | −5.4 |
| LAB-Bench practical | everyday wet-lab skills | 71.1 | 80.7 | 76.8 | −3.9 |
| Biosecurity VCT | dual-use virus-lab troubleshooting | 44.1 | 47.8 | 41.5 | −6.3 |
| VCT | virology, including misusable work | 65.5 | 67.4 | 63.0 | −4.4 |
| WMDP-Cyber | dual-use knowledge, MCQ | 83.2 | 90.1 | 88.1 | −2.0 |
| WMDP-Bio | dual-use knowledge, MCQ | 90.9 | 90.0 | 88.1 | −1.9 |
| WMDP-Chem | dual-use knowledge, MCQ | 87.3 | 85.3 | 84.9 | −0.4 |
| what went up | 4.5 | 4.6 | 4.7 | 4.6 to 4.7 |
| HealthBench Professional, clinical | − | 48.5 | 56.7 | +8.2 |
| BioSecBench Function, agentic analysis | − | 39.9 | 43.3 | +3.4 |
| LatchBio capability suite, agentic analysis | − | 43.3 | 44.5 | +1.2 |
| BioSecBench Refusal | − | 45.6 | 62.4 | +16.8 |
all figures from section 7 and section 5 of the Grok 4.7 card; capability rows measured without production safeguards, per the card.
The three suites the card describes as not weapons-relevant fell 9.2, 5.4 and 3.9 points. The three dual-use multiple-choice probes fell 1.9, 0.4 and 2.0. The subtraction landed harder on everyday protocol work than on the dual-use knowledge it was aimed at. ProtocolQA has slid two releases running, 87.0 to 79.6 to 70.4, 16.6 points on troubleshooting ordinary PCR and cloning steps.
The rows that rose are the agentic ones: reading real experimental files with tools (LatchBio, BioSecBench Function) and clinical conversation (HealthBench). Read together, the model does more with files and tools and recalls less from memory. That is consistent with what the card says was done: RL environments for the work, data selection for the knowledge.
On August 9 this newsletter covered the Stanford and Arc phage paper under the title The Guardrail Was Deleted Data. Here is a lab writing that guardrail into its own model card and printing the cost. The difference from a refusal is durability. A refusal can be prompted around, and UK AISI showed in July that refusal training reverses with a fine-tune. A capability that was never trained in cannot be reversed by the user. If you route wet-lab or protocol questions through Grok, the recall number moves with each release, and it has moved down twice.
One claim to file against this table. The launch post says Grok 4.7 tops LatchBio's biosafety benchmark at 62.4 percent. That is the Refusal component of BioSecBench. On the Function component, recovering biosecurity-relevant conclusions from data, Grok 4.7's 43.3 trails Astra at 46.0 and Opus 5 at 44.8. The card says the three components should be read together; the launch post read one.
|
| @@ 06 · what holds for xAI @@ |
The card discloses the declines it keeps in plain sentences: self-harm compliance is higher, general refusals went 0.93 to 1.10, CVE-Bench went 39.8 to 37.7 at high and 36.6 at xhigh while the section summary claims small cyber gains. A changelog exists at all, which is more than a silent replacement. Every jailbreak row improved: standard 0.04 to 0.01, StrongReject 3.9 to 2.0, long-horizon 1.0 to 0.65. Child safety stays at 0.0; bio refusal recall holds at 100 and chem slips from 100 to 99.9. MASK-Rectified reads 0.00 with the 4.5 baseline untouched.
Independently, Artificial Analysis measures a real step on agentic knowledge work: 1657 Elo on AA-Briefcase, up 111, behind Opus 5 and Fable 5.1; 56 on the Coding Agent Index, up 9. Price held at 2.00 and 6.00 while the base model grew, and the August 13 cliff math applies unchanged.
The Legal Agent Benchmark row is the card's widest lead, 19.6 against 6.7 for Fable 5.1 and 2.5 for Sol, on a Vals AI run with Harvey's every-criterion grading. Two cautions stay in view: four in five matters still fail, and Fable 5 scores 11.3 on the same chart, above its successor at 6.7, which says as much about the benchmark's variance as about any model.
|
| @@ 07 · what to do with a card that moves @@ |
| + | Store the PDF and its hash when you cite a model card. The URL is not a version. The 4.6 card's changelog is the courtesy; the Wayback snapshot NOPE links is the proof. |
| + | Compare inside one document, at one effort, in one harness. The launch table pairs 4.7 xhigh with 4.6 high; the card has 4.6 xhigh. Use the card. |
| + | Price the task, not the token. 81,000 output tokens at 6.00 per million is 0.49 before a single input token. Meter one representative job on your own trace and route on that number. |
| − | Stop treating a peer's absence from a chart as neutral. Three models at 74 on DeepSWE are missing from a chart where 71.0 ranks second. |
| − | Stop assuming a science regression is a bug that the next release fixes. This card calls it a design choice, and design choices compound. |
The most useful sentence in the Grok 4.7 card is not in the Grok 4.7 card. It is the changelog line in the Grok 4.6 card, and it says the number you read is the number until it is not.
|
sources read in full: Grok 4.7 model card (sept 21) · Grok 4.6 model card (rev. aug 17) · launch post · model page · Artificial Analysis · Datacurve DeepSWE · NOPE Insights · BenchLM · VentureBeat. Subsection counts are mine.
ResearchAudio.io · frontier models, read against their own receipts
|
|