Sep 9, 2026
GPT-Image-2 paints objects onto a render, SAM3D lifts them. The evaluation has no numbers.
Sep 7, 2026
Same 10 in and 50 out per million tokens. One burns 3.3x the tokens per index task and bills 2.4x. One cache line sits 4x apart.
Sep 5, 2026
A study turns software knowledge into tested instructions for research agents. The gains are substantial. Understanding where they come from takes a closer look.
Sep 4, 2026
Scope violations fell to 0%. Reasoning monitorability fell too. Your agent can get a 403.
Sep 2, 2026
Anandkumar's physics model is graded by the equation itself, not by labels.
Aug 31, 2026
The full stack reached 33, at twelve times the cost and 108 times the tokens.
Aug 29, 2026
Fabrication detection went from 5 of 36 to 33 of 36. What each layer bought.
Aug 27, 2026
57 on the intelligence index, 50.2 tokens a second, and a rate that ends September 9.
Aug 24, 2026
Measured: 1.80 against a 1.50 chance baseline across 300 tokens.
Aug 21, 2026
It paused two weeks instead. The phrase that kept the clause shut.