Aug 4, 2026
Ten proofs. About 2,000 dollars in tokens. The strongest verification any machine-produced mathematics has cleared, and three questions it still leaves open.
Sixteen browser-local tools, twenty-one deployment and cost field notes, and a four-test starter kit for AI evidence, reliability, ROI, voice systems, GPU sizing, and Codex configuration.
Aug 3, 2026
Full history scored 71.0%. Pruning plus a summary hit 91.6% on 62.7% fewer tokens.
It's the difference between an agent that gets sharper every week and one that confidently repeats the same mistake at 3am. There's already a name for it. Almost no one's applied it to release and incident work yet.
Jul 31, 2026
141,006 runs reviewed. Three companies compromised. The sandbox had internet.
Jul 29, 2026
Kimi K3 just went public: 2.8 trillion parameters, 896 experts with only 16 awake per token, and a million token context made possible by deleting most of its attention. The size makes the headline. The accounting is the story.
Jul 28, 2026
A 2.5x efficiency claim, a 96-shard download, and fine print.
Jul 27, 2026
The same self-verification wins benchmarks and stalls 24-hour agent runs.
Jul 25, 2026
Anthropic's Economic Index shipped as a connector. The delegation data inside.
Jul 24, 2026