Evidence-led AI briefings and free decision tools for engineers, researchers, and product builders.
Sep 11, 2026
DeepSeek claimed to beat Anthropic and OpenAI
Sep 9, 2026
GPT-Image-2 paints objects onto a render, SAM3D lifts them. The evaluation has no numbers.
Sep 7, 2026
Same 10 in and 50 out per million tokens. One burns 3.3x the tokens per index task and bills 2.4x. One cache line sits 4x apart.
Sep 5, 2026
A study turns software knowledge into tested instructions for research agents. The gains are substantial. Understanding where they come from takes a closer look.
Sep 4, 2026
Scope violations fell to 0%. Reasoning monitorability fell too. Your agent can get a 403.