Evidence-led AI briefings and free decision tools for engineers, researchers, and product builders.
Sep 7, 2026
Same 10 in and 50 out per million tokens. One burns 3.3x the tokens per index task and bills 2.4x. One cache line sits 4x apart.
Sep 5, 2026
A study turns software knowledge into tested instructions for research agents. The gains are substantial. Understanding where they come from takes a closer look.
Sep 4, 2026
Scope violations fell to 0%. Reasoning monitorability fell too. Your agent can get a 403.
Sep 2, 2026
Anandkumar's physics model is graded by the equation itself, not by labels.
Aug 31, 2026
The full stack reached 33, at twelve times the cost and 108 times the tokens.