Logo
ResearchAudio
Search
Latest
Free tools
Community
Login
Subscribe free
Logo
ResearchAudio

Archive

More capable coding agents are better at hiding distributed attacks, not worse.

Jul 3, 2026

And the strongest 4-monitor ensemble the paper could build still lets 47% of gradual attacks through.

Read More

Meta FAIR built a self-grading data factory: the same model writes the questions, judges the answers, and accepts the examples.

Jul 2, 2026

Autodata says agents replace data scientists. The paper uses three different AI models in a closed loop, and the same model grades the test it designed the training data for. We read the math. Here is what the headline is hiding.

Read More

Anthropic shipped Claude Sonnet 5. The system card says it is worse at cybersecurity than 4.6.

Jul 1, 2026

The launch pricing is $2 / $10 per million tokens, and ends August 31, 2026. Then the price doubles.

Read More

Meta’s Brain2Qwerty v2: more data, real-time, and still a long way from a patient

Jun 30, 2026

Meta fine-tuned an LLM on noisy neural data. No surgical implant.

Read More

OpenAI Designed a Chip in Nine Months

Jun 26, 2026

Built for language model inference, with the lab's own models accelerating the work.

Read More

The best AI memory agent fails 19.9% of the time at the easiest job.

Jun 25, 2026

The new GateMem benchmark scores memory agents on a multiplicative scale. The best LLM+memory system in the test still leaks or forgets 1 in 5 times.

Read More

Sakana's Fugu routes to GPT-5.5, Claude Opus, and Gemini. It beats all three.

Jun 24, 2026

A 7B-parameter model from a Tokyo lab is now the strongest publicly accessible coding agent and it does not generate a single token.

Read More

OpenAI Found the Bugs. Humans Read Every One.

Jun 23, 2026

Inside the Daybreak pipeline: judging agents, false positives, and a 23-year-old flaw.

Read More

Stop prompting Claude. Start writing loops.

Jun 22, 2026

The person prompting Claude Code is the bottleneck. The loop is no longer optional.

Read More

DiffusionGemma writes a paragraph the way Stable Diffusion paints a face

Jun 19, 2026

Google DeepMind open-weighted a 26B model that skips token-by-token generation entirely. The speed claim is real. The benchmark gaps are also real.

Read More
Load more

ResearchAudio

Evidence-led AI briefings and free decision tools for engineers, researchers, and product builders.