Logo
Search
Login
Sign Up
Logo

AI Models

Astra and Fable 5.1 Share a Rate Card, Not a Bill

Sep 7, 2026

Same 10 in and 50 out per million tokens. One burns 3.3x the tokens per index task and bills 2.4x. One cache line sits 4x apart.

Read More

24GB Is Not 24GB: The Local LLM VRAM Worksheet

Aug 8, 2026

Model weights are the floor. Context, concurrency, runtime, and usable memory decide whether the deployment fits.

Read More

How Much VRAM Do 7B and 13B Models Need?

Aug 7, 2026

Six weight-and-cache scenarios show why parameter count alone cannot answer the GPU question.

Read More

gpt-oss 20B vs 120B Hardware Requirements: VRAM, RAM, and Context Math

Aug 6, 2026

The official checkpoints fit in 16 GB and 80 GB. Context, concurrency, and runtime reserve decide whether your deployment does.

Read More

LLM API Cost Calculator

Aug 4, 2026

A vendor-neutral formula for turning token assumptions into a monthly API estimate.

Read More

Kimi K3 Ships Open, With a Revenue Catch

Jul 28, 2026

A 2.5x efficiency claim, a 96-shard download, and fine print.

Read More

Google Deleted the Encoder Out of Gemma 4 12B

Jun 5, 2026

Vision is one matrix multiply. Audio drops the encoder entirely. It still runs on a 16GB laptop.

Read More

ResearchAudio

Evidence-led AI briefings and free decision tools for engineers, researchers, and product builders.