Logo
ResearchAudio
Search
Latest
Free tools
Community
Login
Subscribe free
Logo
ResearchAudio

AI infrastructure

24GB Is Not 24GB: The Local LLM VRAM Worksheet

Aug 8, 2026

Model weights are the floor. Context, concurrency, runtime, and usable memory decide whether the deployment fits.

Read More

How Much VRAM Do 7B and 13B Models Need?

Aug 7, 2026

Six weight-and-cache scenarios show why parameter count alone cannot answer the GPU question.

Read More

gpt-oss 20B vs 120B Hardware Requirements: VRAM, RAM, and Context Math

Aug 6, 2026

The official checkpoints fit in 16 GB and 80 GB. Context, concurrency, and runtime reserve decide whether your deployment does.

Read More

Voice AI Latency Budget: Fast and Slow Models in Parallel

Aug 4, 2026

Map endpointing, transcription, model branches, TTS, and playout before optimizing the wrong stage.

Read More

Prompt Caching Cost Calculator

Aug 4, 2026

Cache reads are cheap. Cache misses may not be. Calculate the hit rate where caching starts paying for itself.

Read More

Google Deleted the Encoder Out of Gemma 4 12B

Jun 5, 2026

Vision is one matrix multiply. Audio drops the encoder entirely. It still runs on a 16GB laptop.

Read More

PagedAttention: The Memory Revolution in LLM Serving

Jan 20, 2026

Treating the GPU memory like an operating system

Read More

ResearchAudio

Evidence-led AI briefings and free decision tools for engineers, researchers, and product builders.