In partnership with

Join 29,000+ Marketers at the Email Conference Everyone Talks About

What do Molly Ringwald, Dan Levy, Amy Porterfield & a world record have in common? They'll all be at GURU Conference 2026. 100% Free. 100% Virtual.

If you're obsessed with marketing like we're obsessed with marketing, GURU is the must-attend conference of the year. We'll be covering all things email marketing: B2B, B2C, newsletters, deliverability, email design, AI & more.

You can expect to walk away with new email strategies, the very latest digital trends, and how to step up your email performance. But don't worry, we also like to have fun. This year's theme is rom-com, so there will be DJs, meet-cutes, and a cutest pet contest. (Start prepping your dog's headshot now.)

Don't miss out. Join us Nov 12th & 13th for the largest virtual & free email marketing conference, powered by Constant Contact.

ResearchAudio.ioSEPTEMBER 29, 2026
THE BUILDER'S EDITION

CLAUDE SONNET 5.5 / LAUNCH ANALYSIS

The price stayed.
The bill could shrink.

Why Sonnet 5.5 deserves a fresh look at your agent's economics, and what to check before switching.

The annoying part of using a coding agent is often the lap it takes around a simple problem. It reads, searches, calls another tool, changes direction, and eventually produces the patch you wanted ten minutes ago.

That is the lens I would use for Sonnet 5.5.

Anthropic released the model on September 28. It reports 30%+ faster output generation and up to 30% lower cost per task versus Sonnet 5, with the savings coming from fewer tokens. These are Anthropic’s results; your workload needs its own test. Launch announcement.

01 / THE ECONOMICS

The rate card only tells half the story

Both Sonnet 5 and Sonnet 5.5 charge the same base token rates. Pricing comparison.

INPUT / 1M TOKENS
$2
Sonnet 5 and 5.5
OUTPUT / 1M TOKENS
$10
Sonnet 5 and 5.5

Sonnet 5.5 also lists $0.20 per million cache-read tokens, a 1M-token context window, and a 128K-token standard output limit. Model specifications.

An agent’s bill depends on how much work it does along the way. Repeated searches, extra reasoning, unnecessary edits, and failed attempts all have a cost. A model can become cheaper to operate without a discount on any individual token.

For a product team, I would track one number above the rest:

COST PER ACCEPTED RESULT

Total model spend

Tasks that pass your acceptance checks

ResearchAudio’s suggested evaluation metric. Include failed attempts and retries in total spend; track tool costs and human review time separately.

A smaller bill is useful only if enough of the work still passes. If your agent finishes quickly but leaves you repairing its output, the saving can disappear outside the API invoice.

02 / THE RESULTS

A large jump, with a narrow interpretation

Two coding evaluations in Anthropic’s launch table show the scale of the change:

Terminal-Bench 4.0

Sonnet 5
 
10.3%
Sonnet 5.5
 
70.6%

CursorBench 4.0

Sonnet 5
 
34.1%
Sonnet 5.5
 
55.5%

Bars use a 0–100% scale. Values are from Anthropic’s launch table, not a ResearchAudio evaluation. Scores come from specific evaluation setups and should not be read as universal task-success rates.

Those results make Sonnet 5.5 worth testing on real repositories. They do not tell you whether it will handle your build system, respect your conventions, or avoid touching unrelated files.

I would begin with work that has a clear finish line: a reproducible bug, an isolated UI change, a test repair, or a document transformation with a fixed template. Save a separate evaluation bucket for ambiguous work, where deciding what to do is part of the task.

03 / THE CONTROL

Effort is part of the model choice

Anthropic’s guidance starts well-specified agentic tasks at medium effort, with high for harder or longer work. The API default is high. Effort has been recalibrated, so carrying over a Sonnet 5 setting does not preserve the same amount of thinking. Effort guidance.

There is a practical wrinkle: the prompting guide says low effort can skip verification, while low and medium can pause for user input during longer tasks. Prompting guidance.

For an unattended workflow, make completion explicit. Tell the agent what counts as finished, which checks must pass, and when it should stop for help. Then measure whether it actually follows those instructions.

04 / BEFORE YOU SWITCH

The API changes deserve a separate pass

Sonnet 5.5 is available through the Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry. The Claude API model ID is claude-sonnet-5-5. Existing integrations should check the following changes.

01Thinking settings. disabled becomes between_tools for turning off up-front thinking. Use it at high effort or below.
02Tool selection. Forced tool_choice modes any and tool now error. Review auto selection and strict tool use.
03Conversation replay. Replaying thinking blocks after editing earlier history can fail. Keep conversations append-only and review model-switching behavior.
04Computer use and advisors. The older computer-use tool is unsupported on the Claude API and Google Cloud. Some older advisor models are also rejected.
05Streaming progress. Between-tool updates arrive in thinking blocks. Review thinking.display if your interface shows those updates.

Selected checks, with platform-specific details in Anthropic’s full change list.

One behavior change also matters for security workflows: Anthropic says higher-risk cybersecurity requests visibly fall back to Sonnet 5. Routine bug fixing remains supported. Safeguard details.

05 / THE TAKEAWAY

Run a small, boring bake-off

Take 20 recent tasks from your own backlog. Keep the prompts, tools, and acceptance checks fixed. Compare your current setup with Sonnet 5.5 at medium and high effort.

Record accepted results, total token spend, elapsed time, retries, and human fixes. Repeat close results before treating a small sample as a decision.

My read: Sonnet 5.5 is most interesting as a candidate for the work you run repeatedly. If it reaches your quality bar with less waiting and fewer billed tokens, that improvement compounds every time the workflow runs.

The question for your next evaluation: what does one accepted result actually cost?

SOURCE NOTES

Verified September 29, 2026. Launch figures are vendor-reported; the evaluation suggestions and interpretation are ResearchAudio’s analysis.
Launch announcement · Model specifications · API changes · Effort settings · Prompting guide

Deep / ResearchAudio.io