3 min read

Claude Sonnet 5.5: One Million Tokens for Less

Claude Sonnet 5.5AnthropicLLM

Anthropic released Claude Sonnet 5.5 as the primary Sonnet model for its API. It offers a one-million-token context window, up to 128,000 output tokens and pricing from $2 per million input tokens. Its significance lies in combining long context, adaptive thinking and claimed savings of up to 30% versus Sonnet 5.

What Anthropic actually released

What stands out here is not the version number, but the set of practical capabilities Anthropic has combined in Sonnet 5.5. On the official Claude Sonnet 5.5 page and in the Claude Platform model overview, the company positions it as the primary Sonnet for API use and a balance of speed and intelligence. This is no longer an experimental showcase; it is a working entry point into the model family.

The context window is one million tokens, while maximum output reaches 128,000 tokens. The model supports adaptive thinking, with the default effort level set to high. In the API, it uses the identifier claude-sonnet-5-5, making migration explicit rather than silently replacing a model under an existing name.

At launch, base pricing was $2 per million input tokens and $10 per million output tokens. Batch API provides a 50% discount in both directions. For prompt caching, reads cost $0.20 per million tokens, standard writes cost $2.50, and writes with a one-hour TTL cost $4.

Anthropic says a typical workload can cost up to 30% less than Sonnet 5, even though the models share the same base rates. That means the company ties the main benefit not to a new list price, but to greater operating and billing efficiency.

Where the release genuinely changes the equation

For API teams, this is a meaningful upgrade rather than a simple model-catalog reshuffle. A million-token context is relevant for agents, large codebases and long document chains, while 128,000 output tokens remove some constraints on generating substantial artifacts. Still, a large limit alone does not guarantee that the model uses the entire context equally well.

I would test three things first:

  • how accurately it retains facts from different parts of a long prompt;
  • what happens to latency and token consumption with effort set to high;
  • whether the claimed savings persist in real agent loops using tools and caching.

The strongest fit is for workloads where context capacity and prompt reuse matter more than the lowest possible cost per request. The familiar risk remains: a capable model running at high effort can turn input savings into a large bill for extended reasoning and output. The real question with Sonnet 5.5 is not the one-million-token number, but how predictably that capacity performs under load.

We previously covered how Claude Sonnet can support parallel code-review agents and expose race conditions in pull requests. That practical use case helps clarify where a new Sonnet release may affect engineering workflows.