3 min read

Claude Opus 5.5 explains errors more clearly, but loves TLDR

Claude Opus 5.5Anthropicпромпт-инжиниринг

User feedback indicates that Claude Opus 5.5 explains bugs more clearly than Opus 5 and can feel generous on weekly usage limits. Anthropic lists a one-million-token context window and up to 128,000 output tokens. However, the model’s recurring TLDR summaries remain a noticeable stylistic issue.

The biggest change is not visible only in benchmarks

What stands out here is not another score increase, but something more practical: according to user feedback, Claude Opus 5.5 has become much clearer at explaining errors. Where Opus 5 repeatedly hinted at a bug in vague, almost “cosmic” language without getting to the point, the newer version was able to state the problem plainly.

The official picture also points to a substantial update. Claude Platform documentation and Anthropic’s release notes describe the model as a tool for long-running agentic work in software development and knowledge tasks. It offers a one-million-token context window, a maximum output of 128,000 tokens, and always-on adaptive thinking controlled through the effort setting.

In Anthropic’s published materials, Opus 5.5 scores 66.4% on Terminal-Bench 4.0, compared with 52.3% for Opus 5. On GDPval-AA v2.1, it is listed at 1846 Elo versus 1708. These are not direct language-quality tests, but they align well with the idea that the model handles complex instructions and multi-step workflows more reliably.

At launch, pricing was $4 per million input tokens and $20 per million output tokens. One user also reports not coming close to the weekly limit. That experience should not be generalized, though: usage depends heavily on context length, operating mode, and the number of agentic iterations.

Clarity improved, but stylistic habits remain

For practical work, this is a meaningful upgrade. A clear explanation of a bug is more valuable than an elegant hint about it. When a model identifies the cause of a failure faster and explains it in ordinary language, fewer follow-up questions are needed and important details are less likely to disappear behind confident verbosity.

But there is a familiar stylistic cost. In the discussion, another user calls TLDR one of the most frequent words in recent Opus outputs. That is anecdotal rather than a measured model trait, yet the signal for prompt engineers is obvious: define the response format explicitly, especially when an automatic summary disrupts the intended structure.

The first thing I would test is not the polish of a single answer, but consistency across a long sequence of tasks: does the model stay clear, avoid repeating summaries, and keep details from being buried under TLDR? The key question around Opus 5.5 is no longer whether it can reason, but how consistently and disciplinedly it presents that reasoning.

We previously covered Anthropic’s reversal of hidden Claude query downgrades and what it revealed about transparency in model behavior. The Opus 5.5 debate similarly shows why output quality changes need clear scrutiny rather than blind trust.