AI Models

Claude Opus 5.5 nerf claims, checked: how to tell a real change from a bad day

Has Opus 5.5 got worse since launch? What can be checked, what cannot, and a three-step test to run on a question of your own.

Multi Chats Team
October 9, 2026 · 8 min read

"To state it plainly: We never reduce model quality due to demand, time of day, or server load." Anthropic wrote that in a postmortem published on 17 September 2025. Before you file it under corporate denial, read the two paragraphs above it. The same document admits that between August and early September that year, "three infrastructure bugs intermittently degraded Claude's response quality", and that the first complaints were "difficult to distinguish from normal variation in user feedback". In one breath, the company told the people who felt Claude getting worse that they were right, and that nobody had turned the dial on purpose.

A year later the argument is back with a new model attached. Claude Opus 5.5 launched on 22 September 2026. A week later, on 29 September, a thread on r/ClaudeAI that went on to draw more than 400 comments asked whether it had entered its "nerfed" phase. Here is what can be checked and what cannot, as of 2 October 2026.

The tracker everyone cited has not ruled yet

That 29 September thread was an update on LiveNerf, an independent, self-funded project that has been asking Opus 5.5 the same frozen set of questions once a day since 24 September. The comments did not wait for it. An automatically generated summary at the top of the thread read the mood as "a resounding YES, Opus 5.5 has been nerfed." The author's own post said something else: "we won't have that data until day 20 after the model release."

Two days earlier, a company called BridgeMind had posted its own retest on X with the verdict "No nerf detected." One week, then, produced a yes, a no and a "come back later", and the most careful of the three was the one admitting it did not know.

LiveNerf's project page is unusually frank about its limits. On 2 October it had 9 of its 30 daily runs in the bag, and its first possible verdict falls around 24 October. The questions are 78 exam-style problems from science, maths and general academic test banks, picked because Opus 5.5 gets them right only some of the time, which is where a drop would show. Its sanity checks are the most useful numbers in this whole debate.

Lowering the effort setting from high to medium cut the model's output by about a quarter and its score by 4.2 points, give or take 3.9. Splitting identical runs into two halves, with nothing changed at all, produced a gap of 6.4 points, give or take 3.6. And swapping in the older Opus 5 could not be told apart from Opus 5.5 with the samples it had. Sit with that: on a panel built for exactly this job, chance moved the score more than a whole effort level did, and a different model slipped through unnoticed. Its own framing of the question is still the fairest sentence on the subject:

It could also mean nothing happened and people are pattern-matching on noise.

None of this proves Opus 5.5 is unchanged. It does put three disappointing answers on a Tuesday evening in perspective.

What can actually make an answer worse?

Quite a lot, and most of it is checkable.

Start with effort, the setting that decides how long the model thinks before answering. Anthropic's developer documentation says an API request that does not name a level runs Opus 5.5 at medium, where the same request on Opus 5 ran at high. Claude's own app tucks an Effort option into the model menu next to the send button, and Anthropic's help article says each model's recommended level is marked "Default" there. If you changed nothing when 5.5 arrived, check that menu before blaming anyone. Thinking itself can no longer be switched off on Opus 5.5, so the level is the only dial left.

Then there is the switch you may not have noticed. According to Claude's help centre, Opus 5.5 runs safety checks on every request, including on files, memories and search results it reads, and a narrow set of flagged requests is rerun on a less capable model, Opus 5 or Opus 4.8. You get a notice and the reply is labelled, but "the model picker stays on the less capable model for the rest of the conversation." Dual-use virology and toxicology questions can trigger it; everyday health questions, Anthropic says, are fine. If one long chat went downhill after a single odd question, check which model the later replies are labelled with.

Long chats are the next suspect. The model rereads everything above your question each time, and answers tend to get vaguer as that pile grows. We have covered why long chats degrade before; for this test, the same question in a fresh chat is fair and the same question at message eighty is not.

Then the instructions you never see. Claude's apps send Opus 5.5 a long set of instructions with every conversation, which Anthropic publishes and says it updates "periodically". On 2 October 2026 the Opus 5.5 page showed one version, dated 22 September, but the overview says newer models get one entry each and does not say how a later edit would be recorded, so the page cannot settle the argument. Every chat app adds instructions of its own, ours included.

That leaves the theories nobody outside can test, the ones LiveNerf lists: "quantization, a smaller model behind the same name, lower effort, or routing changes." Anthropic's model documentation answers part of that: since the 4.6 generation a model ID "maps to a single, fixed model snapshot", and Anthropic "does not update the weights or configuration of an existing model ID". The same page concedes the rest: the machinery around the model, "the request router, safety classifiers, and sampling logic", can change, and occasionally "infrastructure updates produce minor differences in observable behavior". From the outside you cannot tell which kind of week you are having, and that includes the people most certain on Reddit.

One question, checked three ways

Pick a question with an answer you can mark. Here is ours:

"Four of us rented a cottage for five nights. Ana paid the €900 booking and Sam paid the €60 cleaning fee. Sam stayed only three nights. Split the rent by nights stayed and the cleaning fee equally. Who owes whom?"

The right answer: Ben and Cleo each owe Ana €265, and Sam owes Ana €105 (any set of payments that nets out the same also counts). The rent works out at €50 per person per night, the cleaning fee at €15 each, and Sam's €60 is already paid. Slips to watch for: splitting the cleaning fee by nights, or forgetting Sam already paid it.

Check one is for chance. In a new chat, with the same model and effort, ask three times. Mark each answer right or wrong. Two right and one wrong tells you the question sits on the edge for this model, and last night's bad answer may well have been just that.

Check two is for effort and length. Ask again in a fresh chat one effort level higher, then paste the same question at the bottom of the long chat where things went wrong. If the fresh chat passes and the long one fails, the length is the likelier culprit.

Check three is a second opinion. Run it on a different model. If both fail, suspect the wording of your question. If only Opus fails, save the prompt, note the date and run the same three checks in a week. That gives you the thing LiveNerf exists to build: a baseline of your own.

Doing the checks in MultiChats

On a paid MultiChats plan, the effort level is yours to set, and Opus 5.5 offers Auto, Low, Medium and High there. On the website the app chooses the effort for each message by default; turn off Intent Detection in Settings, under Tool Preferences, and the effort button comes back. The phone apps always show it. Each reply is labelled with the model you chose and, whenever a level was set by you or by the app, the effort, so "was that medium or high?" has an answer on screen. More on the setting in our reasoning effort explainer.

On the website, the More menu under a reply also shows its token count, including reasoning tokens. LiveNerf treats that number as its early warning, since a model thinking less often shows up there before its scores move. For check three, Try Again opens a new chat from that point on any model you pick, and Re-generate replaces the reply where it stands. One caveat of our own: when Anthropic's safety checks flag a request, we let it be rerun on Opus 4.8 rather than fail, and the label under that reply still reads Opus 5.5.

Our verdict, for now

As of 2 October 2026, nobody has shown that Opus 5.5 got worse, and the open, day-by-day test that started at launch cannot say so yet. Equally, the 2025 postmortem shows that quality can drop without anyone deciding it should, and that early reports sound like noise. Both camps are right to keep receipts.

So treat a bad answer as one data point. Run the three checks, save the prompt and try it again in a week; if the drop survives that, you hold more evidence than most of that thread did. LiveNerf's first possible verdict is due around 24 October, and Anthropic's sentence about never lowering quality under load will be easier to believe, or to doubt, once it lands.