Skip to content

Do AI models get worse after launch?

Is ChatGPT or Claude getting worse? What LiveNerf and documented regressions show, and how to check quality changes in your own workflow.

Editorial collage of evaluation layers under an inspection lens, with official Claude and OpenAI logos.
≈ 9:06

A new model solves something your previous assistant struggled with. You start trusting it. A week later, it fails a simple task, then gives another disappointing answer. It is tempting to ask whether the provider has made it worse since launch.

An AI tool can perform worse after release. That alone does not establish that its provider deliberately reduced the model’s capability. Understanding the change means separating the model, the software using it and the conditions of the test.

Devsplainers’ video on whether Claude and GPT get nerfed raises this question through early monitoring of Opus 5.5. We used it as a starting point, then checked documentation, evaluation protocols and acknowledged incidents. Sources checked on October 5, 2026.

What does it mean to “nerf” an AI model?

In games, a nerf reduces a character’s or ability’s power. Applied to AI, the word often combines different complaints: fewer correct answers, shorter replies, more refusals, less reasoning or a different model answering the request.

That makes the debate harder to resolve. If an assistant produces code that no longer passes tests, the problem matters. But its cause could be the application’s instructions, the context it supplies or an agent update. Lost usefulness and the reason for it are separate questions.

A precise claim is easier to investigate: “This configuration used to solve these tasks, and now it solves fewer.” The claim that all models become worse a week after release needs much broader evidence.

The model can stay fixed while the service changes

Weights are the values a model learns during training. Anthropic says each model ID pins a version: weights stay fixed, while routing, safety classifiers and sampling logic can change. Some older convenience aliases follow different rules.

Pinning a model therefore helps without freezing the entire experience. A chat application or coding agent adds instructions, tools and context management. Comparing two Code versions may also measure changes in that software.

The Claude API’s safety-refusal fallback is another documented mechanism. When configured, a request can pass to another model; the response identifies which model served it. This does not mean every request uses a less capable model or that server overload triggers that mechanism. Record the model that answered whenever the service exposes it.

Original LetBrand diagram. Recording these layers helps compare conditions; it does not identify the cause of a failure on its own.

What Claude users report and what LiveNerf can measure

On October 3, 2026, Tony (@EnvolDev) posted a comparison of Eiffel Tower scenes generated with Opus 5.5. He prefers the launch result, describing less detail in the buildings, trees and river while seeing an improvement in the tower itself. He also cautions that one run does not prove a nerf. This individual experience can motivate a test; it does not measure the model’s overall quality.

LiveNerf tracks Opus 5.5 through Code with fixed settings. Its panel contains 78 questions selected because the model sometimes gets them right and sometimes wrong. Answers are scored without an AI judge.

Its pre-registered protocol uses the first ten days as a baseline. It requires a change of at least three points and a 99% interval excluding zero in both subsequent ten-day windows, alongside a control-model check. A bad day does not automatically become a degradation claim.

The evaluation card limits what a result means. It does not cover every task, web chat, long sessions or tool use, nor does it identify the cause of a change. “Not detected” means the test did not find it within its scope and sensitivity, not that nothing changed.

Under the protocol checked on October 5, a complete verdict from both subsequent windows is not yet due. A daily curve or viral thread cannot replace that result. Good performance on this panel would not guarantee that your coding workflow remained unchanged either.

Documented quality regressions do exist

In its September 2025 postmortem, acknowledged three infrastructure bugs that degraded responses: routing problems, output corruption and token selection. The company attributed the incidents to bugs and denied cutting quality because of demand or load. Degradation is documented; the causal explanation is the provider’s account.

also rolled back an April 2025 GPT-4o update that made excessively agreeable. Agreement may feel pleasant while making the tool less useful for honest criticism. An update can worsen a behavior without producing a uniform decline across all capabilities.

Independent monitoring offers another example. Marginlab documented five days of lower results for Claude Code with Opus 4.7 in May 2026. Its table aligns the decline and recovery with CLI versions; its authors suspect an agent issue. That interpretation fits their observations, but it is not confirmation from or proof of an intentional downgrade.

LetBrand chart from Marginlab’s May 29, 2026 table. The table supplies no confidence intervals. Alignment with CLI versions does not establish causation.

The table shows recovery starting on May 27, before Opus 4.8 launched on May 28. That distinction matters when interpreting the decline. Marginlab’s May 29 announcement on X links to its analysis and data; the proposed cause remains the authors’ hypothesis.

These cases make a good argument for listening to user reports and keeping records. They also show why different regressions should not automatically receive the same explanation.

Why is ChatGPT getting worse—or feeling worse?

The question “Is ChatGPT getting worse?” often combines two comparisons: today’s performance versus launch, and the end of a conversation versus its beginning. It also helps to separate , the application, from the GPT model answering the request. The mechanisms described above should not automatically be attributed to .

A long conversation accumulates earlier decisions, corrections and scattered requirements. LLMs Get Lost In Multi-Turn Conversation reported an average 39% performance drop across six generation tasks when instructions were spread over several turns rather than supplied completely at once. This describes that experiment, not every chat.

Before comparing today’s response with launch week, try the task in a fresh conversation containing the full brief. Improvement gives you a clue about context. It does not establish that the provider caused the difference.

Answers also vary. As a mathematical illustration, if each attempt had a 10% failure probability and attempts were independent, three consecutive failures would have a 0.1% probability. A million three-attempt sequences would produce about a thousand such streaks in expectation. These are illustrative assumptions, not or GPT measurements; real tasks and failures may be correlated.

Frustration still deserves attention when a losing streak is possible. Without more data, it cannot distinguish chance, changed settings and a service regression.

How to check your own workflow

You do not need to reproduce a research lab. You need a repeatable comparison and a clear definition of failure. Here is a practical starting point:

  1. Keep 20–50 representative tasks. Use calculations with known answers, extraction tasks with required fields or code changes that must pass tests. This is a working suggestion, not a sample size that guarantees statistical significance.
  2. Fix the conditions. Save the full prompt, input files, requested model, reasoning effort and application version. Use fresh sessions for isolated tests; evaluate long conversations separately if they matter to your work.
  3. Define scoring before running. “The code passes these tests” or “These fields are present” is easier to compare than “The answer seems intelligent.” For writing, use specific criteria and human review without revealing which version wrote each answer where practical.
  4. Repeat and keep outputs. Record dates, correct answers, errors, , duration, refusals and the effective model where available. Do not add harder tasks only to the second batch.
  5. Decide how to respond to a sustained decline. Check settings, report reproducible examples or use an already tested alternative. An operational threshold helps you act; it does not establish a cause or intent on its own.

Token counts are an additional clue, not an intelligence score. Less output can reflect less reasoning, shorter prose or a different task. Read it alongside accuracy and settings. When comparing alternatives, LetBrand’s AI price comparator helps estimate costs; use your own tasks to assess quality.

Fictional example by LetBrand. It defines a task and correct answer, not an observed model test.

For an extraction task like this, check JSON validity and whether the fields match the input separately. Keep the same document and scoring rule in both batches. This example defines a correct answer; it is not an observed or GPT result.

So, do models get worse after launch?

They can become worse at particular tasks, and documented incidents exist. The evidence reviewed here does not establish that every model receives a deliberate post-launch downgrade. Fixed weights do not rule out every change either: the tool you use contains other moving parts.

Keep an example when something fails today that worked yesterday. Check settings and repeat it under comparable conditions. Those records help you ask the provider for an explanation and decide whether the service still works for you.