Open to workNew YorkGet in touch

LLM Evaluation

Livenerf: how it tests whether Claude Opus 5.5 got nerfed, and what it can't see

A pre-registered benchmark tracks Opus 5.5 daily through Claude Code, detecting about 7.5-point drops and effort cuts, but it cannot yet distinguish a swap to Opus 5.

Published
September 30, 2026
Read
7 min
Author
Samir Sengupta
Livenerf daily benchmark tracking Claude Opus 5.5 accuracy and output tokens against a launch-week baseline
Note 074 / 074Daily note · Written from 2 sources

the short version

  • The panel keeps only the 78 of 2,336 screened questions that Opus 5.5 sometimes gets right, because questions it always passes or always fails cannot show a change.
  • In validation, output tokens exposed lowered effort more clearly than accuracy did: effort low cut tokens 62% but accuracy only 8.3 ± 4.5 points.
  • A swap from Opus 5.5 to Opus 5 was not distinguishable at 99% (−3.8 ± 6.3 points), so a null result does not rule out a same-family model swap.
  • The earliest possible verdict is around 2026-10-24, and it will apply to Opus 5.5 as served through Claude Code on a subscription, not to the raw API.

Livenerf is a pre-registered benchmark that runs a fixed 78-question panel against Claude Opus 5.5 once a day through headless Claude Code (claude -p). It compares each question's score with that same question's launch-week baseline, and tracks output tokens per sample as a secondary signal. A regression counts only if a 99% interval excludes zero in two consecutive 10-day windows, the change is at least 3 points, and a claude-opus-5 control arm does not show the same move. At one run a day, it can detect an accuracy change of about 7.5 points per window. In validation, however, it could not distinguish Opus 5.5 from Opus 5.

Opus 5.5 shipped on 2026-09-22. The ninjahawk/livenerf repo on GitHub started its series on 2026-09-24 at 22:10 UTC, about 2.5 days after launch. The author states the motivation plainly: nobody has had a clean day-0 baseline to check against, so nerf arguments end up as "vibes versus vibes."

As of 2026-09-29, 6 of 30 daily runs had been collected, none missed. Each ran the full 90 samples on harness hash 461391b6fce64167 with Claude Code CLI 2.1.280 pinned, and day 5 ran with the budget guard overridden once. Days 1 to 10 form the baseline, and the first possible call lands around 2026-10-24.

What livenerf measures, and why 78 questions

Sampling parameters are gone and thinking cannot be turned off, so Opus 5.5's outputs cannot be made deterministic. The benchmark fixes everything else instead: frozen prompts, a pinned CLI, exact graders and raw logs kept forever. It then measures drift statistically over thousands of samples. It is built on Inspect, the UK AI Security Institute's open-source eval framework, and its statistics follow Anthropic's Adding Error Bars to Evals.

The author screened 2,336 GPQA Diamond, MMLU-Pro, competition-math and AIME 2025–26 questions with 4 samples each. Opus 5.5 got about 93% right on the first try, and 97% of the questions were always right or always wrong. Under a common logit-shift model, a question with pass rate p carries information p(1−p) per sample, so a question the model always gets right cannot show a drop. The 78 questions that were sometimes right became the panel. Per the README, this step also filters out memorized questions.

The author measured the resulting selection bias. Questions chosen for being sometimes right look closer to 50/50 than they are. On fresh samples, their pass rate rose from 54.7% to 62.0%, so the power calculation uses the fresh rates rather than the screening rates.

Lower effort shows in tokens more than accuracy

Before the baseline began, the rig was checked against a known degradation, effort medium against high, and an A/A check confirmed that the error bars are honest. Validation passed its pre-registered criterion. The README reports that lower effort shows up much more clearly in tokens than in accuracy:

  • Effort low: −62% output tokens, −8.3 ± 4.5 points of accuracy.
  • Effort medium: −26% output tokens, −4.2 ± 3.9 points of accuracy.

This is why the output token count per sample is the secondary signal the author says they care most about. In the README's words, if a model quietly starts thinking less, the token count is where it shows up first, often before accuracy moves at all. For teams building their own monitoring, the implication is to log token usage alongside scores on every sample.

How the decision rule filters out noise

The primary metric is the per-item paired score difference against baseline, averaged over items, with clustered standard errors following Miller 2024. Pairing each question with its own baseline removes question difficulty from the comparison. Scoring is exact-match only. An LLM judge is excluded because the judge would drift too.

The control arm is designed to catch harness and platform changes. Every day, claude-opus-5 runs the panel's GPQA questions through the same harness, and if both models move together, the change is attributed to the harness or platform, not to Opus 5.5. Every call is also hermetic: a short frozen system prompt, no tools, no MCP servers, no settings or hooks, no CLAUDE.md or memory, one turn, and a fixed empty working directory. The README gives the reason for pinning the CLI:


a changed harness looks exactly like a changed model.
livenerf README

PREREGISTRATION.md was committed before any series data, so its git timestamp is public. The livenerf.analysis step applies the decision rule mechanically. It tests for change in either direction, because launch week could be the worst week: a new serving stack, capacity strain, launch bugs. The README notes that the 2025 quality incidents turned out to be infrastructure bugs, not deliberate downgrades.

Where the instrument cannot see

The most important limit is in the validation data. Swapping in Opus 5 was not distinguishable from Opus 5.5 at 99%: −3.8 ± 6.3 points of accuracy and −23% tokens. The README says directly that the instrument cannot detect a same-family model swap of that size in a validation's worth of samples. A 10-day window has about 2.5 times as many samples, but that has not been shown to be enough. A "no change" result would therefore not rule out a same-family swap of that size, which is close to the "smaller model behind the same name" scenario the README lists among possible nerfs.

There are three other limits:

  • Answer keys: a report-only audit of the 78 panel questions and 2 later-excluded questions found 8 keys that look wrong and 30 ambiguous questions. Nothing was dropped, and a pre-registered sensitivity analysis reruns the result without them.
  • Serving path: the safety classifier sometimes answers with Opus 5 or refuses biology and some math questions. Those samples are rejected and counted, and the questions they touched are excluded.
  • Scope: the benchmark measures Opus 5.5 as served through Claude Code on a subscription (v0 runs on a Claude Max plan). The README says this is not the same thing as the raw API model. Moving to the API would mean passing --model anthropic/claude-opus-5-5, with no change to any task or scorer.

What Hacker News commenters are pushing back on

Skeptics in the thread made two points. One commenter argued that the perceived drop always arrives around a week after launch, attributed it to hedonic adaptation, and said the many benchmark attempts to show nerfing never persist. Another linked xkcd 882 and noted that running more experiments produces more statistically significant-looking results. The repo's design targets both points: a day-0 baseline, a single pre-registered primary metric, a 99% threshold, a two-consecutive-window requirement, and a commitment to publish null results and improvements.

On why the windows are 10 days, a reply in the thread described it as a tradeoff:


The 10-day window is mainly a tradeoff between sensitivity and detection speed.
Commenter on Hacker News

The same commenter said the window rolls forward daily and that its length may be revised once there is enough longitudinal data. That raises a question the sources do not answer. It is not clear whether repeated daily looks at a rolling window inflate the false-positive rate relative to the nominal 99%.

Other commenters asked whether API or cloud-hosted Claude is affected. One per-token Enterprise user reported never noticing slowdowns or a gradual decline in quality. The benchmark cannot settle that question, because it deliberately measures only the subscription path. Another anecdote, about more frequent permission prompts, also falls outside a benchmark that runs with no tools.

What production monitoring can copy

Several of the methods carry over to any team tracking a hosted model:

  • Pin the client binary and store a copy where the auto-updater cannot reach it.
  • Make calls hermetic.
  • Pass effort explicitly, never relying on the default.
  • Grade by exact match outside the model.
  • Keep raw logs append-only, with CLI version, harness git SHA and task and item hashes recorded.

The budgeting is concrete. The daily run of the whole panel uses about 3.6% of the weekly plan. An attempt is skipped if the weekly meter is at or above 75% or the 5-hour meter is at or above 60%, and it retries every hour until the day's run is in. The main caveat is sensitivity. The detectable change is about 7.5 points per window, while the effort-medium degradation in validation moved accuracy by only 4.2 points, so a regression of that size could go unflagged on accuracy alone.

What to watch before 2026-10-24

The first Results row is expected after day 20, and the first possible verdict around 2026-10-24. Points to track:

  • Whether output tokens per sample shift before accuracy does.
  • Whether the control arm moves alongside Opus 5.5.
  • How the sensitivity analysis changes the result once the 8 suspect keys and 30 ambiguous questions are excluded.
  • Whether 10-day windows turn out to have enough power to catch a change the size of an Opus 5 swap. That question is still open.

Questions this raises

How does livenerf decide if Claude Opus 5.5 got nerfed?

A regression counts only if a 99% interval excludes zero in two consecutive 10-day windows and the change is at least 3 points. The claude-opus-5 control arm must also not show the same move. The rule was pre-registered before any series data was collected.

Can livenerf detect if Anthropic swapped Opus 5.5 for a smaller model?

Not yet shown. In validation, swapping in Opus 5 was not distinguishable from Opus 5.5 at 99%, with -3.8 ± 6.3 points of accuracy and -23% tokens. A 10-day window has about 2.5 times as many samples, but that has not been shown to be enough.

Why does livenerf track output tokens?

In validation, lower effort showed up much more clearly in tokens than in accuracy. Effort low cut output tokens by 62%, while accuracy fell only 8.3 ± 4.5 points. If a model quietly starts thinking less, token count often moves before accuracy does.

These daily notes are drafted by a model I run and operate myself - the same kind of pipeline this site is about - from sources published in the previous 24 hours, and every one lists what it read. The longer essays, the talks and the preprint are mine, written by hand.

More notes

Building something on this?

I ship production LLM, RAG and agentic systems for a living - the infrastructure behind the things these notes are about. Open to roles, contract work and research collaboration.