Skip to content

Claude Sonnet 5.5 for code review: More catches than Sonnet 5, in half the time

by
Hendrik Krack

Hendrik Krack

September 28, 2026

12 min read

CodeRabbit model evaluation cover with white “Claude Sonnet 5.5” and “More catches than 5, in half the time” lettering over an orange grid on a dark background.

Anthropic has released Claude Sonnet 5.5, the second model in the Claude 5.5 family, at the same price as Sonnet 5 and with a promise of 30% faster output and up to 30% lower cost per task. We ran it through CodeRabbit’s review pipeline the same way we tested Opus 5.5 earlier this month. The question for most teams is not whether Sonnet 5.5 beats Opus. It is whether it fixes the one thing that held Sonnet 5 back as a reviewer: clean comments, but too many bugs left uncaught.

On our 13 hardest known-bug cases, it does. Sonnet 5.5 caught 6 of 13 issues through actionable comments, against 4 for Sonnet 5, at 41.2% actionable precision (40.0% for Sonnet 5), with 17 reported comments to Sonnet 5’s 15. It did this in about half the wall-clock time, and at list prices the Claude model calls cost about 60% less per review, twice the saving Anthropic advertises. Four of its catches were bugs Sonnet 5 waved through; it missed two that Sonnet 5 caught, so this is a “different misses” story as much as a “more catches” one.

Thirteen cases is a small set, and we say so throughout. A second, larger run on 44 open-source pull requests confirms the speed and comment-volume side of the story. Together they make this the first Sonnet release we would consider for the main review pass rather than only for comment quality.

How the Sonnet line has evolved in our code-review evaluations, from Sonnet 4 to Sonnet 5.5. Qualitative direction only.

We have put every Sonnet on the bench since Sonnet 4, so this release fits a pattern. Sonnet 4.5 added reasoning depth, with a hedging habit. Sonnet 4.6 became the highest-coverage Sonnet we had measured, catching about 63% of known issues, but at 29% precision it commented on everything. Sonnet 5 traded that coverage for precision: cleaner comments, but coverage fell to about 50% and nitpicks multiplied. Sonnet 5.5 is the first release in the line to move coverage back up without giving the precision back.

What’s new in Claude Sonnet 5.5

Anthropic positions Sonnet 5.5 as the faster, lower-cost complement to Opus 5.5: strongest at well-scoped everyday work such as fixing bugs and producing documents, slides, and spreadsheets, while Opus 5.5 stays the model for open-ended work that needs sustained judgment. Five things from the launch announcement matter for teams building review and coding workflows:

  • Same price, fewer tokens. Sonnet 5.5 costs $2 per million input tokens, $10 per million output tokens, $0.20 per million cache reads, and $2.50 per million cache writes, unchanged from Sonnet 5 and half of Opus 5.5’s $4 and $20. Anthropic says it typically needs far fewer tokens for the same work, up to 30% less per task, and generates output 30%+ faster. Our review runs show a larger gap than that; see the cost section below.

  • Large gains on agentic coding benchmarks. On Anthropic’s launch table (below), Sonnet 5.5 jumps from 10.3% to 70.6% on Terminal-Bench 4.0, ahead of Opus 5.5, and lands within about three points of Opus 5.5 on most other rows; AA-Briefcase is the widest gap, at 11 points. Our code-review results follow that shape, except that on our hardest cases the gap to Opus 5.5 is wider.

  • Thinking is on by default, and effort is the dial. Adaptive thinking is the default; effort (low, medium, high, xhigh, max) controls cost and depth, with Medium the default in the Claude apps and High on the Claude Platform. Thinking can still be turned off at the lower effort levels, and teams that ran Sonnet 5 with thinking off need to move to the new between_tools setting when they migrate. Our thinking-off run below uses that path.

  • API behavior carried over from Opus 5.5. Forced tool use is retired in favor of structured outputs, and tool definitions can be added or changed mid-conversation without invalidating the prompt cache or earlier thinking blocks. Anthropic’s guidance also notes the model follows instructions literally, so phrasing like “minimize tool calls” is obeyed to the letter, and that at low effort it can report a code change as done without running a check unless the prompt asks for one.

  • Safeguards and deployment. Sonnet 5.5’s cyber capabilities are comparable to Opus 5’s, so it is the first Sonnet to ship with Opus-style cyber safeguards: routine bug fixing is unaffected, but higher-risk security requests fall back to Sonnet 5. Biology safeguards are unchanged from Sonnet 5.

Anthropic’s launch benchmarksSonnet 5.5Sonnet 5Opus 5.5
Terminal-Bench 4.0 (agentic coding)70.6%10.3%66.4%
FrontierCode 1.1 Main (agentic coding)52.1% at Xhigh · 46.2% at Max42.4%54.4%
GDPval-AA v2.1 (knowledge work)184414491846
AA-Briefcase v1.1 (knowledge work)181113591822
Humanity’s Last Exam, with tools64.5%54.9%67.7%
OSWorld 2.1, partial (computer use)80.1%57.0%81.8%
Chartography, no tools (chart recognition)61.6%15.6%64.4%

For reviewers, the practical change is what we measured below: Sonnet 5.5 spends far fewer tokens per review call than Sonnet 5 and still finds more.

What we tested

We used two benchmarks. Signal is the same 13 harder cases we used for the Opus 5.5 evaluation: each is a real pull request from an open-source project (Elasticsearch, Puma, vLLM, Cilium, axios, and Next.js) with one verified issue the reviewer should catch. Nine are difficulty 3, three are difficulty 4, and one is difficulty 5. OSS August is our broader set, 85 known issues across 44 open-source pull requests, weighted toward logic errors, API misuse, race conditions, null references, and security issues.

ConfigurationWhat it is
Sonnet 5.5, thinking onLow, medium, and high effort for the trivial, junior, and senior review cohorts respectively, with adaptive thinking left on. Run on both benchmarks.
Sonnet 5.5, thinking offSame effort ladder with thinking disabled, which the model allows at these three effort levels. Signal only.
Sonnet 5Same effort ladder, run the same day as the Sonnet 5.5 configurations. Both benchmarks.

All runs replayed the same recorded file summaries, walkthroughs, and layer grouping from a frozen cassette, so they reviewed identical inputs; the review model and its thinking setting are what changed between them. On Signal, an independent judge scored every comment with three votes against the known issue, and only majority-PASS comments count.

Known issues caught is the share of the 13 cases where at least one regular actionable comment passed (outside-diff and nitpick comments excluded). Actionable precision is passing actionable comments divided by all actionable comments. Reported comments is the post-pipeline count after verification, deduplication, and filtering; someone still has to read each one.

TL;DR

Signal · 13 patternsKnown issues caughtCaught incl. outside-diffActionable precisionReported commentsMean time per review
Sonnet 5.5, thinking on6/13 · 46.2%8/1341.2%175:27
Sonnet 5.5, thinking off5/13 · 38.5%7/1338.5%135:58
Sonnet 54/13 · 30.8%4/1340.0%159:55
Opus 5.5 Standard (from our Opus 5.5 post)8/13 · 61.5%10/1366.7%21—
Opus 5.5 Max (from our Opus 5.5 post)10/13 · 76.9%10/1352.0%25—

Figure 1. Known issues caught through actionable comments on the 13 Signal cases. Squares include findings outside the changed lines. Opus 5.5 figures are from our September evaluation of the same cases.

Sonnet 5.5 with thinking on catches half again as many known issues as Sonnet 5 on this set, at the same precision, with two more comments. Sonnet 5 remains the precise-but-quiet reviewer we described in June, and on this harder set that quietness cost it: it caught the fewest issues of any configuration.

Opus 5.5 caught 8 of 13 (Standard) and 10 of 13 (Max) in the same cases. Sonnet 5.5 does not close that gap, and we come back to it below.

What Sonnet 5.5 means for your code reviews

Same precision, more of it

Sonnet 5.5 and Sonnet 5 landed at almost the same actionable precision, 41.2% versus 40.0%, so the extra catches did not come from posting more speculative comments. Sonnet 5.5 posted 17 reported comments to Sonnet 5’s 15, with fewer nitpicks and no comments labeled critical, where Sonnet 5 labeled two critical and one of those was wrong. The two models write comments of similar quality; Sonnet 5.5 simply lands more of them on the bug.

On 44 real pull requests, it is quieter and twice as fast

The Signal set is small, so we also ran Sonnet 5.5 and Sonnet 5 on our Open-Source benchmark: 85 known issues across 44 pull requests.

OSS August · 44 PRsReported commentsCritical / major / minorNitpicksMean time per reviewMedianTotal for 44 reviews
Sonnet 5.5, thinking on1114 / 50 / 5796:335:444 h 49 min
Sonnet 514614 / 87 / 453013:3113:499 h 55 min

Figure 2. OSS August benchmark, 44 pull requests: time per review and post-pipeline comment volume. Judge scoring pending.

Sonnet 5.5 posted 24% fewer comments than Sonnet 5 with a milder severity mix, and a third of the nitpicks. Sonnet 5 labeled 14 comments critical to Sonnet 5.5’s four, and added 30 nitpicks, the same nitpick-heavy pattern we reported in June. The judged Signal results say Sonnet 5’s extra comments did not translate into more catches; the OSS judge run will tell us whether that holds at scale, and whether Sonnet 5.5’s lower volume costs it any coverage. Until then, read the volume numbers as workload, not quality.

The latency result needs no such caveat. Across 44 reviews, Sonnet 5.5 averaged 6:33 per review against 13:31 for Sonnet 5, and 5:44 against 13:49 on the median review. Sonnet 5 also produced four generations that ran longer than ten minutes; Sonnet 5.5 produced none.

Want to see what this looks like on your own code? Try CodeRabbit on your next PR. It is free to start and takes about two minutes to connect a repo.

Thinking on or off?

We ran Sonnet 5.5 twice on Signal: once with adaptive thinking on across the effort ladder, once with it turned off. Unlike Opus 5.5, which rejects explicit thinking toggles, Sonnet 5.5 allows thinking off at low, medium, and high effort (the setting Anthropic’s migration guide now calls between_tools), so this is a configuration teams can actually deploy.

Sonnet 5.5 · SignalKnown issues caughtActionable precisionReported commentsNitpicksMean time per review
Thinking on6/1341.2%1725:27
Thinking off5/1338.5%1335:58

Thinking on won by one case and 2.7 points of precision, while producing four more comments. The two configurations disagreed in both directions: thinking-on caught the vLLM config-context bug and a streaming tool-call serialization case that thinking-off missed, while thinking-off caught the Elasticsearch terms-enum case through a regular comment that thinking-on only reached outside the diff.

The cost of thinking was smaller than we expected. On the core review calls, thinking-on produced about 5,800 output tokens per call versus 2,900 with thinking off, on identical input, and mean latency per review was 31 seconds shorter, which is within noise but shows the extra tokens did not slow reviews down. The default should be thinking on.

What a review actually costs

Sonnet 5.5 and Sonnet 5 share a price list, so the cost difference between them is entirely a token difference. We priced the Claude model calls in each run at Anthropic’s published rates ($2 input, $10 output, $0.20 cache read, $2.50 cache write per million tokens). The shared smaller models that handle summaries and verification are the same on both sides and are left out.

Claude model calls at list priceSignal (13 reviews)Per reviewOSS August (44 reviews)Per review
Sonnet 5.5, thinking on$6.16$0.47$20.32$0.46
Sonnet 5.5, thinking off$5.37$0.41——
Sonnet 5$15.06$1.16$50.95$1.16

On both benchmarks, Sonnet 5.5’s Claude calls cost about 40% of Sonnet 5’s for the same reviews, a saving of roughly 60% per review. Anthropic’s launch figure is “up to 30% less per task”; a code-review workload, where Sonnet 5 read the same files repeatedly and wrote long deliberations, sits well beyond that. Thinking on added about 15% to the bill over thinking off and bought one more catch and a little more precision. The difference comes from token usage.

Figure 3. Cost of the Claude model calls per review at Anthropic’s list prices, Signal and OSS August.

Figure 4. Average input and output tokens per core review call on Signal (22 calls per run).

Avg per core review callSignal: inputSignal: outputSignal: thinking wordsOSS: inputOSS: outputOSS: thinking words
Sonnet 5.5, thinking on110.7k5.8k46487.3k5.7k523
Sonnet 5.5, thinking off110.7k2.9k0———
Sonnet 5247.5k21.6k2,771191.5k23.8k3,143

Signal has 22 core review calls per run, OSS August 84. Input is gross prompt size (uncached tokens plus cache reads and writes), standardized across providers.

The two benchmarks agree. Per review call, Sonnet 5 reads more than twice as much as Sonnet 5.5, writes about four times as much, and thinks roughly six times as many words, on both sets. On Signal it also caught fewer bugs. Whatever Anthropic changed between the two releases, the new model reaches its conclusions with far less deliberation.

Across the whole pipeline, including the smaller models that handle summaries and verification, Sonnet 5 consumed 27% more total tokens than Sonnet 5.5 on Signal and 49% more on the 44 OSS reviews, and more than twice the output tokens on both.

Latency is where Sonnet 5.5 separates most clearly from its predecessor, and the larger run makes the point more firmly than the small one. On Signal, mean time per full review was 5:27 against 9:55 for Sonnet 5; on the 44 OSS reviews, 6:33 against 13:31. Over both sets, Sonnet 5 needed just over twelve hours of wall-clock time for 57 reviews; Sonnet 5.5 needed six.

Figure 5. Mean and median wall-clock time per full review on Signal.

Sonnet 5.5 against the Sonnet line and Opus 5.5

Our test sets, judges, and pipeline versions change between evaluations, so the numbers below are two different kinds of comparison. The Signal rows are a same-set comparison: identical 13 cases and identical recorded inputs, run within days of each other. The earlier Sonnet rows are historical context from our previous posts and should be read as direction, not as a leaderboard.

ModelEvaluationKnown issues caughtActionable precisionComments and noise
Sonnet 4.5Oct 2025, 25 hard PRsclosed much of the gap to the flagship of the time35% comment-levelhedged in a third of its comments
Sonnet 4.6Jun 2026, our standard benchmark of the time (Sonnet 5 review)about 63%about 29%the noisiest Sonnet we measured
Sonnet 5Jun 2026, same set as 4.6about 50–51%38–40%nitpick-heavy
Sonnet 5Signal, this post4/13 · 30.8%40.0%15 comments, 3 nitpicks, 2 critical
Sonnet 5.5, thinking onSignal, this post6/13 · 46.2%41.2%17 comments, 2 nitpicks, 0 critical
Opus 5.5 StandardSignal, Sep 20268/13 · 61.5%66.7%21 comments
Opus 5.5 MaxSignal, Sep 202610/13 · 76.9%52.0%25 comments
Opus 5.5 StandardOSS August, 80 cases, Sep 202663.8%38.6%127 comments

Three things stand out. First, within the Sonnet line, 5.5 is the first release to improve coverage without losing precision: 4.6 had coverage but not precision, and 5 had precision but not coverage. Second, on the same 13 Signal cases, Opus 5.5 caught 8 to 10 issues where Sonnet 5.5 caught 6, at higher precision. On the hardest cases the flagship is clearly the stronger reviewer, and Sonnet 5.5 does not close that gap. That is a wider margin than Anthropic’s launch benchmarks show, where Sonnet 5.5 lands within about three points of Opus 5.5 on most rows, but it fits Anthropic’s own framing that Opus 5.5 stays clearly stronger on complex, open-ended work that needs sustained judgment. Hard review cases are exactly that kind of work. Third, Sonnet 5.5 gets its result cheaply: it finishes reviews in about half the time of Sonnet 5, at half of Opus 5.5’s list price and 40% of Sonnet 5’s actual cost per review.

Those are different roles. If missing a bug on a high-risk change is the expensive failure, Opus 5.5 is worth its extra tokens. If you are choosing the model for the review pass that runs on every pull request, fast and with comments developers will actually read, Sonnet 5.5 is the first Sonnet we would put on that list.

How it builds: a side-by-side with Opus 5.5

We also gave Sonnet 5.5 and Opus 5.5 the same long prompt in side-by-side Claude Code sessions: build “Brick Studio,” a brick-model designer with an Alpine Chalet demo model, and record a 45-second showcase video. Sonnet 5.5 finished first, in 29 minutes 27 seconds, against 44 minutes 50 seconds for Opus 5.5, so it was about 1.5× faster. The results were close to identical, with Opus 5.5 a little higher in fidelity. This is one run, not a benchmark. The sessions shared a working folder, so Sonnet 5.5 moved into an isolated subfolder partway through, and neither model received the reference screenshot. It does match the review results above: Opus 5.5 is the stronger model on the hardest work, and Sonnet 5.5 gets close in much less time.

How to evaluate Sonnet 5.5 yourself

Thirteen cases is enough to see direction, not to settle a decision. Before switching:

  • Run it on pull requests from your own repositories with known outcomes, and compare which bugs it catches and misses against the model you run today. Sonnet 5.5 and Sonnet 5 disagreed on six of our 13 cases; our two Sonnet 5.5 configurations disagreed with each other on three.

  • Test thinking on and off. In our runs thinking on caught more, was more precise, and was not slower; it cost four more comments and about twice the output tokens per review call.

  • Measure the complete review, not the model call. Verification agents, retries, and summary models all add tokens and time. Track input, output, cache-read, and cache-write tokens separately.

  • Read the comments that pass and the ones that do not. Several of Sonnet 5.5’s misses were plausible findings on the right file, which is harder to triage than noise.

Our verdict

Sonnet 5.5 fixes the specific weakness that kept Sonnet 5 off our main review path. On our hardest cases it catches half again as many known issues as Sonnet 5 at the same precision, and on 44 broader pull requests it posts a quarter fewer comments, a third of the nitpicks, and finishes in less than half the time, at about 40% of the cost.

It is not an Opus 5.5 replacement, four of our 13 hard cases defeated every Sonnet configuration we ran, and the coverage numbers for the larger set are still to come. But for the review that has to run on every pull request, fast and with comments developers will actually read, Sonnet 5.5 is the most convincing Sonnet we have tested. As our VP of AI, David Loker, put it in Anthropic’s launch announcement, Sonnet 5.5 “shows better judgment than Sonnet 5 across different levels of complexity” while spending far fewer output tokens, and Sonnet 5’s habit of reaching for web search too often is gone. We are moving simple and moderate reviews over now, and more in the coming weeks.

Get started with CodeRabbit: connect your repo, get your first review in minutes. Free to try, no credit card required.


Methodology notes. Signal is a 13-case subset of harder known-bug patterns drawn from real open-source pull requests; OSS August is 85 known issues across 44 open-source pull requests, for which only comment counts, latency, token usage, and cost are reported here because judge scoring was pending. A case counts as caught when at least one regular actionable comment is accepted by a majority of three judge votes. Precision measures the share of actionable comments that pass the judge for the target issue, not developer acceptance. All runs replayed identical upstream summaries from a recorded cassette. Individual judge calls matter at this sample size: on the Puma case, a finding from Sonnet 5 was accepted while a near-identical finding from Sonnet 5.5 was not, and one of Sonnet 5.5’s seven passing comments passed on a two-to-one vote; excluding it gives 35.3% precision. Token figures are averaged across generation calls with input standardized as uncached tokens plus cache reads and writes. Dollar figures are recomputed from each run’s token counts at Anthropic’s published list prices for Sonnet 5.5, which are unchanged from Sonnet 5, and cover the Claude model calls only; the evaluation harness itself used placeholder rates for the pre-release model.

Share

Share on RedditShare on XShare on LinkedIn
CR_Flexibility.

Frequently asked questions

Catch the latest, right in your inbox.

Add us to your feed.

GetStarted in2 clicks.