Skip to main content
Enterprise AI•9 min read

The Mid-Tier Model Now Matches the Flagship

September 29, 2026•By Ajan Kanagalingam

Anthropic released Claude Sonnet 5.5 on September 28 and reports it landing two points below Opus 5.5 on a test of real-world work across occupations, at $2 per million input tokens against the flagship's $4. The headline everyone will write is that the cheap model caught up. The more useful detail is buried in the same announcement, and it is a setting most businesses have never touched.

What Anthropic published

Sonnet 5.5 is the second model in the Claude 5.5 family, described as running 30% or more faster and costing up to 30% less for most work than Sonnet 5. Pricing is unchanged at $2 per million input tokens, $10 per million output and $0.20 per million cache reads. The saving comes from needing fewer tokens to do the same work rather than from a rate cut, which is a different mechanism and a more durable one.

On capability, the company reports 70.6% on Terminal-Bench 4.0, an agentic coding evaluation, against Sonnet 5's 10.3%. It reports scoring two points below Opus 5.5 on GDPval-AA, and says that on CursorBench its best score lands within about two points of Opus 5.5.

Every number here is Anthropic's, testing its own models. That is the most favourable arrangement available to a vendor and it does not make the figures wrong. It does mean the comparison to a competitor in the same announcement, where Anthropic says Sonnet 5.5 at High effort matches GPT-6 Sol's best score on FrontierCode for about a fifth of the cost per task, deserves the heaviest discount of anything published.

One figure deserves a second look rather than a headline. A jump from 10.3% to 70.6% on Terminal-Bench is enormous, and a starting point of 10.3% says the previous model was close to useless on that particular evaluation. That is a statement about Sonnet 5 on one benchmark as much as about Sonnet 5.5 in general, and the pattern of large jumps on saturating evaluations is why we keep returning to why benchmarks mislead.

The dial nobody adjusts

Reasoning effort controls how much computation a model spends on a request, typically from low through medium and high to max. It is invisible in a chat window, it is set differently by different products, and it changes your bill by an order of magnitude.

Anthropic's reported comparisonCost per task
Low or Medium effort, beating Sonnet 5's best scoreAbout a tenth
High effort on FrontierCode, 10 points above Sonnet 5 at the same settingAbout one fifteenth
High effort, matching GPT-6 Sol's best on FrontierCodeAbout one fifth

The operational detail in that announcement matters more than any of the percentages. Anthropic notes that Medium is the default in the Claude apps while High is the default on the Claude Platform. Same model, same task, different default, materially different cost, depending only on where you called it from.

Most businesses have never looked at this. It sits alongside the long-context meter and channel markups we covered in the fine print under the price cuts as a place money leaves without any decision being made.

The default that is costing you

Most small businesses picked a model once, chose the best one they could afford, and never revisited it. That was reasonable when the tiers were clearly separated by capability.

It is expensive now, because the gap between tiers on ordinary work has narrowed while the price gap has not. Summarising a call, drafting a follow-up, extracting four fields from an invoice: none of that needs a flagship at maximum effort, and most businesses are running all of it through whatever they set up in 2025.

The work that still justifies the top tier is real and narrow. Long multi-step reasoning, ambiguous inputs where being wrong is expensive, anything where you cannot easily check the answer. Everything else is a candidate for a cheaper tier or a lower setting.

An afternoon that settles it

Pick one workload you run regularly. Something with volume, where you know what a good answer looks like.

Run twenty real instances across two tiers and two effort settings. Four combinations, twenty cases each. Not invented tests, actual work from last week.

Score the output and read the cost off the bill. Both numbers, side by side. Published rates will not tell you this because token consumption varies with the setting, which is the whole point.

Route by task, not by policy. The answer is usually a mix rather than a single choice, and a mix is only slightly more work to operate. Keeping that switch cheap is the argument in hedging against model lock-in, and it pays off here rather than only in an outage.

Do it once a quarter. Two flagship releases landed in a single day last week and a mid-tier model closed most of the gap to a flagship this week, so whatever you concluded in June is unlikely to still be right.

Frequently Asked Questions

What did Anthropic announce?

Claude Sonnet 5.5, released September 28, 2026 as the second model in the Claude 5.5 family. Anthropic describes it as running 30% or more faster and costing up to 30% less for most work than Sonnet 5, positioned as a faster, lower-cost complement to Opus 5.5. Pricing is unchanged from Sonnet 5 at $2 per million input tokens and $10 per million output, with the saving coming from needing fewer tokens per task rather than from a rate cut.

Is the mid-tier model really as good as the flagship?

On some measures, per the vendor. Anthropic reports Sonnet 5.5 scoring two points below Opus 5.5 on GDPval-AA, a test of real-world work across occupations, and says that on several evaluations at Max effort it performs comparably to Opus 5.5. On CursorBench its best score is within about two points. These are Anthropic’s own figures comparing its own models, which is the most favourable possible testing arrangement, so treat them as a direction rather than a settled result.

What is a reasoning effort setting?

A dial that controls how much computation a model spends on a request, typically running from low through medium and high to max. It is the least-used cost lever in most businesses because it is not visible in a chat window and defaults differ by product. Anthropic notes that Medium is the default in the Claude apps while High is the default on the Claude Platform, so the same model can cost you very different amounts depending on where you are calling it from.

How much can the effort setting save?

A great deal, on Anthropic’s figures. It reports that at Low or Medium effort, Sonnet 5.5 beats Sonnet 5’s best score for about a tenth of the cost per task, and that at High effort on the FrontierCode evaluation it scores 10 points higher than Sonnet 5 at the same setting for roughly one fifteenth of the cost per task. The general principle holds across providers: routine steps rarely need the highest setting, and turning it down where it does not matter is a cost control nobody has to approve.

Should we switch to a cheaper model tier?

Test rather than switch. Take one workload you run regularly, run it at two or three effort settings and on two tiers, and compare both quality and the actual cost on your bill. Most businesses find some tasks genuinely need the top tier and most do not, and the mix matters more than the choice. That test takes an afternoon and is worth more than any published benchmark.

Check the dial, not just the tier

We test your real workloads across model tiers and effort settings, report quality against actual billed cost, and route each task to the cheapest thing that does it properly.

Related Articles

Enterprise AI

Your AI Agent Can Now Log In as You

September 17, 2026Read more →
Enterprise AI

A 35B AI Model on a Phone, With an Asterisk

September 13, 2026Read more →
Enterprise AI

OpenAI’s New Model Is Built to Use Your Computer

September 4, 2026Read more →
AK
Ajan Kanagalingam
Founder & ChatGPT Consultant, ChatGPT.ca

Ajan leads the ChatGPT.ca team: 200+ custom GPT builds and automation projects for 50+ businesses across 20+ industries. Based in Markham, Ontario. PIPEDA-compliant solutions.

Stay ahead of AI in Canada

Weekly case studies, new tools, and ROI playbooks for Canadian SMEs. One email, zero spam.