Skip to main content
Trends & Strategy9 min read

Both Labs Cut Prices. Your Bill May Not Drop.

September 22, 2026By Ajan Kanagalingam

Anthropic and OpenAI both shipped cheaper models today, hours apart. Anthropic says Claude Opus 5.5 matches its Fable 5.1 flagship at around 40% less to run than Opus 5. OpenAI says GPT-6 Sol and Luna come in 50% below the promotion currently running on GPT-5.6. Both companies say they are passing efficiency savings to customers. Whether that reaches your invoice depends on three things neither announcement leads with.

What was actually announced

Anthropic's Opus 5.5 is described as its most efficient model, performing at the level of Fable 5.1 while costing roughly 40% less to run than Opus 5, which launched in July. It is the first of a family, with Sonnet 5.5 and Haiku 5.5 expected over the coming weeks. Anthropic says it is passing the savings on through price cuts and rate limit increases.

OpenAI's GPT-6 Sol and Luna are offshoots of the GPT-6 Astra flagship released earlier this month, positioned as versions of Astra for everyday work. OpenAI attributes the reductions to improvements in caching and inference. Fortune reported that OpenAI's pricing comes out below Anthropic's, based on each company's own announcement rather than on independent testing.

Two flagship releases on one day, both cheaper, is genuinely good news for buyers. The structure underneath is where the money actually moves.

Three things that decide your invoice

FactorWhat to check
The baselineIs the cut measured against list, or against a promotion
The meterWhether long-context requests bill at a higher rate
The channelWhat the same model costs through your reseller or marketplace

The baseline. OpenAI's 50% is against a promotion currently running on GPT-5.6, and published rates list that Sol promotion as available at least through November 21, 2026. A promotional price has an end date by definition. Budgeting a year against a rate with a stated expiry is how a projected saving becomes a surprise increase in December.

The meter. This is the one almost nobody checks. Published OpenAI pricing carries a separate long-context meter where the current flagship tier moves from $5 to $10 per million input tokens and from $30 to $45 on output, with cached input doubling alongside the input rate. If your workload involves feeding in large documents, you may be paying the second set of numbers while planning against the first.

The channel. Reporting this week indicated Moonshot's Kimi K3 arriving on Amazon Bedrock at roughly 1.76 times OpenRouter pricing for the same model. One example rather than a rule, and enough to establish that where you buy is now a pricing variable alongside what you buy.

Cheaper is not the only direction

The same week also produced reporting that xAI's Grok 4.7 outperforms Claude on enterprise analysis while roughly doubling the bill. That is the pattern worth internalising: the spread between the cheapest and most expensive way to get a task done is widening, not collapsing toward a single falling price.

Which means the useful question stopped being which model is cheapest. It is which model is cheapest for the specific task, at the context length you actually send, through the channel you actually buy from. Three variables, and a headline rate answers none of them. We made the related argument about top-tier pricing in the best model now costs more.

What to do this week

Find your real cost per task. Not per million tokens. Take one workload you run regularly, look at an actual bill, and divide. That number is comparable across providers in a way that published rates are not.

Check whether you are on a promotional rate, and when it ends. Put the date in a calendar. This takes four minutes and prevents the most common AI budgeting surprise.

Check your context lengths against the meter threshold. If you routinely send long documents, you may already be on the higher tier. Trimming what you send is often cheaper than switching models, and it is the same discipline as cutting instructions that no longer earn their place.

Price your channel. If you buy through a cloud marketplace or reseller, compare against direct. Marketplaces can be worth a premium for billing consolidation, data agreements and procurement simplicity. That should be a decision, not an accident.

Keep switching cheap. With two flagship releases in a day, the model you standardise on today is unlikely to be the right one in six months. Abstracting your integrations so a swap is a configuration change rather than a project is the structural version of this advice, set out in hedging against model lock-in. The spend-control side is in turning on your spend controls.

Frequently Asked Questions

What did Anthropic and OpenAI announce on September 22?

Anthropic released Claude Opus 5.5, which it says performs at the level of its Fable 5.1 flagship while costing around 40% less to run than Opus 5 from July, and described it as the first of a family with Sonnet 5.5 and Haiku 5.5 expected over the coming weeks. OpenAI released GPT-6 Sol and GPT-6 Luna, offshoots of its GPT-6 Astra flagship, positioned for everyday work, with API pricing it describes as 50% below the promotion currently running on GPT-5.6.

Are the two price cuts comparable?

Not directly, and the difference matters. Anthropic’s roughly 40% is measured against Opus 5, a previous list price. OpenAI’s 50% is measured against a promotional rate currently running on GPT-5.6, and a discount against a promotion is a different claim from a discount against list. Fortune reported that OpenAI’s pricing comes out lower than Anthropic’s based on each company’s announcement, which is a comparison of two announcements rather than an independent analysis.

Why might a price cut not reduce my bill?

Three reasons. Promotional rates carry end dates, so the number you budgeted against may revert. Long-context requests are billed on a separate and higher meter at some providers, so the headline rate applies only to short requests. And the same model costs different amounts depending on where you buy it, since resellers and cloud marketplaces add margin. Your invoice is decided by all three, not by the announcement.

What is long-context pricing?

Several providers publish one rate for ordinary requests and a higher one once a request exceeds a context threshold. Published OpenAI rates show the current flagship tier moving from $5 to $10 per million input tokens and $30 to $45 output on the long-context meter, with cached input doubling alongside the input rate. If your workload sends large documents, you may be paying the second number while budgeting against the first.

Does buying through a cloud marketplace cost more?

Often, and it is worth checking rather than assuming. Reporting this week indicated Moonshot’s Kimi K3 arriving on Amazon Bedrock at roughly 1.76 times OpenRouter pricing for the same model. That is one example and not a general rule, but it establishes that channel is now a pricing variable. Marketplaces can be worth a premium for billing consolidation, data agreements and procurement simplicity, which is a decision to make deliberately rather than by default.

Know your cost per task, not per token

We calculate what your AI workloads actually cost across providers and channels, find the promotional rates about to expire, and keep switching cheap.

Related Articles

Trends & Strategy

Your Website May Soon Need a Version for Agents

August 26, 2026Read more →
Trends & Strategy

Why Your Next Laptop Costs More: AI Is Hitting Hardware Prices

June 25, 2026Read more →
Trends & Strategy

AI Is Getting More Reliable, Not Just More Capable

June 19, 2026Read more →
AK
Ajan Kanagalingam
Founder & ChatGPT Consultant, ChatGPT.ca

Ajan leads the ChatGPT.ca team: 200+ custom GPT builds and automation projects for 50+ businesses across 20+ industries. Based in Markham, Ontario. PIPEDA-compliant solutions.

Stay ahead of AI in Canada

Weekly case studies, new tools, and ROI playbooks for Canadian SMEs. One email, zero spam.