Both Labs Cut Prices. Your Bill May Not Drop.
Anthropic and OpenAI both shipped cheaper models today, hours apart. Anthropic says Claude Opus 5.5 matches its Fable 5.1 flagship at around 40% less to run than Opus 5. OpenAI says GPT-6 Sol and Luna come in 50% below the promotion currently running on GPT-5.6. Both companies say they are passing efficiency savings to customers. Whether that reaches your invoice depends on three things neither announcement leads with.
What was actually announced
Anthropic's Opus 5.5 is described as its most efficient model, performing at the level of Fable 5.1 while costing roughly 40% less to run than Opus 5, which launched in July. It is the first of a family, with Sonnet 5.5 and Haiku 5.5 expected over the coming weeks. Anthropic says it is passing the savings on through price cuts and rate limit increases.
OpenAI's GPT-6 Sol and Luna are offshoots of the GPT-6 Astra flagship released earlier this month, positioned as versions of Astra for everyday work. OpenAI attributes the reductions to improvements in caching and inference. Fortune reported that OpenAI's pricing comes out below Anthropic's, based on each company's own announcement rather than on independent testing.
Two flagship releases on one day, both cheaper, is genuinely good news for buyers. The structure underneath is where the money actually moves.
Three things that decide your invoice
| Factor | What to check |
|---|---|
| The baseline | Is the cut measured against list, or against a promotion |
| The meter | Whether long-context requests bill at a higher rate |
| The channel | What the same model costs through your reseller or marketplace |
The baseline. OpenAI's 50% is against a promotion currently running on GPT-5.6, and published rates list that Sol promotion as available at least through November 21, 2026. A promotional price has an end date by definition. Budgeting a year against a rate with a stated expiry is how a projected saving becomes a surprise increase in December.
The meter. This is the one almost nobody checks. Published OpenAI pricing carries a separate long-context meter where the current flagship tier moves from $5 to $10 per million input tokens and from $30 to $45 on output, with cached input doubling alongside the input rate. If your workload involves feeding in large documents, you may be paying the second set of numbers while planning against the first.
The channel. Reporting this week indicated Moonshot's Kimi K3 arriving on Amazon Bedrock at roughly 1.76 times OpenRouter pricing for the same model. One example rather than a rule, and enough to establish that where you buy is now a pricing variable alongside what you buy.
Cheaper is not the only direction
The same week also produced reporting that xAI's Grok 4.7 outperforms Claude on enterprise analysis while roughly doubling the bill. That is the pattern worth internalising: the spread between the cheapest and most expensive way to get a task done is widening, not collapsing toward a single falling price.
Which means the useful question stopped being which model is cheapest. It is which model is cheapest for the specific task, at the context length you actually send, through the channel you actually buy from. Three variables, and a headline rate answers none of them. We made the related argument about top-tier pricing in the best model now costs more.
What to do this week
Find your real cost per task. Not per million tokens. Take one workload you run regularly, look at an actual bill, and divide. That number is comparable across providers in a way that published rates are not.
Check whether you are on a promotional rate, and when it ends. Put the date in a calendar. This takes four minutes and prevents the most common AI budgeting surprise.
Check your context lengths against the meter threshold. If you routinely send long documents, you may already be on the higher tier. Trimming what you send is often cheaper than switching models, and it is the same discipline as cutting instructions that no longer earn their place.
Price your channel. If you buy through a cloud marketplace or reseller, compare against direct. Marketplaces can be worth a premium for billing consolidation, data agreements and procurement simplicity. That should be a decision, not an accident.
Keep switching cheap. With two flagship releases in a day, the model you standardise on today is unlikely to be the right one in six months. Abstracting your integrations so a swap is a configuration change rather than a project is the structural version of this advice, set out in hedging against model lock-in. The spend-control side is in turning on your spend controls.
Frequently Asked Questions
What did Anthropic and OpenAI announce on September 22?
Anthropic released Claude Opus 5.5, which it says performs at the level of its Fable 5.1 flagship while costing around 40% less to run than Opus 5 from July, and described it as the first of a family with Sonnet 5.5 and Haiku 5.5 expected over the coming weeks. OpenAI released GPT-6 Sol and GPT-6 Luna, offshoots of its GPT-6 Astra flagship, positioned for everyday work, with API pricing it describes as 50% below the promotion currently running on GPT-5.6.
Are the two price cuts comparable?
Not directly, and the difference matters. Anthropic’s roughly 40% is measured against Opus 5, a previous list price. OpenAI’s 50% is measured against a promotional rate currently running on GPT-5.6, and a discount against a promotion is a different claim from a discount against list. Fortune reported that OpenAI’s pricing comes out lower than Anthropic’s based on each company’s announcement, which is a comparison of two announcements rather than an independent analysis.
Why might a price cut not reduce my bill?
Three reasons. Promotional rates carry end dates, so the number you budgeted against may revert. Long-context requests are billed on a separate and higher meter at some providers, so the headline rate applies only to short requests. And the same model costs different amounts depending on where you buy it, since resellers and cloud marketplaces add margin. Your invoice is decided by all three, not by the announcement.
What is long-context pricing?
Several providers publish one rate for ordinary requests and a higher one once a request exceeds a context threshold. Published OpenAI rates show the current flagship tier moving from $5 to $10 per million input tokens and $30 to $45 output on the long-context meter, with cached input doubling alongside the input rate. If your workload sends large documents, you may be paying the second number while budgeting against the first.
Does buying through a cloud marketplace cost more?
Often, and it is worth checking rather than assuming. Reporting this week indicated Moonshot’s Kimi K3 arriving on Amazon Bedrock at roughly 1.76 times OpenRouter pricing for the same model. That is one example and not a general rule, but it establishes that channel is now a pricing variable. Marketplaces can be worth a premium for billing consolidation, data agreements and procurement simplicity, which is a decision to make deliberately rather than by default.
Know your cost per task, not per token
We calculate what your AI workloads actually cost across providers and channels, find the promotional rates about to expire, and keep switching cheap.
Related Articles
Your Website May Soon Need a Version for Agents
Why Your Next Laptop Costs More: AI Is Hitting Hardware Prices
AI Is Getting More Reliable, Not Just More Capable
Ajan leads the ChatGPT.ca team: 200+ custom GPT builds and automation projects for 50+ businesses across 20+ industries. Based in Markham, Ontario. PIPEDA-compliant solutions.