There Is a New AI Model Every Three Weeks Now
Google shipped Gemini 3.7 Flash about three weeks after 3.6 Flash, with notably better coding scores and a 50 percent introductory price cut. Three weeks. That is the interval now, across providers, and it creates a problem nobody had to solve two years ago: the model you carefully evaluated and standardized on last month may have been superseded twice before you finished rolling it out. The answer is not to move faster. It is to decide, in advance, what would actually make you move at all.
Two ways to get this wrong
The first is chasing. Every launch looks compelling, every benchmark chart goes up and to the right, and it feels irresponsible to be running last month’s model. But each switch costs real time: re-testing your actual workflows, re-tuning prompts calibrated to the previous model’s quirks, retraining people on changed behaviour, and absorbing whatever regressions arrive uninvited. Switch every three weeks and you will spend all your time switching and none of it compounding. The second failure is the mirror image: standardize once, never revisit, and discover your model is being deprecated with a deadline attached. Both come from having no policy.
Decide your triggers before the announcements
The fix is unglamorous: write down what would justify a move, before the next launch makes the case for you.
| Worth switching for | Not worth switching for |
|---|---|
| A capability you genuinely cannot do today | A higher benchmark score |
| A real cost drop at your actual volume | A headline price comparison |
| Your current model being deprecated | A competitor announcing they switched |
| Measurably better results on your own tests | General sense that it is newer |
That first right-hand row is the one that catches people. Benchmark gains routinely fail to show up in your specific work, which is exactly the trap we described in why AI benchmarks mislead. A model that is better on average can still be worse at your task.
Put the review on the calendar
Reactive evaluation means every launch announcement gets a vote on your week. Scheduled evaluation means none of them do. Quarterly is plenty for most businesses. In that session, ask three questions: has anything shipped that hits one of our written triggers, is our current model still supported, and would we choose it again today? That is twenty minutes, four times a year, and it catches both the genuine improvement and the looming deprecation. It also pairs naturally with the upkeep review we recommend for the AI you have already built, since both are checking whether yesterday’s decision still holds.
Make switching cheap and the policy gets easy
Most of the pain of upgrading is self-inflicted, created earlier by decisions that made your setup hard to move. Three habits fix that. Keep prompts and business context stored separately from any one provider, so a switch means repointing rather than rebuilding, which is the practical payoff of treating prompts as company property. Keep a handful of real examples from your own work that you can run against any candidate to compare like for like. And be wary of building anything critical on a capability only one provider offers, the core of hedging against vendor lock-in.
The calm position
A three-week release cycle sounds like pressure, and it is only pressure if you have no policy. With written triggers, a scheduled review, and portable prompts, the next launch becomes something you glance at rather than something you react to. You will skip most of them, and that is correct. Every release you skip on purpose is time your team spent getting better at the work instead of migrating between tools. The advantage in a fast-moving market does not go to whoever adopts fastest. It goes to whoever wastes the least motion.
Frequently Asked Questions
How fast are models actually shipping now?
Fast enough that the release calendar has stopped being an annual event. Google shipped Gemini 3.7 Flash roughly three weeks after 3.6 Flash, with meaningfully better coding scores and a 50 percent introductory price cut. That is not an isolated case: across providers, 2026 has seen a steady stream of point releases, price adjustments, and capability jumps landing weeks apart rather than quarters. For a business, this means the model you carefully evaluated and standardized on last month may have been superseded twice by the time you finished rolling it out.
Should we upgrade every time something better appears?
No. Chasing every release is its own failure mode, and an expensive one. Each switch costs re-testing, re-tuning prompts that were calibrated to the old model, retraining people on changed behaviour, and absorbing whatever unexpected regressions arrive with the new version. If you switch every three weeks you will spend all your time switching and none of it compounding. The businesses that do well here are not the ones running the newest model. They are the ones who decided, in advance, what would actually justify a move.
What should trigger an upgrade then?
Set the triggers before the announcements, so you are responding to your needs rather than to marketing. Three good ones: a capability you currently cannot do at all becomes possible, your costs would drop meaningfully at your actual usage volume rather than in a headline comparison, or your current model is being deprecated and you have no choice. Notably absent from that list is "it scored higher on a benchmark," because benchmark gains frequently do not show up in your specific work. Write your triggers down, then ignore everything that does not hit one.
How do we avoid getting stuck on something stale?
By making the review scheduled rather than reactive. Put a recurring check in the calendar, quarterly is plenty for most businesses, and in that session ask three questions: has anything appeared that hits one of our triggers, is our current model still supported, and would we choose it again today knowing what we now know? That rhythm catches genuine improvements without letting every launch announcement disrupt your week. The failure mode it prevents is the common one: standardizing on something in year one and never revisiting until it is deprecated out from under you.
What makes switching easier when we do decide to?
Portability, designed in early. Keep your prompts and business context stored separately from any one provider rather than embedded in a single vendor’s interface, so moving means repointing rather than rebuilding. Keep a small set of real examples from your own work that you can run against any candidate model to compare like for like. And avoid building critical processes on a capability only one provider offers, unless the advantage is worth the dependence. Do that and an upgrade becomes a decision you make calmly rather than a project you dread.
Stop reacting to every AI launch
We help Canadian businesses set model upgrade triggers, run candidate models against their real work, and build setups that make switching cheap when it is worth it.
Related Articles
A New AI Model Every Week: Stop Chasing Them
The AI You Rely On Now Passes a Government Check
Kimi K3: An Open AI Model Rivals the Frontier
AI consultants with 100+ custom GPT builds and automation projects for 50+ Canadian businesses across 20+ industries. Based in Markham, Ontario. PIPEDA-compliant solutions.