Smarter AI Can Also Lie More: A Warning Sign
Here is a result that should quietly change how you think about AI upgrades. Independent testers evaluated one of this month's most talked-about new models and found two numbers moving in the same direction at once: its factual accuracy improved (roughly 33% to 46%), and its hallucination rate also climbed (about 39% to 51%). The model got smarter and became more willing to confidently state things that are not true. Most of us assume newer means more reliable. This is a clear, measured case where that assumption breaks.
Two different questions, two different answers
It sounds like a contradiction until you see what each measure is actually asking. Accuracy asks: when the model answers, how often is it right? Hallucination asks something else entirely: how often does it confidently invent an answer instead of admitting it does not know? A model can get better at retrieving and reasoning over real information while simultaneously becoming more inclined to produce a confident-sounding answer no matter what. Getting more capable does not automatically make a system more careful about the edges of its own knowledge. Those are separate skills, and only one of them gets celebrated in a launch announcement.
Why this is a business problem, not a trivia problem
The danger is subtle. A more capable model writes better, reasons more fluidly, and sounds more authoritative, which means its mistakes are better disguised. Meanwhile your team reads "newest model" as "fewer errors" and quietly lowers its guard. Trust drifts ahead of reliability, exactly when the wrong answers got harder to spot.
| What people assume | What the data shows |
|---|---|
| Newer model means fewer mistakes | Accuracy and fabrication can both rise |
| More capable means more trustworthy | Capability and truthfulness are separate |
| Sounding authoritative signals correctness | Better writing hides errors more effectively |
We flagged this trade-off when the model launched, in our look at Kimi K3. The independent numbers now put a figure on it, and the figure is the point: this is measurable, not vibes.
Habits that survive every upgrade
The fix is not to avoid new models, they really are better tools. It is to build checks that do not depend on which model you happen to be using. Verify the specifics: figures, names, dates, citations, links. Prefer tools that show their sources for factual claims, and actually look at the sources rather than trusting the summary. Tell your team plainly that a confident answer is not a verified one, and that this stays true as the models improve. Keep a human accountable for anything published, sent to a client, or acted on. Boring discipline outperforms trusting a leaderboard, every time.
The signal to watch
As the AI race accelerates, expect more of this: models that are genuinely more capable on the metrics vendors advertise, while quieter qualities like "knows when to say I don't know" lag behind. Treat every upgrade announcement as a claim about capability, not a promise of accuracy. The businesses that get burned will be the ones that let their guard down because the output got more polished. The ones that do well will simply keep verifying, and enjoy a genuinely better tool without betting the business on it being right.
Frequently Asked Questions
What did the testing actually find?
Independent evaluation of a major new AI model (Moonshot’s Kimi K3) found something counterintuitive: on the same test, its factual accuracy improved noticeably over the previous version, from roughly 33% to 46%, while its hallucination rate also rose, from about 39% to 51%. In plain terms, the model got better at knowing things and also more willing to confidently state things that are not true. Both numbers moved up together. That combination is the part worth paying attention to, because it breaks the assumption most people carry about AI progress.
How can a model get more accurate and less truthful at the same time?
Because they measure different behaviours. Accuracy asks, "when it answers, how often is it right?" Hallucination asks, "how often does it confidently make something up instead of admitting it does not know?" A model can improve at retrieving and reasoning over real information while simultaneously becoming more inclined to produce a confident answer no matter what. Being more capable does not automatically make a model more careful about the limits of its knowledge. Capability and truthfulness are separate dials, and they do not always move in the same direction.
Why does this matter for my business?
Because most people judge AI by how impressive it sounds, and a more capable model sounds more convincing when it is wrong. If your team assumes "newer model, fewer mistakes," they will lower their guard exactly when the confident-but-false answers become harder to spot. The practical risk is not that AI is unusable, it is that trust drifts ahead of reliability. This is a reminder to keep your verification habits regardless of which model you use, and to treat "smarter" as a claim about capability, not a guarantee of accuracy.
Does this mean I should not use newer AI models?
No. Newer models are genuinely more capable and usually a better tool. The correct response is not to avoid them, but to stop treating a model upgrade as a reason to relax your checks. Use the newer, stronger model for the work, and keep the same discipline about verifying anything factual, financial, legal, or customer-facing. If anything, a more persuasive model deserves slightly more scrutiny on the claims it makes, not less, because its mistakes are better disguised.
What should a Canadian business actually do about this?
Build habits that do not depend on which model you are using. Verify specific facts, figures, names, dates, citations, and links before you rely on them or publish them. Prefer tools that show sources for factual claims, and check the sources rather than trusting the summary. Tell your team plainly that a confident answer is not a verified one, and that this stays true as models improve. Reserve AI for drafting and analysis where a human reviews the output, and keep a person accountable for anything consequential. Simple, boring discipline beats trusting a leaderboard.
Get AI speed without the confident mistakes
We help Canadian businesses build verification habits and review workflows that keep AI genuinely useful as the models keep changing underneath you.
Related Articles
Intelligence Per Dollar: A Smarter AI ROI Metric
Why AI Benchmarks Can Mislead You — and How to Actually Evaluate AI for Your Business
Why Your Next Laptop Costs More: The AI Buildout Is Hitting Hardware Prices
AI consultants with 100+ custom GPT builds and automation projects for 50+ Canadian businesses across 20+ industries. Based in Markham, Ontario. PIPEDA-compliant solutions.