Skip to main content
Data & Analytics6 min read

Judge AI on What It Finds, Not How It Talks

August 14, 2026By ChatGPT.ca Team

Perplexity reported that its Search as Code offering beat both OpenAI and Anthropic on four of five search benchmarks. Set aside the scoreboard for a second, because the interesting part is what it implies. The companies with the most impressive conversational models are not automatically the best at finding things. Those are separate skills. And for most of what businesses actually ask AI to do, the second one matters far more than the first.

Most business questions are lookups

Notice what your team actually asks AI. Very little of it is "write me something beautiful." Most of it is: what are the current rules for this, what did that supplier announce, what are competitors charging, what changed in this regulation, what is the standard practice here. Every one of those is a lookup. The model’s prose style contributes nothing; its ability to find current, correct, relevant sources contributes everything. If you are buying AI mainly for questions like these, you are buying a research tool that happens to be conversational, and you should evaluate it that way.

Why polish makes bad sourcing worse

Here is the uncomfortable interaction between the two skills. A beautifully written answer built on the wrong three web pages is more dangerous than a clumsy answer built on the right ones, because the fluency is what convinces you not to check. Strong writing on weak retrieval is the worst combination available, and it is unfortunately a common one. This is a cousin of the problem we covered in AI inventing its sources, except here the sources are real, they are simply the wrong ones, which makes it much harder to spot.

The warning signs

Weak retrieval has a recognizable signature once you know to look for it.

SymptomWhat it usually means
Confident answer, vague sourcingIt is reasoning, not looking things up
Citations that do not support the claimSources fetched, then largely ignored
Out-of-date facts stated as currentNo freshness weighting in retrieval
Different sources on a repeat questionRetrieval is unstable, so trust is too

The one-hour test

Do not evaluate this from a leaderboard. Take five real questions from your own work where you already know the correct answer and where the good sources are, and deliberately include one where the obvious top web result is out of date or misleading, because that is the question that separates tools. Run all five through each candidate and check three things: did it find the right sources, was the information current, and does the answer actually follow from what it cited? An hour of that tells you more than any benchmark, for the same reason we argue benchmarks mislead: they measure someone else’s questions.

The Canadian trap

This one catches businesses here constantly, so make it one of your five test questions. Tools trained and tuned largely on American sources will confidently return US tax rules, US regulations, US pricing, and US employment law when you ask a general question, with no signal that the answer does not apply to you. Ask about something Canadian-specific where you know the right answer, and see whether the tool gets there or quietly hands you the American version. And regardless of which tool wins, keep the standing rule: anything consequential gets its sources checked by a person before it becomes a decision or reaches a client.

Split your evaluation

The practical upshot is to stop looking for one best AI and start matching tools to jobs. Judge the tools your team uses for drafting and thinking on reasoning and writing. Judge anything doing lookups on retrieval, using your own questions. They may well be different products, and that is fine, a slightly awkward tool that reliably finds the truth beats an elegant one that reliably sounds right. It is the same principle as making sure your data is ready for AI: the quality of what goes in decides the value of what comes out.

Frequently Asked Questions

What is the difference between chat quality and research quality?

Chat quality is how well a model reasons and writes with what it already has. Research quality is how well it goes out, finds the right sources, and grounds its answer in them. They are separate skills, and a tool can be excellent at one and mediocre at the other. Perplexity recently reported beating both OpenAI and Anthropic on four of five search benchmarks with its Search as Code offering, which is a useful reminder: the company with the most impressive conversational model is not automatically the one that will find the right answer for you.

Why does this distinction matter for business use?

Because most business questions are lookups, not essays. What are the current rules for this, what did this supplier announce, what are competitors charging, what changed in this regulation. For all of those, the model’s eloquence is irrelevant and its ability to find current, correct, relevant sources is everything. A beautifully written answer built on the wrong three web pages is worse than a plain one built on the right ones, because the polish makes it more persuasive. If your AI is doing lookups, you are buying a research tool that happens to talk.

How do I actually test research quality?

Ask questions you already know the answer to. Take five real questions from your own work where you know the correct answer and where the good sources are, ideally including one where the obvious web result is out of date or wrong. Run them through each tool and check three things: did it find the right sources, did it use current information, and did the answer actually follow from what it cited? That takes an hour and tells you more than any leaderboard, because it is measuring performance on your questions rather than someone else’s.

What are the warning signs of poor retrieval?

Watch for confident answers with vague or missing sourcing, citations to pages that do not actually support the claim, and a tendency to reach for whatever ranks highest rather than what is authoritative. Out-of-date information presented as current is especially common and especially costly in areas like tax, regulation, and pricing. Another tell is inconsistency: ask the same question twice and see whether you get materially different sources. A tool with weak retrieval will feel fine until the day it quietly gets something important wrong.

What should a Canadian business do about this?

Split your evaluation. Judge the tools your team uses for drafting and thinking on reasoning and writing; judge anything doing lookups on retrieval, with your own test questions. Be especially careful with Canadian-specific queries, since tools trained and tuned on largely American sources routinely return US rules, US pricing, and US regulations with total confidence. And regardless of the tool, keep the rule that anything consequential gets its sources checked by a person before it becomes a decision or reaches a client.

Make sure your AI finds the right answer

We help Canadian businesses test AI tools on retrieval quality with their own questions, catch the American-answer trap, and set review rules that hold.

Related Articles

Data & Analytics

Your Data Isn't Ready for AI (and That's the Problem)

August 5, 2026Read more →
Data & Analytics

Quality Data Is Becoming AI’s Real Bottleneck — and Your Hidden Advantage

July 6, 2026Read more →
Data & Analytics

AI Can Be Fooled by Tiny Changes: Why Robustness Matters Before You Automate

June 26, 2026Read more →
AI
ChatGPT.ca Team

AI consultants with 100+ custom GPT builds and automation projects for 50+ Canadian businesses across 20+ industries. Based in Markham, Ontario. PIPEDA-compliant solutions.

Stay ahead of AI in Canada

Weekly case studies, new tools, and ROI playbooks for Canadian SMEs. One email, zero spam.