Skip to main content
Security & Compliance6 min read

AI Safety Testing Is Moving Outside the Vendors

August 17, 2026By Ajan Kanagalingam

Two stories landed in the same week, pointing in opposite directions. The Verge reported, citing the Financial Times, that OpenAI disbanded its preparedness team at the end of July and folded its risk work into existing groups. At roughly the same time METR, which stress-tests frontier AI systems from the outside, raised 71 million dollars to do considerably more of it. Put those together and you can see the shape of something: the checking is migrating out of the companies that build these systems and into institutions with no commercial interest in the result.

Be fair about the restructure

It is worth resisting the dramatic reading. Companies reorganise constantly, and moving risk specialists into product teams can genuinely place expertise closer to the decisions that matter, which is an argument organisations make in good faith across many industries. The reason it still deserves attention is structural rather than moral. A dedicated team with its own reporting line can say no in a way that becomes considerably harder for a specialist sitting inside a team with a shipping deadline. The change might turn out to be an improvement. The point is that from the outside, nobody can tell, and the people affected are downstream.

You inherit these decisions

A handful of companies build the models underneath almost every AI tool a Canadian business touches. Their internal testing standards travel downstream into whether your tool behaves predictably, resists being manipulated by content designed to hijack it, and fails in a way you notice rather than silently. You did not participate in any of those decisions and you cannot verify them. That is not a reason for anxiety. It is a reason to pay attention to which vendors submit to outside examination, because that is the one signal available to you.

Four questions worth asking

You do not need technical expertise to ask these, and the answers separate serious vendors from the rest faster than any feature grid.

AskWhat a good answer looks like
Who tested this from outside?A named third party and a published result
How are we told about model updates?Advance notice, with a way to stay on a version
What are the known failure modes?A specific list, not a reassurance
Who do we call when it misbehaves?A named channel and a response time

The third row is the most revealing. A vendor who cannot name a single situation where their product performs poorly either does not know or will not say, and both answers tell you something. It sits alongside the questions we set out for the aftermath of an incident, except these are the ones you ask beforehand.

Independent does not mean suitable

There is a tempting shortcut here that is worth naming and refusing. External evaluation tells you a model has been examined for broad and serious failure modes by someone with no reason to flatter it. It tells you nothing about whether the tool suits your process, your data, or your obligations under Canadian privacy law. Those decisions stay with you permanently. What independent testing does is raise the floor beneath you, which is real value, and it is not the same thing as a certificate that the tool is right for what you plan to do with it.

Where this is heading

Every maturing industry ends up with third-party assurance, because self-certification eventually stops persuading anyone. Financial audit, food safety, building inspection, electrical certification: none of these began as external functions, and all of them moved outside once the stakes grew and the conflict of interest became obvious. AI is early in that same progression, and the money now flowing to independent evaluators is what the beginning of it looks like. The practical consequence for buyers arrives sooner than the regulation: within a year or two, "independently evaluated" becomes a normal procurement question, and the vendors who cannot answer it start losing deals to the ones who can. Asking early costs you nothing and puts you ahead of your own supply chain, which is the same logic behind treating the emerging AI security standard as a buyer checklist.

Frequently Asked Questions

What happened?

Two things landed in the same week and they point in opposite directions. The Verge, citing the Financial Times, reported that OpenAI disbanded its preparedness team at the end of July, folding its risk work on areas such as biological and cyber threats into existing groups. Meanwhile METR, an organisation that stress-tests frontier AI systems from the outside, raised 71 million dollars to do more of it. Read together, they describe a shift in where the checking happens: less of it inside the companies building the systems, more of it in institutions with no commercial stake in the answer.

Is folding a safety team into other teams necessarily bad?

Not automatically, and it is worth being fair about this. Companies restructure constantly, and integrating risk specialists into product teams can genuinely put expertise closer to the decisions, which is an argument organisations make sincerely in many industries. The reason it still deserves attention is structural rather than moral. A dedicated team with its own reporting line can say no in a way that a specialist embedded in a shipping team finds much harder, particularly under deadline. The change may be fine. It is not nothing, and buyers do not have visibility either way.

Why should a small business care about frontier AI safety?

Because you inherit these decisions without participating in them. The models underneath the tools your business uses are built by a handful of companies, and their internal testing standards flow downstream into whether a tool behaves predictably, resists manipulation, and fails safely. You will never audit a frontier model yourself, and nobody expects you to. What you can do is treat evidence of external scrutiny as a purchasing signal, in the same way you would treat a food safety certification without personally inspecting the kitchen.

What should we actually ask vendors?

Four questions that a serious vendor can answer and a weak one cannot. Has this system been evaluated by anyone outside your company, and can we see what they published? What happens when your model is updated, and how will we be told? What are the documented failure modes, meaning the situations where you know it performs badly? And who do we contact, with what response time, when it behaves unexpectedly in our environment? None of these require technical expertise to ask, and the quality of the answers separates vendors more reliably than any feature comparison.

Does independent testing mean the tool is safe for our use?

No, and conflating the two is the common mistake. External evaluation tells you the underlying model has been examined for broad, serious failure modes by someone with no incentive to be flattering. It says nothing about whether the tool is appropriate for your specific process, your data, or your regulatory obligations. Those remain your responsibility and always will. The value of independent testing is that it raises the floor beneath you, which is genuinely worth something. It does not replace deciding what you are willing to let the tool do.

Ask the hard questions before you sign

We help Canadian businesses evaluate AI vendors on evidence rather than marketing, covering external testing, update policy, failure modes, and support.

Related Articles

Security & Compliance

The Next Laptop You Buy Ships With AI Already On It

August 15, 2026Read more →
Security & Compliance

AI That Watches Your Screen So You Stop Repeating Yourself

August 14, 2026Read more →
Security & Compliance

AI Reasoning Traces Leaked Real Passwords and Keys

August 12, 2026Read more →
AK
Ajan Kanagalingam
Founder & ChatGPT Consultant, ChatGPT.ca

Ajan leads the ChatGPT.ca team: 200+ custom GPT builds and automation projects for 50+ businesses across 20+ industries. Based in Markham, Ontario. PIPEDA-compliant solutions.

Stay ahead of AI in Canada

Weekly case studies, new tools, and ROI playbooks for Canadian SMEs. One email, zero spam.