AI Monitoring: Watch What It Does, Not What It Says
If your AI tool stopped working tomorrow, you would know within an hour. If it kept working perfectly and started being wrong, you might not know for six months. Almost every business has monitoring for the first situation and none for the second, which is why AI problems in small companies are usually discovered by a customer rather than by anyone inside.
Two different questions
Is it running is an infrastructure question and somebody else usually answers it. Is it still right is a business question and nobody answers it unless you decide to.
Your vendor knows whether their service responded and how quickly. They have no idea whether this month's summaries were accurate for your business, whether the tone still suits your customers, or whether the tool has started handling a category of work you never intended it to touch. That is not negligence on their part. They own the service, you own whether it is doing your job correctly.
Why output drifts without anything breaking
| Cause | Why you will not notice |
|---|---|
| The model behind it gets updated | Often without notice, behaviour shifts slightly |
| Your inputs change | The business moved, the tool did not |
| People widen how they use it | Nobody decided, it just spread |
All three are gradual, and gradual is what defeats casual observation. Nobody spots a slow decline in a thing they see every day, which is why this has to be a deliberate check rather than a general intention to pay attention.
Four things worth watching
A real sample, read properly. Ten actual outputs, read the way you would read a new employee's work. Not skimmed.
Volume. A sudden change in how much the tool is doing almost always means something upstream changed, and it is the earliest signal you get.
Overrides. How often a person corrected or rejected the output. If that rate is climbing, something has shifted, and if it is zero you should check whether anyone is reviewing at all.
Cost. Runaway processes show up in spend before they show up anywhere else, which is why spend is a quality signal and not just a finance one.
Fifteen minutes, weekly or monthly depending on volume, in a recurring calendar entry. No software required.
Sample the awkward cases
This is the change that catches the most and costs the least.
Most people sample randomly, which means they mostly review typical cases, which the tool handles well. Meanwhile it may be doing badly on the ten percent that are unusual: the awkward client, the non-standard job, the enquiry in the second language, the situation nobody thought about when it was set up.
Because those are rare, each error looks like a one-off and no pattern emerges. Deliberately pulling your sample from the unusual cases surfaces it immediately. It is also where the differential treatment we described in AI bias in a small business tends to hide.
Log what it did, not what it reported
For anything that takes actions rather than producing text, the distinction sharpens. A summary of what an agent did is a claim. A log of what it actually touched is a record.
Keep the record. It is the difference between being able to answer a question afterwards and having to reconstruct it, and it is the specific gap behind businesses that cannot trace what their agents have been doing. Same principle as knowing who can explain a piece of work, which we covered in tracking work when AI writes it.
Decide the threshold in advance
Monitoring only helps if something happens when the number moves. Otherwise you have built a habit of looking at things.
So write down what would make you act. Override rate above a certain point. Two customer complaints about the same tool in a month. Any output that reached a customer and was wrong. Then decide what acting means, which is usually narrowing what the tool handles rather than switching it off entirely. That pairs with having a defined stop condition, and keeps the review from becoming the kind of ritual nobody responds to.
Frequently Asked Questions
What is AI monitoring?
Checking that an AI system is still doing the right thing, which is a different question from whether it is running. Uptime monitoring tells you the service responded. Quality monitoring tells you the responses were still correct, still on-brand, and still within the boundaries you set. Almost every business has the first and almost none have the second, which is why AI problems in small businesses tend to be discovered by a customer rather than by anyone internally.
Why does output quality drift?
Three reasons, and none of them involve anything breaking. The model behind your tool gets updated, sometimes without notice, and behaves slightly differently. Your own inputs change, because the business changed, so the tool is now handling cases it was never tested on. And people gradually widen how they use it, feeding it work nobody considered when it was set up. All three are gradual, which is exactly why they are invisible without a deliberate check.
What should a small business actually watch?
Four things, checked weekly or monthly depending on volume. A sample of real output, read properly rather than skimmed. Volume, because a sudden change in how much the tool is doing usually means something upstream changed. Exceptions, meaning cases where a human overrode or corrected it. And cost, because a runaway process shows up in spend before it shows up anywhere else. None of that needs software; a recurring calendar entry and fifteen minutes covers it.
Is not the vendor monitoring this?
They monitor their service, not your outcomes. A vendor knows whether their system is responding and how fast. They have no idea whether the summaries it produced this month were accurate for your business, whether the tone suits your customers, or whether it started handling a category of work you never intended. That gap is not negligence, it is the natural division: they own the service, you own whether it is doing your job correctly.
What is the most common failure people miss?
Silent degradation on the cases that matter least often. A tool handles the common situations well, and everybody judges it on those, while quietly doing badly on the ten percent that are unusual. Because those are rare, nobody notices a pattern, and each individual error looks like a one-off. Sampling deliberately across the unusual cases rather than the typical ones catches this, and it is the single highest-value change to how most businesses review AI output.
Catch AI drift before your customers do
We help Canadian businesses set up lightweight AI monitoring: what to sample, what to log, and what triggers a change.
Related Articles
Does Claude Opus 4.8 Answer the AI ROI Question?
COBOL to Java Migration: What Does It Actually Cost in 2026?
AI Is Getting a Standard Way to Control Equipment
Ajan leads the ChatGPT.ca team: 200+ custom GPT builds and automation projects for 50+ businesses across 20+ industries. Based in Markham, Ontario. PIPEDA-compliant solutions.