Data Loss Prevention When Staff Use AI Every Day
Somebody in your business pasted a customer email into an AI tool this week to get help drafting a reply. They were not being careless. The approved path was slower, the customer was waiting, and the tool is genuinely good at that. Multiply it by everyone, daily, and you have a data loss problem that no policy document has ever solved.
Why the old approach does not transfer
Traditional data loss prevention watched files: email attachments, USB drives, uploads. Those are discrete events with clear boundaries, and you can inspect them.
What leaves through a chat window is prose. High volume, continuous, and almost always well-intentioned. Nobody is exfiltrating anything. They are pasting the context needed to get a useful answer, which happens to include a name, an address, and an account number. Same objective as before, considerably harder surface.
Blocking makes it worse
The instinct is to prohibit the tools. It fails predictably, and it fails in the direction that costs you most.
People switch to a personal phone. The behaviour continues, your visibility drops to zero, and you lose whatever productivity the tools were providing. You also signal to your team that you would rather not know, which is not the culture you want when something eventually goes wrong.
This is the same dynamic behind shadow AI: suppression relocates the risk rather than removing it.
What actually reduces exposure
| Do this | Because |
|---|---|
| Give them a good sanctioned tool | Most leakage is the approved option being worse |
| List specific examples, not categories | Nobody knows what counts as sensitive data |
| Redact before pasting, as a habit | Five seconds, removes most of the exposure |
The middle row deserves emphasis. A policy saying do not share confidential or sensitive information tells people nothing, because everyone draws that line differently. Do not paste a customer's full name and address together is a rule someone can follow at speed.
Software can now help
This part is newer than most people realise. Small guard models have become good enough and fast enough to sit between a person and a cloud AI service, flagging or stripping personal information before anything leaves the device.
Perplexity shipped one this week at around 0.6 billion parameters, built for exactly this. The size is the interesting part: small enough to run locally without noticeable delay, which is what makes it usable rather than a checkpoint people route around.
It is not a complete answer. A guard model catches recognisable patterns like names, addresses, and account numbers, and it will miss information that is sensitive because of context rather than format. But this is a real control that did not practically exist for a small business a year ago, and it is worth asking your vendors whether they offer something similar.
The obligation applies whether you approved it or not
Under PIPEDA you stay accountable for personal information handed to a third party, and an AI vendor is a third party. Quebec Law 25 raises the bar further for businesses handling personal information there.
The part worth internalising: that accountability does not depend on whether you sanctioned the tool. Staff using something you never approved still creates your obligation, which is the strongest practical argument against pretending shadow usage is not happening. The wider framing is in keeping AI use PIPEDA-compliant.
Start by asking, not auditing
Before buying anything, ask your team which AI tools they actually use and what they paste in. Ask it in a way that makes honesty safe, because the answer is only useful if it is true.
You will usually learn two things. There are more tools in use than you thought, and the reason is nearly always that the sanctioned option is slower or worse at that specific job. Fixing that is cheaper than any control you could buy, and it makes every subsequent control easier to enforce. The training half of this is covered in what your team actually needs to know, where the short list of what never goes in is one of the four things worth teaching.
Frequently Asked Questions
What is data loss prevention in an AI context?
Stopping sensitive information from leaving your business through the tools your team uses to get work done. Traditional data loss prevention watched email attachments and USB drives. The AI version watches what gets typed or pasted into a chat window, which is a harder problem because the volume is high, the intent is almost always innocent, and the information arrives as ordinary prose rather than as a file somebody attached. Same objective, different surface.
Why does blocking AI tools fail?
Because it moves the activity to a personal phone where you cannot see it at all. Staff paste customer details into AI tools for the same reason they always used workarounds: the sanctioned path is slower and the deadline is real. Blocking removes your visibility without removing the behaviour, and it also removes the productivity you were presumably trying to capture. Every organisation that has tried a hard block has ended up with the same problem plus a worse relationship with its own team.
What actually reduces the risk?
Three things in order of value. Give people a sanctioned tool that is genuinely good, because most leakage happens when the approved option is worse than the unapproved one. Write a short, specific list of what must never be pasted, phrased as examples rather than categories. And redact by habit, meaning teach people to strip names and account numbers before pasting rather than after. That last one is a five-minute habit that removes most of the exposure at almost no cost.
Can software catch it automatically?
Increasingly yes, and this is newer than most people realise. Small guard models now run locally, fast enough to sit between a user and a cloud AI service and flag or strip personal information before anything leaves the device. Perplexity recently shipped one at around 0.6 billion parameters for exactly this purpose. It is not a complete answer, since a guard can miss context-dependent sensitivity, but it is a real control that did not practically exist for small businesses a year ago.
What does Canadian law require here?
Under PIPEDA you remain accountable for personal information you hand to a third party, which includes an AI vendor, and Quebec Law 25 raises the bar further for businesses handling personal information there. Practically that means knowing which tools your team uses, having an agreement with the vendors that matter, and being able to describe what leaves your business. Note that the obligation attaches whether or not you approved the tool, which is the strongest argument against pretending shadow usage is not happening.
Keep customer data out of the wrong tools
We help Canadian businesses set practical AI data rules, choose sanctioned tools people will actually use, and meet their privacy obligations.
Related Articles
When AI Invents Its Sources: A Professional Risk
AI Browser Agents Are Getting Safer to Use
Always-On AI: When Devices Record Everything
Ajan leads the ChatGPT.ca team: 200+ custom GPT builds and automation projects for 50+ businesses across 20+ industries. Based in Markham, Ontario. PIPEDA-compliant solutions.