Quality Assurance When AI Writes the First Draft
Most small business quality assurance was a sentence rather than a system: somebody reads it before it goes out. That worked, and it worked because producing the thing took long enough that reading it was the cheap part. Google raised Gemini's output ceiling from 64,000 tokens to a million this week, which is a useful marker for a change that already happened. The reading is now the expensive part.
Three approaches, pick one per workflow
| Approach | Fits when | Fails when |
|---|---|---|
| Review everything | Low volume, high consequence | Volume grows and nobody says so |
| Review by risk | Mixed output, clear consequences | The risk categories are never written down |
| Sample | High volume, similar items | Nobody tracks what the sampling found |
Most businesses are nominally in row one and actually in a fourth state, which is reviewing whatever somebody has time for on the day. That is the state with the worst properties of all three, because it feels like full review and provides the coverage of nothing in particular.
The list that volume does not change
Five categories get read regardless of how much of them there is. Anything that reaches a customer. Anything that moves money. Anything forming part of a contract. Anything making a decision about a person. Anything that cannot be undone.
If the volume in those categories has outgrown your ability to read it, the answer is to produce less of it, not to review less of it. The consequence of being wrong did not fall when the production cost did, and a business generating more customer-facing output than it can check has increased its risk while feeling more productive.
Everything outside those five is a candidate for sampling. Internal drafts, research notes, first passes that a person will rework anyway. Those are where the speed gain should be taken.
Sampling without fooling yourself
Select at random, not by convenience. Left to choose, a reviewer picks the short, recent and straightforward items. Errors concentrate in the long, complex and unusual ones, so convenience sampling systematically checks the material least likely to be wrong. Every fifth item, or the first after each hour, beats judgment here.
Record what you found, not just that you looked. One line per review: item, problem found or none, severity. Without that record you have an activity, not a control.
Let the defect rate set the rate. Start at 20% of a given output type. Finding problems in one in five means sample more. Two hundred reviewed with two trivial issues means you can reduce. A percentage chosen once and never revisited is a ritual.
Sample the boring workflow too. The one that has run fine for six months is the one nobody watches, and vendors update models underneath you. Output degrading quietly is the failure mode we described in the AI gap in your continuity plan.
Fluency removed the cues
Reviewing human work came with free signals. A rushed draft looked rushed. Someone out of their depth wrote hesitantly. Typos clustered where attention lapsed. None of that was in the brief and all of it told you where to look.
AI output is uniformly fluent. A paragraph built on a wrong figure reads exactly like a correct one, so the surface no longer points at the problem and the reviewer has to check content rather than scan for signals.
That makes reviewing slower per item at precisely the moment there are more items, and it is why reviewer fatigue matters more than it used to. After an hour of competent, plausible output a person stops reading and starts confirming. Shorter sessions and rotating the reviewer catch more than long ones, and the deeper problem of volume without judgment is in the workslop problem.
Check the right unit
A document that is 95% correct is 100% wrong if the incorrect 5% is the total. Reviewing at the level of the document tends to produce a judgment about tone and completeness, which is not the thing most likely to be wrong.
Decide which specific fields or claims matter and check those. In extraction work that means scoring per field rather than per document, which we set out in extracting data from PDFs. In written work it usually means the numbers, the names, the dates and anything stated as a commitment.
Write that list into the procedure for the task so the reviewer knows what they are checking before they open the file, which is one of the things a good standard operating procedure exists to say.
Frequently Asked Questions
How do you do quality assurance on AI output?
Pick one of three approaches per workflow and say which you are using. Review everything, which stays correct for low-volume high-consequence work. Review by risk, where anything reaching a customer or affecting money gets read and internal drafts do not. Or sample, where you read a fixed proportion and track how often you find a problem. The failure is having no stated approach, which in practice becomes reviewing whatever somebody has time for.
What percentage of AI output should be reviewed?
There is no universal number, and the useful question is what your defect rate tells you. Start at 20% of a given output type, record how often you find something wrong, and adjust. If you are finding problems in one in five sampled items, sample more. If you have read 200 and found two trivial issues, you can reduce. A sampling rate chosen once and never revisited is a ritual rather than a control.
What should always be reviewed regardless of volume?
Anything that reaches a customer, moves money, forms part of a contract, makes a decision about a person, or cannot be undone. That list is short and it is not negotiable by volume. If the volume of that category has grown past your capacity to read it, the answer is to produce less of it rather than to review less of it, because the consequence did not change when the production cost fell.
How do I sample without fooling myself?
Select at random rather than by convenience, and resist the pull toward the short, recent and easy items. Convenience sampling systematically misses the long, complex and unusual cases, which is exactly where the errors concentrate. A simple rule such as every fifth item, or the first item after each hour, beats a reviewer choosing which ones to look at.
What is reviewer fatigue and why does it matter now?
Attention degrades when someone checks many similar items that are mostly fine. After an hour of reading plausible, competent output, a reviewer stops reading carefully and starts confirming. It matters more now because AI output is uniformly fluent, which removes the surface cues that used to flag a weak piece. Shorter review sessions and rotating who reviews catch more than longer ones.
Decide what gets read before volume decides for you
We set review rules by consequence, design sampling that catches real problems, and track the defect rate so your QA is a control rather than a habit.
Related Articles
Your AI Can Now Send the Email, Not Just Draft It
Asynchronous AI: When Your AI Keeps Working After You Log Off
Appointment Reminders That Cut No-Shows
Ajan leads the ChatGPT.ca team: 200+ custom GPT builds and automation projects for 50+ businesses across 20+ industries. Based in Markham, Ontario. PIPEDA-compliant solutions.