Performance Reviews When AI Wrote the Work
Review season arrives and the usual evidence has stopped meaning much. Someone produced forty proposals this year instead of twenty. Someone else wrote three reports a week. Neither number tells you what it used to, because the first draft of most of it came out of a model in ninety seconds, and the person who produced twice as much may simply have pasted twice as often.
What stopped being evidence
| No longer tells you much | Still tells you something |
|---|---|
| Number of documents produced | What they chose to work on first |
| Speed of first draft | Errors they caught before anything left |
| Length or polish of written work | What they escalated instead of guessing |
| Responsiveness in a chat window | How the work held up three months later |
The right-hand column has a useful property. None of it is inflated by a faster draft, and none of it can be produced by a tool on someone's behalf. It is also harder to observe, which is why it needs writing down during the year rather than reconstructing in December.
The error-catching question
If you ask one new question this cycle, ask what they caught.
A model produces confident, plausible, fluent output that is sometimes wrong. The person who reads it, notices the figure is off, and checks before it reaches a customer is doing the job that now matters most. The person who forwards it unchecked produced the same volume and created a liability.
That distinction is invisible in any output metric and obvious the moment you ask for examples. Someone who can name three things they caught this year is telling you they read their own work. Someone who cannot may simply not have been keeping track, and that is worth a follow-up rather than a conclusion.
Do not assess AI adoption
Some organisations have started scoring people on whether they use AI. It is an understandable impulse and it rewards the wrong thing.
Adoption is activity. Score it and you get activity, including from people whose work was better before. The careful colleague who uses AI for two tasks and refuses it for the third, because they have tested it and it was not good enough, is exercising exactly the judgment you want and will score badly on an adoption measure. The general version of that failure is in your AI will optimise for the score you set.
Assess the outcome. Whether someone got there with a model, a spreadsheet or a pencil is their business.
Writing the review itself
A performance review is one of the few documents where the specifics are the entire value. Ask a model to write one and you get balanced, fluent, professional paragraphs that could describe any competent person in that role.
Employees recognise that instantly. A review that reads as generated tells someone their manager did not spend twenty minutes thinking about them, which is a worse message than anything the review contained. A shorter, plainer document with two real examples in it lands better than three polished pages of nothing.
Using an assistant for structure and tone is reasonable. Give it your own notes and ask it to organise them, then check that every specific in the output came from you. The failure to avoid is the fluent generic paragraph, which is the same problem as workslop appearing in a document where it does real damage.
The conversation that belongs before the review
If half of what someone did last year is now handled by a tool, reviewing them against their old job description is unfair and common.
Have the role conversation separately and first. What the job is now, what the remaining half consists of, and what it should become. Then review against that. Rolling both into one meeting produces a discussion where the person is defending their performance and their existence at the same time, and nobody does their best thinking in that position. The wider shift is covered in how AI is redrawing job descriptions.
Start the note-keeping now
Everything above needs specifics, and specifics need to be recorded when they happen. Three or four notes per person through the year, each a sentence: what they caught, what they decided, what they flagged.
That file turns a December review from an hour of reconstruction into twenty minutes of writing, and it is the only input to this process that no tool can supply for you. If you also monitor how work gets done, keep that separate from this and read employee monitoring with AI first, because surveillance data makes a poor substitute for observation.
Frequently Asked Questions
How should performance reviews change now that staff use AI?
Stop treating output volume as evidence. When a model produces the first draft of everything, the number of documents someone produced says more about their tooling than their contribution. Assess judgment instead: what they chose to work on, what they caught before it went out, what they escalated, and how their work held up. Those are observable and they are not inflated by a faster draft.
Should managers use AI to write performance reviews?
For structure and tone, carefully. For the substance, no. A review is one of the few documents where the specifics are the entire value, and a model asked to write one will produce fluent, balanced, generic paragraphs that could describe anyone. Employees recognise that immediately, and a review that reads as generated damages trust more than a shorter, plainer one written by hand.
Is it fair to assess someone on how well they use AI?
It is fair to assess the outcome, and risky to assess the tool use directly. Someone who produces good work slowly without AI is performing. Someone producing fast, plausible, unchecked output is not, whatever the volume says. Judge the result and the judgment behind it rather than adoption, because adoption metrics reward activity and tend to penalise the careful people you most want to keep.
What should we document during the year?
Specific instances rather than impressions. The call they handled well, the error they caught, the decision they made when the answer was unclear, the thing they flagged that nobody else noticed. Three or four noted through the year produce a better review in twenty minutes than an hour of trying to remember in December, and they are the raw material no assistant can supply.
How do we review someone whose role has been partly automated?
Review the role as it exists now rather than as it was written. If half of what someone did is now handled by a tool, the honest conversation is about what the remaining half is and what the role becomes, and that conversation belongs before the review rather than inside it. Assessing someone against a job description the business has quietly changed is unfair and it is also the most common version of this problem.
Redefine the role before you review it
We help Canadian businesses work out what each role actually consists of once AI has absorbed part of it, so reviews measure the job people are doing now.
Related Articles
Co-op Students Are How Small Firms Get AI Help
Change Management for an AI Rollout That Sticks
Billable Hours When AI Does Part of the Work
Ajan leads the ChatGPT.ca team: 200+ custom GPT builds and automation projects for 50+ businesses across 20+ industries. Based in Markham, Ontario. PIPEDA-compliant solutions.