Root Cause Analysis: Where AI Helps, Where It Guesses
OpenAI is currently working through roughly 50 petabytes of its own records to establish what its models did on the internet during training and evaluation. The company says the review will take months, and press coverage citing OpenAI has put the cost at around half a million dollars a day. That is root cause analysis with an effectively unlimited budget and complete records. Most businesses investigating a recurring problem have the opposite: a meeting, three recollections and no log.
Why investigations stop early
The first explanation that fits ends the search. Someone offers a cause that accounts for what everyone remembers, the room relaxes, and the meeting moves to actions. Nobody checks whether that cause explains the other four times this happened, because the other four times are not in front of anyone.
Two habits make this worse. Investigations that implicitly ask who tend to stop as soon as a name appears, since the answer feels complete. And human error is treated as a terminal cause rather than as the point where the real question starts, which is why the person could make that error at all, and why the system accepted it.
A useful test before closing an investigation: does the proposed cause explain every occurrence, including the ones you did not investigate? If you cannot answer because you have no record of the others, that is the finding.
What AI contributes, and what it fabricates
| Task | How well it works | What to watch |
|---|---|---|
| Building a timeline from records | Very well, and nobody does it by hand | Timestamps in different zones |
| Finding similar past incidents | Well, if the tickets exist | Matching on vocabulary, not on mechanism |
| Spotting what is missing | Reasonably, when asked directly | It needs to be asked, every time |
| Naming the cause | Produces a fluent answer regardless | Sequence presented as causation |
| Explaining intent | Invention, every time | Why someone did something is not in a log |
Naming the cause is where most of the damage happens. Given a sequence of events, a model will supply a causal story linking them, because that is what the text it learned from does. The story will be coherent and it will be delivered with the same confidence as the timeline above it, which is the reviewing problem described in quality assurance when AI writes the first draft.
The practical response is to split the request. Ask for the timeline and the gaps, accept that output, then do the causal reasoning yourself with the timeline in front of you. Asking for both in one prompt gets you a conclusion with a timeline arranged to support it.
Feed it the raw records
Give the model ticket history with timestamps, system logs covering the window, the email and chat thread, and the procedure as actually written. Whatever summaries exist have already had the unexplained details stripped out by the person who wrote them, and those details are frequently the whole investigation.
Withhold your theory. A model told what you suspect will find support for it, which feels like confirmation and is closer to compliance. Ask the open question first, read what comes back, then test your theory against it as a second step.
Where the records live in five systems that were never meant to talk, that is the familiar obstacle covered in breaking down data silos, and it is worth solving once rather than reassembling by hand during every incident.
Ask what the evidence cannot settle
Three questions keep an investigation open longer than it would otherwise stay. What would have had to be different for this not to happen? Which of those things does anyone control? And what did we expect to catch this, that did not?
The third one is the most productive and the least asked. Almost every recurring problem passed through a control that was supposed to stop it. A check nobody runs, an alert routed to a mailbox nobody reads, an approval that has become a formality. Finding the control that failed usually explains the repetition better than finding the trigger does.
Expect at least one answer of the form we assumed somebody else checked that. Those are worth writing down even when they are uncomfortable, because they are the ones that produce a durable fix rather than a reminder.
Fixes that survive the attention fading
Sort proposed actions by what they change. Retraining, reminders and asking people to be careful change what an attentive person does on a good day. A changed default, a required field, a removed permission, a reordered step or an automated check changes what happens when nobody is paying attention, which is when the incident occurred.
If the only actions on the list are in the first group, the investigation found a trigger and not a mechanism. That is also the difference between fixing a process and speeding one up, which we set out in what happens when AI scales a broken process.
Make the next investigation cheaper
The reason OpenAI can conduct a review at this scale is that the records existed. Most of what makes a small-business investigation painful is decided months earlier, by what gets logged and whether anyone can retrieve it.
After each investigation, note the one thing you wished had been recorded and start recording it. Over a year that converts a guessing exercise into a reading exercise. Writing the procedure down in the first place, as described in standard operating procedures, also gives you something to compare the actual sequence against, which is often where the gap becomes obvious.
Recurrence is the measure that matters. Counting investigations closed tells you about meeting throughput, and a simple count of how many problems came back is the number worth putting on a dashboard.
Frequently Asked Questions
Can AI do root cause analysis?
It can do most of the work that surrounds root cause analysis and very little of the analysis itself. Reading across logs, tickets, emails and chat history to assemble a timeline is genuinely useful and is the part that normally does not get done. Deciding which event in that timeline caused the others is a judgment about mechanism, and a language model will produce a fluent, plausible answer whether or not the evidence supports one.
What is the biggest mistake in root cause analysis?
Stopping at the first explanation that fits. The first plausible story ends the search, and in most organisations it also ends the meeting, because a named cause feels like progress. The discipline that separates a real investigation from a tidy one is continuing to look after you have an answer you like, and checking whether the proposed cause explains every instance or only the one in front of you.
Does the Five Whys technique still work?
It works when someone recorded what happened. Five Whys is a prompt to keep asking, and it assumes each answer can be checked against evidence. Applied to an incident nobody documented, it produces five layers of recollection, which tends to terminate at human error because that is where memory runs out. The technique is not the constraint; the record is.
What should I give an AI model to investigate an incident?
Raw records rather than summaries. Ticket history with timestamps, system logs for the window, the email and chat thread, and the relevant procedure as written. Summaries have already had the unexplained details removed by whoever wrote them, and those details are usually the investigation. Ask for a timeline first and withhold your own theory, because stating it tends to produce agreement rather than analysis.
How do I stop the same problem recurring?
Check whether the fix changes what happens by default or only what a careful person does. Retraining, reminders and increased attention address the instance and leave the mechanism intact, which is why the same incident reappears after the attention fades. A change to the system, the form, the permission, the order of steps or the default value is the kind of fix that holds when nobody is watching.
Stop investigating the same problem twice
We help Canadian businesses assemble the records an investigation needs, run the analysis without the guesswork, and pick fixes that hold.
Related Articles
Grant Writing With AI: What It Does Well
Standard Operating Procedures an AI Can Follow
OpenAI’s New Advice: Delete Half Your Prompt
Ajan leads the ChatGPT.ca team: 200+ custom GPT builds and automation projects for 50+ businesses across 20+ industries. Based in Markham, Ontario. PIPEDA-compliant solutions.