Ask Your AI What It Is Likely to Get Wrong
A Goldman Sachs partner building the bank's own AI tools says the system told his team, without being asked, that it is better at sounding thorough than being thorough. Take the account as reported rather than confirmed. Keep the line anyway, because it describes the actual failure mode of AI in professional work more precisely than most formal guidance manages, and it points at something free that almost nobody does.
Why polished work passes
Nobody reviewing a document reads every line with equal attention. You look for the markers: clear structure, complete sections, a confident tone, no obvious holes. Those markers were useful because they correlated with somebody having actually done the work. A person who could not be bothered generally produced something that looked like it.
AI produces the markers at essentially zero cost. The correlation broke and the habit did not, so work now clears review because it resembles work that would have cleared review. That is a different problem from AI being wrong, and it is harder to catch, because everything your eye was trained to check comes back clean.
The prompt worth adding
After any substantive output, add one line:
What in this are you least confident about, and what did you assume that I did not tell you?
You will usually get something useful. Not deep self-knowledge, and it would be silly to describe it that way. What you get is a competent summary of known failure modes applied to your specific request, which turns out to be worth having.
| Ask | What it surfaces |
|---|---|
| What are you least confident about? | The two or three claims to verify first |
| What did you assume? | Context you never supplied, filled in silently |
| What would you need to do this better? | The document or number you forgot to give it |
The middle row is the one people find most useful. Assumptions are invisible in finished text, because a good writer states them as though they were established, and AI writes like a good writer by default.
Read the caveats before the work
Order matters more than it sounds. If you read the output first, you form an impression of competence and then read the caveats as minor footnotes to something you have already accepted.
Read the caveats first and your attention goes to the right places. A fifteen-minute review lands on the two claims that matter instead of spreading itself evenly across twelve pages, most of which were fine. Same time spent, considerably better result, and no new tooling required.
It does not replace knowing your own business
Be clear about the limit here. A model can tell you where it typically goes wrong. It cannot know that your biggest client dislikes a particular phrasing, that the figure came from a price list you superseded in June, or that the rule changed in your province last spring.
So this directs attention rather than replacing judgment. Used as a substitute for review it is worse than not asking at all, because a list of caveats can create the feeling that the work has been checked. That is the same trap as approval processes that become a formality, which we wrote about in human-in-the-loop turning into a rubber stamp.
The junior staff problem underneath
The other half of the Goldman comment was a worry about cognitive atrophy in junior staff. That is worth separating from the tooling question, because it is slower and harder to reverse.
Someone who produces confident-sounding work without doing the underlying thinking does not develop the judgment to tell the difference later, and that compounds. We have covered the two halves of this before: measurable deskilling in people who lean on AI, and what happens to the pipeline when junior work disappears. The caveats habit helps a little here too, because asking a junior to work through the flagged uncertainties is a genuine exercise in judgment rather than a formatting task.
One line, no budget
Most advice about managing AI quality involves a policy, a tool, or a process. This one is a sentence you add to a prompt, and it costs nothing to try this afternoon on the next piece of work that matters.
The reason it works is not that the model is being honest with you in any meaningful sense. It is that you are asking a different question, and a different question directs attention differently. Given that the core problem is polished output passing an inattentive review, redirecting attention is most of the fix. The rest is the thing that was always required, which is somebody who understands the work reading it properly, an argument we made at more length in the workslop quality problem.
Frequently Asked Questions
What was said?
A Goldman Sachs partner working on the bank’s own AI tools described the system warning his team, unprompted, that it is better at sounding thorough than being thorough. The same executive raised concerns about cognitive atrophy in junior staff who lean on it. Treat the account as reported rather than verified. The line is worth carrying regardless, because it names the specific failure mode of AI in professional work more precisely than most formal guidance does.
Why does sounding thorough matter so much?
Because reviewers check for the signals of diligence rather than the diligence itself, and they always have. Clear structure, complete sections, confident tone, no obvious gaps. Those markers used to correlate with someone having done the work, which is why they became useful shortcuts. AI produces the markers perfectly at almost no cost, so the correlation broke while the habit of trusting it did not. Work now passes review because it looks like work that would have passed review.
Can a model actually tell you its own weaknesses?
Partially, and more usefully than most people expect. Ask what it is likely to get wrong on a specific task and you generally get a sensible list: where it may lack current information, where it is guessing at context, which claims it is least confident about, what it would need from you to do better. That is not genuine self-knowledge in any deep sense. It is a decent summary of known failure modes applied to your request, and it costs one extra prompt.
How do I use this in practice?
Add one line after any substantive piece of work: what in this are you least confident about, and what did you assume that I did not tell you? Read the answer before you read the output. It reorders your attention, so instead of scanning polished text for errors you go straight to the two or three places the model has flagged. Most people find the assumptions question more revealing than the confidence question, because the assumptions are invisible in the finished text.
Is this a substitute for reviewing the work?
No, and treating it as one would be worse than not asking. A model can tell you where it typically goes wrong. It cannot know that your biggest client hates the phrasing in paragraph three, that the figure it used is from a superseded price list, or that the regulation changed in your province last spring. Use it to direct attention, not to replace judgment. The value is that it makes a fifteen-minute review land in the right places rather than spreading evenly across the document.
Catch polished-but-shallow before your client does
We help Canadian teams build practical review habits for AI-assisted work, so quality does not depend on how confident the output sounds.
Related Articles
AI Contract Review: What It Catches and What It Misses
AI Email Writer Guide 2026: The Tools That Draft and Send
The AI Formula in Your Spreadsheets Is Being Removed
Ajan leads the ChatGPT.ca team: 200+ custom GPT builds and automation projects for 50+ businesses across 20+ industries. Based in Markham, Ontario. PIPEDA-compliant solutions.