AI Writes Three Times More and Says Less
Head-to-head testing on LM Arena turned up a result that will feel familiar to anyone who uses AI daily: one leading model produced roughly three times as much text as its peers, while scoring lower on linguistic complexity. More words, simpler words. Length had stopped tracking substance. If you have ever asked an AI a direct question and received four paragraphs that restate the question, hedge, list considerations you did not ask about, and then summarize themselves, you have met the problem. It is costing you more than you think.
Why AI pads
It helps to understand that this is not a bug someone forgot to fix. Models were trained to be helpful and thorough, and thoroughness is easy to imitate by adding words. During training, human raters tended to prefer longer answers, so length quietly got rewarded. And a model with no strong view about what matters hedges by covering everything. The result is an answer optimized to look complete rather than to be useful, which is a subtle but expensive difference. Padding is not the model failing. It is the model succeeding at the wrong target.
The three costs
Only one of these shows up on an invoice, which is exactly why the other two go unmanaged.
| Cost | Why it hurts |
|---|---|
| Money | Usage-based plans charge per word generated |
| Attention | Your team reads padding to find the point |
| Judgment | Long, confident answers get questioned less |
That last row is the one worth sitting with. Length reads as effort, and effort reads as authority. A padded answer is not just slower to read, it is more likely to be accepted without scrutiny, which is a quieter cousin of the workslop problem. And the reading time is a direct contributor to the AI management tax we measured recently.
Five instructions that actually work
Because padding is the default, you have to be explicit. Set a hard length ceiling with a number, not a vague request to be brief, since "concise" means nothing to a model that already thinks it is being concise. Ask for the conclusion first and the reasoning after, so you can stop reading once you have what you need. Tell it not to restate your question and not to summarize at the end. Ask for a shaped output, a table, a ranked list, a decision, because structure resists padding in a way prose does not. And when you find a phrasing that produces tight output, save it into your shared prompts, the same discipline as any good business prompt library, so nobody has to rediscover it.
Do not over-correct
Shorter is not automatically better, and forcing terseness on genuinely complex work causes its own failures. A model squeezed too hard will drop the caveat that mattered, or compress its reasoning into something you cannot verify. The target is not minimum words, it is no wasted words. Here is a test that takes ten seconds: read the output and ask which sentences you could delete without losing anything. If the answer is most of them, you have a verbosity problem worth fixing in your prompts. If cutting anything would lose real information, the length was earned. Judge output by what it tells you, not by how much of it there is.
Frequently Asked Questions
What did the testing show?
Head-to-head evaluation on LM Arena found one leading model producing roughly three times as much text as its peers while scoring lower on linguistic complexity. In plain terms: more words, simpler words. That is not automatically bad, since simple language is often a virtue, but the combination is telling. It means length is not tracking substance. You are getting a longer answer that carries about the same amount of thinking, and in some cases less, because the point is now buried in padding rather than stated plainly.
Why do AI models write so much?
Partly because they were trained to be helpful and thorough, and thoroughness is easy to imitate by adding words. Partly because human raters, during training, tended to prefer longer answers, so length got quietly rewarded. And partly because a model with no strong view about what matters hedges by covering everything: restating your question, adding caveats, listing considerations you did not ask about, then summarizing what it just said. None of that is deception. It is what producing an answer looks like when the system is optimizing for looking complete rather than being useful.
What does verbosity actually cost my business?
Three things, and only one of them shows up on a bill. The obvious cost is money: if you use AI through an API or a usage-based plan, you pay per word generated, so three times the output is roughly three times the cost for the same answer. The bigger cost is attention: every padded paragraph is time your team spends reading to find the point. And the subtlest cost is judgment, because length signals effort. A long, confident answer feels more authoritative than a short one, which makes it less likely anyone questions it.
How do I get shorter, better output?
Be explicit, because the default is padding. Set a length ceiling in the prompt, a specific number of words or bullets rather than a vague "be brief". Ask for the conclusion first and the reasoning after, so you can stop reading once you have what you need. Tell it to skip restating your question and to drop the closing summary. Ask for the answer in a shaped format, a table or a decision, rather than prose, since structure resists padding. And build these into your saved prompts so nobody has to remember them every time.
Is shorter always better?
No, and over-correcting causes its own problems. Genuinely complex work needs room, and a model forced to be terse will sometimes drop the caveat that mattered or compress reasoning into something you cannot check. The goal is not minimum words, it is no wasted words. A useful test: read the output and ask which sentences you would delete without losing anything. If it is most of them, you have a verbosity problem. If cutting anything loses real information, the length was earned and you should leave it alone.
Get AI output your team can actually use
We help Canadian businesses set prompt standards that cut padding, lower usage costs, and put the answer in the first line instead of the fourth paragraph.
Related Articles
“Workslop”: When AI Output Creates More Work
Your Team Spends Hours a Week Babysitting AI
You Are Already Paying for AI You Never Use
AI consultants with 100+ custom GPT builds and automation projects for 50+ Canadian businesses across 20+ industries. Based in Markham, Ontario. PIPEDA-compliant solutions.