Skip to main content
Automation6 min read

A Benchmark Asked AI Agents to Run a Shop

September 3, 2026By Ajan Kanagalingam

Qwen published a benchmark this week testing whether AI agents can handle the work of operating an online business end to end. They did badly. That is a useful counterweight to a year of agent marketing, and it needs three qualifiers before anyone quotes it: the shop was simulated, a benchmark measures tasks somebody chose, and Qwen also sells models, so the evaluator has an interest in where the bar sits.

Read the source before the number

A benchmark published by a company that sells models is not disqualified, but it is not neutral either. Evaluators choose the tasks, and task choice determines the result more than anything else does.

So treat the headline as a direction rather than a measurement. That is the same caution we set out in why AI benchmarks mislead, and it applies whether a result flatters the technology or does not. A finding that confirms your scepticism deserves the same scrutiny as one that does not.

Why it is still worth reading

Because the failure shape matches what businesses actually report, and that agreement is more informative than either source alone.

Agents do well on a bounded task and degrade across long chains. An error at step three does not announce itself; it propagates through steps four to nine, each of which behaves reasonably given the wrong input it received. By the end the output is confidently wrong and the cause is buried several decisions back.

Anyone who has run an agent on real work recognises that description. A benchmark reaching the same conclusion is worth something precisely because it arrived independently.

Works wellDegrades
Retrieve, fill, move between two systemsNine steps with no checkpoint
Errors visible within a step or twoErrors that compound quietly
One decision, clearly scopedMany decisions, unsupervised

This is not an argument against agents

Worth being clear, because the easy read is that agents do not work.

Bounded agent work is genuinely valuable. Retrieve information, fill a form, move data between two systems, draft a reply and queue it for review. Those are short jobs where a mistake is visible immediately and costs a minute to fix.

What the benchmark tested is something else entirely: running a business unsupervised across many steps and many decisions. Failing at the second says nothing bad about the first, and conflating them is how businesses talk themselves out of automation that would have worked.

Find the compounding step

Here is the practical response, and it is one question rather than a project.

If a workflow has nine steps, do not ask whether an agent can do all nine. Ask which step is the one where being wrong quietly ruins the following eight. There is almost always exactly one, and everybody who works on the process can name it immediately.

Put a human there. It usually costs a minute and it converts an unbounded failure into a caught one, which is most of the difference between agents being useful and agents becoming a story about a project that did not work. Keeping a log of what the agent actually did, rather than what it reported, is the other half, and that is covered in watching what AI does rather than what it says.

The question for vendors

How many steps does this run before a human sees anything, and what happens when step three is wrong?

A vendor who has thought seriously about reliability answers with specifics: where the checkpoints are, how retries work, how an error surfaces rather than propagating. One who has not will steer back to the demo, which invariably shows a clean run on a well-behaved example.

The demo is not the deployment, and the entire gap between them is what happens when something goes wrong midway. That gap is also where most of the organisational reasons behind AI projects failing tend to live.

Frequently Asked Questions

What was published?

Qwen released an e-commerce benchmark that tests whether AI agents can handle the work of operating an online business end to end, and reports that they perform poorly at it. Three qualifiers belong on that. The environment is simulated rather than a real shop. It is a benchmark, so it measures performance against tasks somebody chose. And Qwen also sells models, which means the evaluator has a commercial interest in where the bar sits. None of that makes the result wrong, and all of it should shape how much weight you put on the number.

Why is the result still worth attention?

Because the shape of the failure matches what businesses actually report. Agents handle a bounded task well and degrade over long chains, where an early mistake compounds silently through every subsequent step. That pattern shows up in a benchmark and in practice, which is more persuasive than either on its own. The specific score matters much less than the fact that end-to-end operation is where things break, since that is precisely what agentic products are marketed on.

Does this mean agents are not useful?

No, and reading it that way would cost you real value. Bounded agent work is genuinely good: retrieve information, fill a form, move data between two systems, draft a reply and queue it for review. Those are three-step jobs where an error is visible immediately. What the benchmark tests is a very different thing, which is running a business unsupervised across many steps and many decisions. Failing at the second says nothing bad about the first.

How should this change what we deploy?

Keep the chains short and put a checkpoint where a mistake would compound. If a workflow has nine steps, the question is not whether an agent can do all nine, it is which step is the one where being wrong quietly ruins the following eight. Put a human there. That usually costs a minute and converts an unbounded failure into a caught one, which is the difference between agents being useful and agents being a story you tell about a project that did not work.

What should we ask vendors about this?

How many steps does this run before a human sees anything, and what happens when step three is wrong? A vendor who has thought about reliability will answer with specifics about checkpoints, retries, and how errors surface. One who has not will redirect to the demo, which usually shows a clean five-step run on a well-behaved example. The demo is not the deployment, and the gap between them is almost entirely about what happens when something goes wrong midway.

Put the checkpoint where it matters

We help Canadian businesses design agent workflows that stay short, surface errors early, and keep a person on the step that compounds.

Related Articles

Automation

AI Agents Just Ran a Full Drug-Discovery Loop. What Autonomous AI Means for Your Business

May 29, 2026Read more →
Automation

How to Automate Legacy Desktop Apps with AI Agents in 2026

Mar 26, 2026Read more →
Automation

AI Agents Are Going Mainstream in 2026, What Canadian Businesses Should Do Now

Mar 2, 2026Read more →
AK
Ajan Kanagalingam
Founder & ChatGPT Consultant, ChatGPT.ca

Ajan leads the ChatGPT.ca team: 200+ custom GPT builds and automation projects for 50+ businesses across 20+ industries. Based in Markham, Ontario. PIPEDA-compliant solutions.

Stay ahead of AI in Canada

Weekly case studies, new tools, and ROI playbooks for Canadian SMEs. One email, zero spam.