Skip to main content
Enterprise AI9 min read

AI Employees: What You Are Actually Buying

September 8, 2026By Ajan Kanagalingam

Products sold as AI employees arrive with a first name, a headshot, a job title and a monthly price that sits neatly below a salary. The framing does a lot of work, and it is worth taking apart before the comparison to payroll gets made. OpenAI published a figure this week that gives a sense of the upper bound: since June, it says, agent runtime across its research organisation has exceeded human labour, at roughly 3.1 agent-workdays for every human workday.

What the OpenAI figure does and does not show

Take the 3.1 seriously and read it carefully. It is OpenAI's own number about OpenAI's own research organisation, which is the most favourable environment for agent work that exists: the staff build the models, the tasks are software and research, and the tolerance for a failed run is high.

It also measures runtime converted into workday equivalents rather than output or value. Agents run in parallel and overnight, so a figure expressed in workdays inflates easily against humans who do neither. A firm doing client work with deadlines and liability sits a long way from that setting.

What it does establish is that agent labour at real scale is happening somewhere, not just in pitch decks. Treat it as the ceiling under laboratory conditions, and your own measurement as the number that matters.

The three products behind the label

Outbound sales. The most crowded category. The agent researches prospects, writes personalised messages, sends sequences and books meetings. It genuinely produces volume. Whether volume helps depends on whether your problem was not enough outreach or not enough good outreach, and for most small firms it was the second.

First-line support. Answers common questions from your documentation, handles simple account actions, escalates the rest. This is the most mature of the three and the easiest to measure, because deflection rate and customer satisfaction are already numbers you track.

Back-office operations. Scheduling, data entry, reconciliation, chasing missing information. Least glamorous, often the best return, because the work is high volume and the correct answer is unambiguous. It is also where verification matters most, which we went into in extracting data from PDFs.

What per-seat pricing hides

These products cost the vendor per token and charge you per seat, and that gap has to be managed somewhere. It usually appears as a usage cap in the contract, expressed in conversations, tasks or credits per month, with an overage rate that is easy to miss.

Work out your real monthly volume before the demo and ask what happens at three times that number, because a busy month is exactly when you want the thing working. The related trap of unpredictable consumption billing is covered in consumption pricing and unpredictable bills.

The second hidden cost is supervision. Somebody reviews the output, handles what the agent could not, and fixes what it got wrong. At five hours a week that is a meaningful fraction of a salary, and it belongs in the comparison rather than in the assumption that the product runs itself.

Four things the label promises and the product does not

Accountability. When a person makes an expensive mistake, someone answers for it and something changes. When an agent does, you have a log and a support ticket. That difference matters most in exactly the regulated and client-facing work people most want to automate.

Knowing when to stop. A junior member of staff who senses a conversation going wrong asks someone. An agent proceeds confidently unless you built the escalation rule, and you will not think of every case that needs one.

Relationships. In small business sales, the person on the other end often buys because they trust someone. An agent can carry the process and cannot carry that, and pretending otherwise annoys the customers you already have.

Noticing that the process changed. When a supplier changes a form or a regulation shifts, a person adapts and mentions it. An agent keeps following the old configuration until someone notices the outputs have quietly gone wrong, which argues for the monitoring approach in watch what it does, not what it says.

The number to ask for

One question separates these products more than any feature list: what fraction of real cases does it finish without a person. Research published this month found that only 2.6 percent of catalogued agent tools complete a recorded work task end to end, with most of the rest assisting instead, which we unpacked in only 2.6% of agent tools finish a whole task.

Assisting is a legitimate and valuable product. It is a different purchase from replacing a role, and the pricing rarely distinguishes between them. Ask for the completion rate on customers who look like you, and ask for a case that failed.

A four-week trial that settles it

1. Pick one workload with a countable volume. Inbound enquiries, invoices, appointment confirmations. Something you can count last month's of.

2. Run it in parallel, not instead. The agent handles the same queue your current process does, for four weeks, with nobody depending on it yet.

3. Record three numbers. Completed unaided, needed a person, and total review minutes. Those three settle the business case, and none of them come from the vendor.

4. Give it a scoped account before it touches anything live. Its own login, access to only the systems the task needs, an approval gate on anything irreversible, and a log. The framework is in three boxes for an AI agent.

Put the three numbers into the ROI calculator alongside the subscription and the supervision hours. Plenty of these products clear that bar comfortably. The ones that do not tend to fail on the review time rather than on the subscription, which is the line item the pricing page never mentions.

Frequently Asked Questions

What is an AI employee?

It is a marketing label rather than a technical category. Behind it sits a packaged AI agent configured for one job function, given a name and often an avatar, connected to a few of your systems, and sold per seat rather than per token. The common versions handle outbound sales prospecting, first-line customer support, and back-office operations such as scheduling or data entry. What varies most between products is how much of the task they finish without a person.

How much does an AI employee cost?

Advertised prices commonly run from a few hundred to a couple of thousand dollars a month per configured agent, which is why the comparison to a salary is made so often. The advertised figure rarely includes the whole cost. Ask about usage caps and overage rates, integration and setup, and how many hours a week one of your people will spend reviewing output and handling what the agent could not. That supervision time is real payroll and it belongs in the comparison.

Can an AI employee replace a person?

It can replace a set of tasks, and the honest question is what share of a role those tasks represent. Products in this category are strongest on high-volume, repetitive, well-defined work with a quick way to check the output. They are weakest at judgment under ambiguity, accountability for an outcome, and adapting when the process changes without anyone updating a configuration. Most businesses find the first year returns hours inside existing roles rather than removing positions.

How do I trial an AI employee properly?

Pick one workload with a countable volume, run it in parallel with your current process for four weeks, and record three numbers: how many items the agent completed unaided, how many needed a person, and how long the review took. Compare that against the same period handled the current way. Judge it on those three numbers rather than on how impressive the demo was, and insist on seeing a case that failed before you sign.

What questions should I ask an AI employee vendor?

Does it complete the task or assist with it, and who performs the last step? What happens when the input is unusual, and what fraction of real cases hit that path? What account does it use in our systems and what can that account reach? What are the usage limits and the overage rate? Vendors who have thought carefully about where their product hands back to a person answer all of these without hesitating.

Test it on your workload, not their demo

We run four-week parallel trials of agent products against real work from your business, and report completion rate, exception rate and supervision cost before you commit.

Related Articles

Enterprise AI

COBOL to Java Migration: What Does It Actually Cost in 2026?

Feb 24, 2026Read more →
Enterprise AI

OpenAI’s New Model Is Built to Use Your Computer

September 4, 2026Read more →
Enterprise AI

AI Limitations: What It Still Cannot Do for You

September 3, 2026Read more →
AK
Ajan Kanagalingam
Founder & ChatGPT Consultant, ChatGPT.ca

Ajan leads the ChatGPT.ca team: 200+ custom GPT builds and automation projects for 50+ businesses across 20+ industries. Based in Markham, Ontario. PIPEDA-compliant solutions.

Stay ahead of AI in Canada

Weekly case studies, new tools, and ROI playbooks for Canadian SMEs. One email, zero spam.