Before you buy an AI agent, answer these five questions
An AI agent decides for itself what needs doing in your business, then does it. That is the whole pitch, and it is also the whole risk. Gartner counted about 130 real ones in mid-2025, out of thousands of vendors selling agents. Five questions sort the real from the rebranded, and all five are about your business rather than the software.
ChatGPTSame prompt, 2 AI models — swipe to compare. Showing 1 of 2.
Made with ChatGPT (ChatGPT Images 2.0 (gpt-image-2))view prompthide prompt
Create an editorial magazine illustration in a hand-painted style with visible soft brushstrokes and subtle oil-painting texture. NOT photoreal, NOT a 3D render. Palette: cool slate-navy and warm cream with selective deep cobalt blue accents with a single bold red accent used as a deliberate callout, cinematic 16:9 widescreen. Composition: a single ornate brass key resting alone on a small cushion of deep cobalt velvet beneath a narrow shaft of warm directional light, and behind it a tall wall of many identical unmarked doors receding into soft shadow, one distant door faintly outlined in red, the key presented carefully and the doors left unexplained, cool raking light, a quiet sense of being handed the key to rooms nobody has shown you, cream and slate-navy tones with deep cobalt accents and one small red detail, no people, no text, no letterforms, no numbers, no numerals, no symbols, no paper, no documents, no charts, no screens, no signage. No people, no readable text, no logos.
HERO for article-23 (LCP), owner web-lane render, picked at review over the local Flux take that shipped in the first draft (now the gallery's second slide). Wins on the brand palette: the callout door reads a true RED here where Flux rendered it orange, the cushion is a real cobalt velvet with gold tassels, and the light shaft is warm cream against slate-navy instead of Flux's uniformly cool blue. Exactly 16:9 (1672x941). Verified text-free at full size. Reads the concept: the key is the product being sold, the identical unmarked doors are the buyer's own systems the seller has never seen, and the single red door is the irreversible action. Must equal featured_image_url; first key in the gallery.
"It's an AI agent. It handles your inbound automatically."
"Handles it how?"
"Automatically."
That exchange is a fair summary of the market right now, and it is not really the salesperson's fault. The word arrived faster than the meaning did.
Gartner, the technology research firm a lot of corporate software budgets get built on, has a name for the resulting mess. They call it agent washing: taking an assistant, a chatbot or an old automation tool, and reissuing it with the new label. Of the thousands of vendors selling agents, they counted about 130 selling a real one in mid-2025.
Gartner counted about 130 real agent vendors in mid-2025, out of thousandsYou do not need to work out which 130. You need five answers, and every one of them is a question about your business rather than a question about the software.
1. What does it do when nobody is watching it?
This is the question the label is hiding. Ask the vendor to name one thing their software does with no person approving it first.
If every answer comes back as it suggests, it drafts, it recommends, it surfaces, then you are being shown an assistant. Which is genuinely useful and worth paying for. It is just a different product at a different price, and you should know which one is on the invoice.
An assistant
- Waits for you to ask
- Hands the work back for you to approve
- Its mistakes stay inside your building
- Usually priced per person per month
An agent
- Decides on its own that something needs doing
- Does it, then tells you
- Its mistakes reach customers before you see them
- Usually priced per action, per outcome, or per run
Gartner's analyst on this, Anushree Verma, put the useful half of it more bluntly than a vendor ever will: “Many use cases positioned as agentic today don't require agentic implementations.”
Plenty of jobs are better done by something that asks first.
2. What can it do that I can't take back?
Everything an agent is allowed to touch falls into two piles, and the piles are not the same size.
One pile is reversible. It tagged the wrong lead, it filed a job under the wrong customer, it wrote a note nobody needed. You fix it in a minute and no one outside the business ever knew.
The other pile leaves the building. It emailed a customer. It issued a refund. It quoted a price. It cancelled tomorrow's appointment.
You cannot un-send an email to a customer, and the agent will not know it should have asked.
So get the list. The one that names each action it can take with no human in between, and where that action lands. Vendors lead with a longer and friendlier list, so ask for this one specifically.
Then say out loud which of those you are comfortable with on a Saturday when nobody is looking at a screen. That is the real permission you are granting, and it is worth writing into the contract rather than discovering later.
3. Which of my numbers does it read?
An agent acts on what it reads. That single fact is why this question sits in the middle of the list rather than at the end.
We have made this point about what AI-ready data actually means for a while, and about why the AI you already pay for stalled out. An agent raises the stakes on it. A dashboard built on numbers that disagree wastes your afternoon. An agent built on the same numbers sends the customer a bill for the wrong amount.
So ask which systems it reads from, and then ask yourself the harder half: would you personally bet on those numbers being right this morning?
If the honest answer is no, do that work first. It pays for itself whether or not you ever sign. Match the build to the mess, same as always.
4. What does this cost in my busiest month?
The software you already buy costs the same in February as it does in July. Agents mostly are not sold that way.
Intercom publishes its price plainly, which is more than most vendors do, so it makes a clean example: its Fin agent is $0.99 per outcome, charged once per conversation. An outcome counts when the customer confirms the issue is resolved, when the agent completes a workflow, or when the customer doesn't ask for more help after it responds.
Read that last one twice. Silence counts as success, which is a reasonable way to measure it and also exactly the number you want to watch.
Usage pricing is a perfectly reasonable way to buy this. It just asks one piece of arithmetic of you first: take your worst week last year, the one where the phone did not stop, and work out what that week would have cost at the quoted rate.
Then ask what happens when you hit the ceiling on your plan, because the answer is rarely nothing.
5. How would I know it did the job right?
We wrote recently that a stopped automation looks exactly like a quiet week. An agent inherits that blind spot and adds a worse one of its own.
An automation does one fixed thing every time its trigger fires, so at least you know what it was trying to do. An agent picked the thing. It can run perfectly, do exactly what it decided to do, and still have done the wrong thing, confidently, in your name.
There is a good measurement of how often an agent finishes the job at all. Researchers at Carnegie Mellon built a fake company, a simulated software business with its own internal websites and data, and turned the best agents available when they ran it loose on the work its staff would handle. The most competitive agent finished 30% of the tasks on its own. Their conclusion is worth keeping in your head during the sales call: simpler tasks often go fine, and longer multi-step ones are still beyond what these systems can do.
That number will improve. It has not improved enough for you to skip the check.
So ask what it produces that a person can look at afterwards. A record of what it decided and why, in language you can read on a phone, with a way to find the ones it got wrong before your customer does. A dashboard of how busy it has been will not tell you any of that.
When the answer is yes

Made with ChatGPT (ChatGPT Images 2.0 (gpt-image-2))view prompthide prompt
Create an editorial magazine illustration in a hand-painted style with visible soft brushstrokes and subtle oil-painting texture. NOT photoreal, NOT a 3D render. Palette: cool slate-navy and warm cream with selective deep cobalt blue accents with a single bold red accent used as a deliberate callout, cinematic 16:9 widescreen. Composition: one tall panelled door standing open in a long wall of identical closed doors, warm light spilling out of the opened doorway across a pale stone floor and revealing a calm ordinary room beyond, a single ornate brass key resting still on the threshold of the open door and not yet carried inside, the closed doors on either side receding into cool shadow, one small red detail on the frame of the opened door, soft warm directional light, a quiet sense of having looked inside before going in, cream and slate-navy tones with deep cobalt accents, no people, no text, no letterforms, no numbers, no numerals, no markings, no paper, no documents, no screens, no signage. No people, no readable text, no logos.
MID for article-23, owner web-lane render, picked over the local Flux take that shipped in the previous commit. Two reasons it wins. (1) REGISTER MATCH: it is the same model and the same painterly hand as the hero, so hero and mid now read as one commissioned pair instead of two different models bolted together — a mismatch that was accepted in the previous commit and should not have been. (2) It shows the room. The Flux take opened the door onto a flat wash of light; this one reveals an ordinary sunlit room with a table, a chair and a plant, which is the actual point of the section — you looked inside, and it is just a room. Also lands a genuine red callout on the doorframe plaque, echoing the hero's red door. Exactly 16:9 (1672x941), no crop needed, all four corners inspected at 3x and clean. The Gemini take was rejected for the usual baked-in sparkle watermark (lower right) plus a red accent that reads as a splotch rather than a considered detail. Local lanes rejected earlier: Flux seed 2312 rendered the open door as closed with an explicit copyright watermark, 2314 dropped the surrounding wall of doors, 2311 needed a 9% bottom crop to remove a cursive signature; the whole SDXL lane was off-brand throughout. All candidates retained in the Dropbox package.
None of this is an argument against buying one. There are jobs where an agent earns its keep quickly, and they have a recognisable pattern: the same decision, many times a week, against a rule you could write down on an index card.
Booking requests. First-pass triage on inbound. Chasing the same three pieces of missing information off every new job. Work where being right ninety-something percent of the time and flagging the rest is a genuine improvement on what happens today, which is that it waits until Thursday.
Verma's own advice on where to start says it better than we can: “They can start by using AI agents when decisions are needed, automation for routine workflows and assistants for simple retrieval.” Three different tools. Most of what gets pitched as the first one is doing the job of the other two.
The businesses that get value out of this are the ones that could answer the five questions before they signed.
Four or more and you are ready to have a serious conversation with a vendor. Fewer, and the gaps are your agenda for the next call: whichever one you couldn't tick is the one that will cost you.
Take the list to the next call. You will know inside ten minutes whether the person on the other end has thought about their product as hard as you are about to think about your business.
And if you get the answers and still cannot tell whether the thing is worth buying, send us what you were pitched and we will tell you what it actually is. Sometimes the answer is that you already own something that does the job.
Cheers, from the boring side of the business,

P.S. One more to try on the next demo, and it is the cheapest of the lot. Ask what the software does when it is not sure. A real answer sounds like a rule: it stops, it flags, it asks a human. If what comes back is a shrug, or that it uses its best judgement, you have just learned the most important thing on this list.
Want help applying this to your business? Start a no-pressure conversation →
Frequently asked questions
- What is the difference between an AI agent and an AI assistant?
- An assistant waits for you and hands its work back to you. You ask it to draft the email, it drafts the email, you decide whether to send it. An agent acts on its own: it decides that an email should go, writes it, and sends it, and you find out afterwards. The interesting part is that both can sit behind the same chat box and look identical in a demo. The question that separates them is what happens when nobody is watching, and it is worth asking plainly because the two are priced very differently.
- Is my business too small for an AI agent?
- The test is volume and repetition rather than headcount. An agent earns its place where the same decision gets made many times a week against a rule you could write down. A shop fielding dozens of booking requests a week has a genuine case; a shop fielding a handful will spend more time supervising the agent than it spends doing the job. That bar sits higher than the one we set for ordinary automation, and deliberately so: an agent has to be worth supervising, and that is a higher bar than worth switching on. What does scale down badly is the recovery cost. A small business feels one bad automated message to a customer far more than a large one does, which is why the question about what it can do irreversibly matters more, not less, at your size.
- How do I know if a vendor is selling a real agent or a rebranded chatbot?
- Gartner calls the rebranding "agent washing" and counted about 130 real ones in mid-2025, out of thousands of vendors claiming agentic solutions. You do not need to work out which 130. Ask the vendor to name one thing the software does without a person approving it first, and to show it happening in your own account rather than in a demo environment. If every answer comes back as "it suggests" or "it drafts" or "it recommends," you are looking at an assistant. That may well be worth buying, at assistant prices.
- What should an AI agent cost?
- Expect the pricing to work differently from the software you already buy. A seat licence costs the same whether the week is quiet or brutal, and agents usually are not sold that way: Intercom, for example, publishes a price of $0.99 per outcome for its Fin agent, charged once per conversation. That means the bill tracks your volume, so it peaks in the month you are busiest and least able to look at it. Before signing, work out what the bill would have been during your worst week last year, and ask what happens when you hit the ceiling.
- Do I need to fix my data before buying an AI agent?
- You need to know which of your numbers it reads, and whether those are right today. An agent acts on what it reads, so a wrong number stops being a reporting problem and becomes a message to a customer, a price, or a refund. This is the same point we make about the data foundation generally, with the stakes raised: reporting on bad data wastes your time, and acting on bad data costs you a customer. If the numbers the agent would read are ones you would not personally bet on, fix those first. That work is useful whether or not you ever buy the agent.



