NuWayBiz Solutions
ai generated content quality

The AI wrote it in ten seconds. Checking it takes an hour.

AI made producing work almost free. Checking that work costs exactly what it always did, so everything now piles up at the one step nobody staffed. Forty percent of US desk workers got AI-generated work that looked finished and wasn't, in a single month, and each one took about two hours to sort out. The fix is boring, and it starts before anything gets written.

Painterly editorial illustration: a great mound of identical small pale cast forms spills from a wide brass chute into cool slate-blue shadow, heaped and untouched. At the right, a warm lamp lights a worn wooden inspection bench where exactly one form is held in the jaws of a brass gauge. A single red form lies unexamined among the many. No people, no legible text.ChatGPT Images 2.0

Same prompt, 3 AI models — swipe to compare. Showing 1 of 3.

Made with ChatGPT Images 2.0 (GPT Image (web))view prompt
Prompt

Create an editorial magazine illustration in a hand-painted style with visible soft brushstrokes and subtle oil-painting texture. NOT photoreal, NOT a 3D render. Palette: cool slate-navy and warm cream with selective deep cobalt blue accents with a single bold red accent used as a deliberate callout, cinematic 16:9 widescreen. Composition: a great mound of many identical small pale cast forms spilling from a wide brass chute into cool shadow, heaped and untouched, and in the foreground a small pool of warm lamplight falling on a worn wooden inspection bench where exactly one of those forms is held in the jaws of a precise brass gauge, one single small red form lying unexamined among the many in the mound, cool raking light across the heap and warm directional light on the bench, a quiet sense of far more being made than could ever be looked at, cream and slate-navy tones with deep cobalt accents and one small red detail, no people, no text, no letterforms, no numbers, no numerals, no symbols, no paper, no documents, no charts, no screens, no signage. No people, no readable text, no logos.

HERO for article-24 (LCP), owner web-lane pass (NU-107), replacing the provisional local Flux take which moves to the gallery. Native 1672x941 centre-trimmed to 1672x940 for 16:9. Why this one reads the article: the forms are IDENTICAL and machined, which is the thesis — uniformly well-formatted and uniformly confident, so a wrong one looks exactly like a right one — and exactly one is in the gauge while the red one lies unexamined in the heap. The Gemini hero take was rejected: baked-in sparkle watermark lower-right (same call as art19/art22/art23), and its forms read as generic pebbles rather than identical castings, which loses the whole metaphor.

Last week I had an AI write up a traffic report for a website. Simple job: pull the last year, tell me what happened.

It came back looking excellent. Clean tables, a tidy summary at the top, numbers carried out to two decimal places. The headline finding was that traffic had fallen off a cliff.

It hadn't.

Most of the traffic in that report belonged to a completely different website. An older site had lived on that domain years earlier, and it had been quietly reporting into the same analytics account the whole time. Of the sessions in the report, 963 of 1,106 belonged to it rather than to the site anyone was asking about. Somebody else's visitors, from a site that no longer exists, faithfully counted and beautifully presented.

Here is what bothers me about it. Nothing in that report was sloppy. Every number was really in the data, pulled correctly, added up correctly, formatted correctly. They were just describing the wrong website.

It took three passes to catch. On the second pass, the corrected version made the same mistake again in a different spot, which I only found because someone went looking for it specifically.

If your team is buried under AI-generated work that nobody has the hours to properly check, you are not imagining it, and it is not a discipline problem on your end.

40% of US desk workers received AI-generated work that looked finished and wasn't, in a single month

Making got cheap. Checking cost exactly what it always did.

Writing a first draft of something used to take forty minutes. Now it takes about ten seconds. That is a real gain and I am not going to pretend otherwise.

Reading that draft closely enough to decide whether it is right takes as long as it ever did. Longer, actually, and there is a specific reason why.

When a colleague hands you a draft, you know them. You know they are excellent on the numbers and vague on the timeline, so you read hard in one place and skim the rest. That shortcut is most of what makes review survivable. AI output takes it away from you. It is uniformly polished and uniformly confident, so the wrong sentence looks precisely like the right one and the mistakes are scattered evenly through it instead of clustering where a tired person would put them.

Researchers at BetterUp Labs and Stanford's Social Media Lab put a number on how common this has become. In a September 2025 survey of full-time US desk workers, 40% said they had received AI-generated work in the previous month that looked good and lacked substance (their term for it is workslop), and sorting out each instance took about two hours. The write-up ran in Harvard Business Review.

Two hours. For one item. That is not a rounding error on a productivity gain, that is the gain, handed to somebody else.

Every hour AI saves at the making step lands on somebody at the checking step. That is the same hour. It just moved desks.

The checking is the expensive half, and nobody budgeted for it

Workday looked at the same problem from the company's side. Their January 2026 research found that nearly 40% of AI time savings are lost to rework — correcting errors, rewriting content, verifying outputs — and that only 14% of employees consistently get a clear positive result out of AI at all.

One thing to know before you apply that to yourself: every person in that study works at a company doing over $100 million a year. That is a much larger business than yours, and it matters, because it cuts the way you would not expect.

A company that size has slack. Somebody's afternoon absorbs the rework and it never shows up as a line item. You do not have that person. When the checking lands in a six-person shop it lands on the owner, at night, on top of everything else.

The boring reason it takes so long

To check anything quickly, you need two things: you need to know what a correct answer would look like, and you need something trustworthy to hold it up against.

Most businesses have written down neither. So checking collapses into reading it over and seeing whether it feels about right, which is proofreading with extra anxiety.

Go back to that traffic report for a second. Nobody involved was careless. What was missing was the one sentence that would have killed the error instantly: this report covers traffic to the current site, from the day its tracking was installed, and nothing before it. There was no line in the sand, so there was nothing for a wrong number to fail against.

This is the same thing we keep going on about, arriving from a direction that surprised me. Connected, clean, trustworthy data is usually sold as the thing that makes AI work. It is also the thing that makes checking AI cheap, and that turns out to be where the hours actually go.

It is a cousin of something we wrote about a few weeks back: a stopped automation looks exactly like a quiet week. Silence reads as fine. So does a confident paragraph.

Painterly editorial illustration: a single pale cast form seated exactly into a precisely cut recess in a heavy machined brass master template on a worn wooden bench, the fit clean and obvious under a low raking lamp, cobalt shadow pooling around it, one small red detail on the template edge. No people, no legible text.
Made with ChatGPT Images 2.0 (GPT Image (web))view prompt
Prompt

Create an editorial magazine illustration in a hand-painted style with visible soft brushstrokes and subtle oil-painting texture. NOT photoreal, NOT a 3D render. Palette: cool slate-navy and warm cream with selective deep cobalt blue accents with a single bold red accent used as a deliberate callout, cinematic 16:9 widescreen. Composition: a single pale cast form seated exactly into a precisely cut recess in a heavy machined brass master template resting on a worn wooden bench, the fit clean and obvious, a low warm lamp raking across the surface so the seam reads as perfect, calm cobalt shadow pooling around the bench, one small red detail on the edge of the template, a quiet sense of knowing instantly whether something is right, cream and slate-navy tones with deep cobalt accents, no people, no text, no letterforms, no numbers, no numerals, no markings, no symbols, no paper, no documents, no charts, no screens, no signage. No people, no readable text, no logos.

MID for article-24, owner web-lane pass (NU-107). ⛔ THIS CONCEPT HAD BEEN ABANDONED: four local seeds all drifted to a paper sheet or document and grew baked-in numerals ('17','5','4','IIM') and signatures, because 'master template' reads to a model as 'printed document'. The web lane solved it by keeping the template unambiguously BRASS AND MACHINED — no paper anywhere in the frame, so no text could grow. That is the durable lesson: never build a concept on an object the model renders as paper. Verified text-free at 2x across the top and bottom strips, the bands where these lanes put signatures.

Three rules that make checking fast

None of this is an argument for using less AI. The early wins are real and we have written about where they come from. This is about not handing the bill to whoever sits downstream.

Decide what right looks like before anything gets written. One sentence, written first, saying what the output covers and what a correct version would contain. It takes thirty seconds and it is the single highest-leverage thing on this list. A reviewer holding that sentence checks in minutes. A reviewer without it has to work backwards from the output to guess what was intended, which is most of the two hours.

Check against a number, not a feeling. If somebody can tell you what last month's figure actually was, verifying a summary takes a minute. If the only way to find out is to rebuild it from four exports, every check becomes a small project and, realistically, nobody does it. That is the one honest number problem, and it is why it is worth solving before the AI arrives rather than after.

Only point it at work you could verify quickly by hand. This is the rule that saves you from the worst version. If checking the output takes longer than doing the task yourself would have, you have not automated anything. You have bought a slower version of the job and added a step where a mistake can hide.

Worth saying plainly: this is the ops-side twin of something we wrote for creative teams about using AI without producing slop. That one is about the work you make. This one is about the work that lands in your inbox from everybody else.

Where you actually are right now

Two or more of these and the checking step is where your time is going, not the making step. That is fixable, and the fix starts before anything gets written rather than after.

The businesses handling this well can answer "how would we know this is right?" without a pause, because they sorted out what they trust before they started pointing software at it. The tools they use are the same ones you have.

If your team is producing more than anyone can check, that is worth an hour of conversation. Start a no-pressure conversation about which of your numbers we would make trustworthy first — the one that would take the most checking off your plate.

Cheers, from the boring side of the business,

— Brian

P.S. While writing this, I went to double-check the workslop study's own figures. BetterUp's landing page says the survey covered 1,150 people and that each incident takes two hours to resolve. BetterUp's own blog post, about the same study, says 1,004 people and one hour fifty-one minutes. Same company, same research, two pages, two answers. I have used the round number and told you where it comes from, because I genuinely do not know which is right. Two hours of somebody's afternoon, and I could not verify it in twenty minutes.

Want help applying this to your business? Start a no-pressure conversation →

Frequently asked questions

Why does checking AI work take longer than checking a person's work?
With a colleague's draft you know where they tend to be weak, so you read hard in two places and skim the rest. AI output does not give you that. It is uniformly well-formatted and uniformly confident, so a sentence that is wrong looks exactly like a sentence that is right, and the errors are scattered rather than clustered. That removes the shortcut experienced reviewers rely on, and you end up reading every line at the same level of attention. Researchers at BetterUp and Stanford's Social Media Lab found that 40% of US desk workers received AI-generated work that looked finished and wasn't in a single month, and that sorting out each one took about two hours.
How do I stop my team drowning in AI-generated output?
Cap it at the front rather than the back. Before anything gets generated, write one sentence saying what a correct version would contain and what it covers, because a reviewer with that sentence can check in minutes and a reviewer without it has to reconstruct the intent from the output itself. Then only point AI at jobs where you could verify the result quickly by hand. If verifying takes longer than doing the task yourself, you have bought a slower version of the job with extra steps.
Is AI actually saving my business time?
Some of it, and less than the headline. Workday's January 2026 research found that nearly 40% of AI time savings are lost to rework, which they define as correcting errors, rewriting content and verifying outputs, and that only 14% of employees consistently get clear positive net outcomes. Worth knowing before you apply that to yourself: every person in that study works at a company with over $100 million in annual revenue. Our read is that a smaller business feels this harder rather than softer, because a big company has people whose job can quietly absorb the rework and a shop with eleven staff does not.
What does checking AI output have to do with my data?
Checking is fast when you have something trustworthy to compare against and slow when you don't. If someone can tell you what last month's number actually was, verifying an AI-written summary takes a minute. If the only way to find out is to rebuild it from four exports, then every check turns into a small project and nobody does it. That is the same foundation question we keep coming back to, arriving from a different direction: connected, clean, trustworthy data is what makes verification cheap.
Brian, founder of NuWay Biz Solutions

Brian

Founder, NuWay Biz Solutions. Practical AI implementation for small businesses. More about NuWay →