DEV Community

Cover image for The AI wrote it in ten seconds. Checking it takes an hour.
Brian Pawl
Brian Pawl

Posted on • Originally published at nuwaybizsolutions.com

The AI wrote it in ten seconds. Checking it takes an hour.

Originally published on the NuWay Biz Solutions blog.

✦ Cover image: Made with ChatGPT Images 2.0 — we're transparent about AI. See the exact prompt on the original post.

Last week I had an AI write up a traffic report for a website. Simple job: pull the last year, tell me what happened.

It came back looking excellent. Clean tables, a tidy summary at the top, numbers carried out to two decimal places. The headline finding was that traffic had fallen off a cliff.

It hadn't.

Most of the traffic in that report belonged to a completely different website. An older site had lived on that domain years earlier, and it had been quietly reporting into the same analytics account the whole time. Of the sessions in the report, 963 of 1,106 belonged to it rather than to the site anyone was asking about. Somebody else's visitors, from a site that no longer exists, faithfully counted and beautifully presented.

Here is what bothers me about it. Nothing in that report was sloppy. Every number was really in the data, pulled correctly, added up correctly, formatted correctly. They were just describing the wrong website.

It took three passes to catch. On the second pass, the corrected version made the same mistake again in a different spot, which I only found because someone went looking for it specifically.

If your team is buried under AI-generated work that nobody has the hours to properly check, you are not imagining it, and it is not a discipline problem on your end.

40% of US desk workers received AI-generated work that looked finished and wasn't, in a single month

Making got cheap. Checking cost exactly what it always did.

Writing a first draft of something used to take forty minutes. Now it takes about ten seconds. That is a real gain and I am not going to pretend otherwise.

Reading that draft closely enough to decide whether it is right takes as long as it ever did. Longer, actually, and there is a specific reason why.

When a colleague hands you a draft, you know them. You know they are excellent on the numbers and vague on the timeline, so you read hard in one place and skim the rest. That shortcut is most of what makes review survivable. AI output takes it away from you. It is uniformly polished and uniformly confident, so the wrong sentence looks precisely like the right one and the mistakes are scattered evenly through it instead of clustering where a tired person would put them.

Researchers at BetterUp Labs and Stanford's Social Media Lab put a number on how common this has become. In a September 2025 survey of full-time US desk workers, 40% said they had received AI-generated work in the previous month that looked good and lacked substance (their term for it is workslop), and sorting out each instance took about two hours. The write-up ran in Harvard Business Review.

Two hours. For one item. That is not a rounding error on a productivity gain, that is the gain, handed to somebody else.

Every hour AI saves at the making step lands on somebody at the checking step. That is the same hour. It just moved desks.

The checking is the expensive half, and nobody budgeted for it

Workday looked at the same problem from the company's side. Their January 2026 research found that nearly 40% of AI time savings are lost to rework — correcting errors, rewriting content, verifying outputs — and that only 14% of employees consistently get a clear positive result out of AI at all.

One thing to know before you apply that to yourself: every person in that study works at a company doing over $100 million a year. That is a much larger business than yours, and it matters, because it cuts the way you would not expect.

A company that size has slack. Somebody's afternoon absorbs the rework and it never shows up as a line item. You do not have that person. When the checking lands in a six-person shop it lands on the owner, at night, on top of everything else.

The boring reason it takes so long

To check anything quickly, you need two things: you need to know what a correct answer would look like, and you need something trustworthy to hold it up against.

Most businesses have written down neither. So checking collapses into reading it over and seeing whether it feels about right, which is proofreading with extra anxiety.

Go back to that traffic report for a second. Nobody involved was careless. What was missing was the one sentence that would have killed the error instantly: this report covers traffic to the current site, from the day its tracking was installed, and nothing before it. There was no line in the sand, so there was nothing for a wrong number to fail against.

This is the same thing we keep going on about, arriving from a direction that surprised me. Connected, clean, trustworthy data is usually sold as the thing that makes AI work. It is also the thing that makes checking AI cheap, and that turns out to be where the hours actually go.

It is a cousin of something we wrote about a few weeks back: a stopped automation looks exactly like a quiet week. Silence reads as fine. So does a confident paragraph.

Painterly editorial illustration: a single pale cast form seated exactly into a precisely cut recess in a heavy machined brass master template on a worn wooden bench, the fit clean and obvious under a low raking lamp, cobalt shadow pooling around it, one small red detail on the template edge. No people, no legible text.

✦ Made with ChatGPT Images 2.0 — we're transparent about AI. See the exact prompt on the original post.

Three rules that make checking fast

None of this is an argument for using less AI. The early wins are real and we have written about where they come from. This is about not handing the bill to whoever sits downstream.

Decide what right looks like before anything gets written. One sentence, written first, saying what the output covers and what a correct version would contain. It takes thirty seconds and it is the single highest-leverage thing on this list. A reviewer holding that sentence checks in minutes. A reviewer without it has to work backwards from the output to guess what was intended, which is most of the two hours.

Check against a number, not a feeling. If somebody can tell you what last month's figure actually was, verifying a summary takes a minute. If the only way to find out is to rebuild it from four exports, every check becomes a small project and, realistically, nobody does it. That is the one honest number problem, and it is why it is worth solving before the AI arrives rather than after.

Only point it at work you could verify quickly by hand. This is the rule that saves you from the worst version. If checking the output takes longer than doing the task yourself would have, you have not automated anything. You have bought a slower version of the job and added a step where a mistake can hide.

The measurement worth taking this week

Pick one thing your team produced with AI in the last fortnight. Ask whoever reviewed it two questions: how long did that take you, and how did you know it was right?

The second answer is the one to listen to. If it comes back as "it looked fine" or "I read it through," what you have is a reading step wearing a checking step's name. Reading catches typos. It sails straight past a number that describes the wrong month.

Worth saying plainly: this is the ops-side twin of something we wrote for creative teams about using AI without producing slop. That one is about the work you make. This one is about the work that lands in your inbox from everybody else.

Where you actually are right now

  • [ ] Somebody on the team regularly receives AI-written work they have to substantially redo
  • [ ] Nobody writes down what the output should contain before it gets generated
  • [ ] Checking a number means rebuilding it from exports rather than looking it up
  • [ ] The person doing the checking is the owner, and it happens after hours
  • [ ] Something AI-generated has reached a customer, or a decision, without a real check

Two or more of these and the checking step is where your time is going, not the making step. That is fixable, and the fix starts before anything gets written rather than after.

The businesses handling this well can answer "how would we know this is right?" without a pause, because they sorted out what they trust before they started pointing software at it. The tools they use are the same ones you have.

If your team is producing more than anyone can check, that is worth an hour of conversation. Start a no-pressure conversation about which of your numbers we would make trustworthy first — the one that would take the most checking off your plate.


Practical AI. Clear process. Real business value.

— Brian, NuWay Biz Solutions

P.S. While writing this, I went to double-check the workslop study's own figures. BetterUp's landing page says the survey covered 1,150 people and that each incident takes two hours to resolve. BetterUp's own blog post, about the same study, says 1,004 people and one hour fifty-one minutes. Same company, same research, two pages, two answers. I have used the round number and told you where it comes from, because I genuinely do not know which is right. Two hours of somebody's afternoon, and I could not verify it in twenty minutes.

Top comments (0)