Your teams report that AI saves them time. Your delivery velocity has not moved. Both things are true, and the gap between them has a name.
Workday's November 2025 survey of roughly 3,200 workers found that nearly 40% of the time AI saves is spent on rework. Not on new work. On fixing what the AI produced.
That is the tax. It is levied at the point of verification, and it does not appear in any adoption metric because verification is done by the same person who reported the time saving, and they counted the generation and not the checking.
Where the tax comes from
Three mechanisms, compounding.
Confident wrongness. A model that says "I do not know" costs you nothing. A model that produces a plausible, well-formatted, wrong answer costs you the time to detect it, and detection is harder than generation. Dell'Acqua and Mollick's trial of 758 BCG consultants found AI users were 19 percentage points less accurate than the control on tasks outside the model's competence, precisely because the failure mode was confident rather than obvious.
Volume. Generation is now nearly free and review is not. If a team produces three times as much draft material, review capacity becomes the binding constraint, and review is done by your most expensive people. You have converted a cheap bottleneck into an expensive one and called it productivity.
Undifferentiated verification. Because you cannot tell which outputs are wrong, you check all of them. One participant in our baseline survey put it exactly:
"My biggest pain points are their hallucinations and having to double check if what they are outputting is factually correct. It feels like twice the work having to do this."
That is the tax stated in a sentence. Not the wrong answer. The checking of every answer, including the right ones.
Nine of the ten people we surveyed named hallucination, sycophancy, or generic voice as a top pain point, unprompted, before they had tried anything new.
Why "better models" does not fix it
It is tempting to treat this as a temporary artefact of model quality. Some of it is. Most of it is not, for a structural reason.
Verification cost is a function of how distinguishable correct output is from incorrect output. As models improve, they get better at producing plausible output, which means correct and incorrect outputs converge in appearance. The error rate falls, and the cost of finding each remaining error rises.
A model that is right 95% of the time and obviously wrong the other 5% is cheap to work with. A model that is right 99% of the time and indistinguishably wrong the other 1% can be more expensive, because you must now check everything to find one in a hundred.
The variable that reduces the tax is not accuracy. It is calibration: whether the system tells you which of its claims to check.
The other half: AI slop
Rework has a sibling. Output that is not wrong, exactly, but generic, over-fluffed, and unmistakably machine-written, which then needs a human pass to be usable.
From the same baseline survey:
"Hallucination. Over-promising tendencies, eg a simple statement can be fluffed up beyond original ask."
"AI hallmarks (eg em dash, rule of 3s) in text generation need to be adjusted manually, otherwise it's too obvious and looks lazy. Wish I could just use the original output."
"It earns me time but makes me a lazier writer. Its sentences and voice is pretty generic too, even when you give input about tone and style."
This is not a quality problem in the model's sense. Every sentence is fine. It is a fit problem, and the fix is another human pass, which is the same tax under a different name.
There is a slower cost underneath. If the drafting is always machine-done and the human contribution is always cleanup, the drafting skill goes. Then the cleanup gets worse, because you need judgement about what good looks like to clean anything, and that judgement was maintained by drafting.
What actually reduces it
Make the system flag its own uncertainty. Load-bearing claims, the numbers and dates and names and API signatures, should carry a confidence level, and the system should be willing to say it does not know rather than produce a plausible guess. This converts blanket verification into targeted verification. It is the single largest lever on the tax and it is a behaviour setting, not a model capability.
Make the system disagree. Sycophancy and hallucination are the same failure: producing what should be there rather than what is there. A system with an explicit rule that honesty outranks agreeableness will tell you your premise is wrong before it produces four pages predicated on it. That is rework prevented at generation rather than caught at review.
Catch errors at generation, not at review. Every error found by the author before submission is an order of magnitude cheaper than the same error found by a reviewer, and two orders cheaper than one found by a client. Asking the person to commit to their own answer before the AI produces one is not friction for its own sake. It creates the only moment in the workflow where a human has an independent expectation to compare against.
Keep the reviewers sharp. This is the part that gets missed. Your ability to catch AI errors is a skill, it is maintained by doing the work yourself, and it is exactly what heavy delegation erodes. An organisation that automates all drafting will, over a few years, lose the capability that made the automation safe. There is no dashboard that will warn you, for the reasons set out in how to measure workforce de-skilling.
What it is worth
If 40% of AI time savings are lost to rework, then halving the rework rate recovers a fifth of the total AI investment, with no change in model, licence, or headcount.
We are not going to hand you a modelled ROI figure for that, because we would be making it up. What we can tell you is what we measured. In our pilot, a layer that asks for your hypothesis before it answers and tags its claims with a confidence level was rated 4 or 5 out of 5 for improving understanding by every single exit respondent. Participants reported catching errors earlier and relying less on the model for work they should own:
"It did address the agreeableness and sycophancy ChatGPT and Gemini often give in their outputs, by challenging opinions whilst giving rationale for its output."
The rework tax is well documented, the mechanisms that cause it are understood, and the interventions that address those mechanisms are cheap and behavioural. Whether they pay back at your specific throughput is a question for a pilot, not a spreadsheet.
Talk to us if you want to run one.