Nearly every AI adoption dashboard measures the same four things: seats provisioned, weekly active users, prompts per user, and self-reported time saved. All four go up when your workforce is quietly getting worse at its job.
That is not a rhetorical point. It is what the underlying studies predict.
Why usage metrics are blind to the thing you care about
In the Bastani PNAS trial, students using an unguarded AI tutor performed better than the control group while they had the tool. Engagement was high. Output quality was high. Every metric a dashboard could see was green.
Then the tool was withdrawn and they sat a closed-book exam. They scored 17% below students who had never used AI.
The degradation was invisible from the outside for the entire period it was accumulating. It only became visible under a condition the dashboard never creates: the tool going away.
The BCG trial found the complementary problem. Consultants using GPT-4 were roughly 25% faster on tasks inside the model's competence and 19 percentage points less accurate on tasks outside it. Throughput and accuracy moved in opposite directions depending on a property of the task that no usage metric records.
So: high adoption, high satisfaction, high throughput, and a workforce that is losing the ability to catch the model when it is wrong. Every leading indicator you have would say this is going well.
What you are actually trying to measure
Three distinct things get muddled under "AI impact".
Throughput. Work per unit time. Easy to measure, usually the only thing measured, and it moves early.
Quality after rework. Workday's November 2025 survey of around 3,200 workers found roughly 40% of AI time savings are lost to rework. If you count the hours saved and not the hours spent fixing, you are booking a gain that never reaches the P&L.
Capability. Whether a person could still do the task well without the tool, and whether they can tell when the tool is wrong. This is the one that determines your organisation's value in three years, and almost nobody measures it.
Capability is a level. Most instrumentation records a derivative with no level attached: a per-interaction impression that someone seems to be growing or slipping. A derivative with nothing to differentiate against is not a measurement. It is a vibe with a timestamp.
A measurement design that would actually work
Anchor to a rubric, not a Likert scale. "Rate your Python 1 to 5" and "the model thinks this person is a 3" are not the same quantity and cannot be subtracted from one another. Define what each level means in observable behaviour, and score self-report and observation against the same anchors. Otherwise the drift you compute is noise.
Observe the opening frame, not the AI-assisted output. This is the trap that ruins most designs, and it is subtle.
A system that adapts to a user's expertise will manufacture the evidence for its own beliefs. If it explains basics to people it thinks are novices, they will look like novices. If it challenges people it thinks are experts, they will produce expert-looking pushback. The measurement error correlates with the prior. The system confirms what it already believed, most confidently where it was most likely wrong.
The only clean signal is the user's own framing of a task, captured before the system responds. How they scope the problem, what constraints they name unprompted, which tradeoffs they surface. That happens before any adaptation can contaminate it.
Require a threshold before you claim drift. A single observation is noise. Something like five or more trusted observations, comparing a median against the stated level, with a full-point gap required before you report anything.
Watch the tail, not the average. Averages hide the people a tool is failing. A mean that says an intervention costs nothing is compatible with a minority of staff quietly abandoning it. Report distributions, and act on the tail.
Create the withdrawal condition. You cannot detect atrophy in a population that never works unaided. Periodic unaided work is not a productivity loss. It is the only instrument you have.
The surveillance problem, and why you do not need it
Everything above sounds like it requires watching individuals. It does not, and building it that way is a mistake on three counts.
It is legally expensive. Individual-level capability profiling of employees pulls an AI deployment toward the heavier end of the EU AI Act's obligations. Aggregate-only reporting, with k-anonymity floors, keeps it light.
It destroys the signal. People who know their expertise is being scored for their manager will perform expertise. You will measure impression management.
And it is not necessary. The decisions an organisation actually makes with this data are structural: which domains are at risk, whether junior staff are developing, whether a team has stopped catching errors. All of those answer at the level of a domain or a cohort. None of them require knowing that Sarah in analytics is slipping at SQL.
The right posture is that administrators set policy and see aggregates, and individuals see their own data. Policy without surveillance. That is how TAOS is built: org-level mode mix, coaching engagement, and atrophy risk by domain, with no individual drill-down and a minimum cohort size before any cell renders.
The uncomfortable part for the business case
There is a version of this pitch that says coaching pays for itself in productivity. Be careful with it.
The productivity argument for expertise is strong, and it now has Anthropic's own data behind it. Their June 2026 analysis of roughly 400,000 Claude Code sessions across about 235,000 users found that domain expertise, not coding background, predicts whether an agent session succeeds. Verified success more than doubled from novices (15%) to experts (33%), while the gap between software engineers and every other occupation was five points. Novices abandoned 19% of sessions; everyone else abandoned 5 to 7%.
Read that as an ROI statement. Expertise is the input that determines what you get back from an AI agent, and unmanaged AI use degrades expertise. The agent is only as good as the judgement pointed at it, and judgement is the thing being quietly spent.
There is a second argument, and it is the one that survives a CFO precisely because it is not dressed up as a productivity claim: you would like your senior people to still be senior in five years. Capability is an asset with a depreciation schedule, and right now nobody is booking the depreciation.
One more thing worth saying plainly. Skill preservation has not yet been demonstrated anywhere in this market with instruments rather than surveys, by us or by anyone else. What our pilot shows is that every participant found being made to think first improved their understanding, and that all of them wanted to keep using it. That is the mechanism the PNAS trial says protects skill, observed outside the lab. Treat any vendor claiming more than that, including a future version of us, as ahead of the evidence.
Where to start
Pick three to five domains that matter to your organisation's future and define behavioural anchors for each. Instrument the opening frame, not the output. Report by domain and cohort, never by person. Build in periods of unaided work. Track the abandonment tail alongside the average.
Then, and only then, look at what the numbers say about whether AI is making your people more valuable or less.
If you would like to see what that looks like running, talk to us, or read the pilot data first. It includes the findings that went against us.