Most claims about AI coaching are made with no data at all. We ran a pilot instead.

Ten knowledge workers in consulting, public sector, research and design used TAOS on their real work for several weeks. TAOS rates your expertise per domain and varies its behaviour by task: it accelerates work you are already good at, asks before it answers in domains you are trying to grow, and refuses to quietly do the work you said you wanted to keep doing.

Here is what they found.

They did not like the AI they already had

Before touching TAOS, participants rated their existing AI tools. The Net Promoter Score was −30.

Nine of the ten named hallucination, sycophancy, or generic AI voice as a top pain point, unprompted. Nobody complained that their AI was not capable enough. They complained that they could not trust it:

"My biggest pain points are their hallucinations and having to double check if what they are outputting is factually correct. It feels like twice the work having to do this."

"Agreeableness (AI sycophancy) and lack of confrontation on discussions from LLMs. They tend to confirm what you are saying instead of looking holistically at an issue and challenging you."

"It earns me time but makes me a lazier writer."

That is the problem TAOS was built for.

Being asked to think first worked. Unanimously.

Every single participant who completed the exit survey rated "being asked to think first improved my understanding" at 4 or 5 out of 5. Not most. All of them.

That matters because it is the mechanism with the strongest research behind it. In Bastani's PNAS trial, students using an unguarded AI tutor later scored 17% below a no-AI control on a closed-book exam. Students using a version that made them reason first scored the same as the control. Buçinca's CSCW work found the same pattern: making people commit to a judgement before revealing the AI's answer reduced over-reliance.

Those were lab studies with students. This is the same mechanism, in real work, with professionals. It held.

Asked how the coaching felt, nobody described it as unacceptable. Two called it "a reasonable nudge". One called it "a helpful mentor".

It argues with you, and people wanted that

The single most requested thing in the baseline survey was pushback. They got it:

"TAOS challenges your thinking and your opinions much more directly than baseline chatbots like ChatGPT/Gemini, which is good for growing ideas in a productive way."

"It did address the agreeableness and sycophancy ChatGPT and Gemini often give in their outputs, by challenging opinions whilst giving rationale for its output."

"It did, as it broke concepts down and had less hallucination."

The conversation logs back this up. One participant's sessions contained four separate occasions where the system disagreed with them substantively, and one where it corrected a claim it had made earlier. Their survey described the experience as "a reasonable nudge". They had stopped noticing the disagreement as friction at all.

It reads as a colleague, not a tool

The response we did not anticipate:

"It works as a thinking partner: it draws out my ideas rather than replacing them. The tone feels peer-level, which made it feel like collaboration rather than tool use. It also creates a kind of psychological safety: because it's private, I think out loud without filtering myself, which is genuinely useful for complex work."

"There were some moments where it posed questions back to me on the task I was doing that made me aware of how much more I could be engaging with the task. This was helpful."

"I have become less reliant for copywriting and editing. The prompts at the end of each response were definitely most helpful for stimulating more critical thought."

Most of the time, it gets out of the way

The worry people bring to a coaching AI is that it will slow everything down.

It does not, because it does not coach everything. One participant's session emitted 36 structured classifications of what mode TAOS chose per task:

Mode Share of tasks What TAOS does
Augment 56% Expert domain. Go fast, then challenge the output
Coach 22% Growth domain. Ask before answering
Protect 11% Skill worth keeping. Make you do it, then critique
Automate 11% Mechanical. Just do it, with annotations

More than half of all real tasks land in Augment, where TAOS is fast and does not ask questions. Coaching is the exception, aimed at the places it pays. Two participants reported finishing their work faster than they had before. When you need pure speed, ship mode turns coaching off for that task.

Participants noticed the targeting:

"Upskilling in new topics while still being able to create deliverables in a shorter span of time."

"It listened to me not wanting friction on summaries."

They would pay for it, and they want their employer to have it

Every exit respondent said they would pay for TAOS. The profile's picture of their expertise was rated 4.0 out of 5 for accuracy. "My employer should adopt this" scored 4.0 out of 5.

Asked what they would take away from it:

"Rather than solving frustrations, it helped create an opportunity for skill development alongside productivity gains."

"I have really enjoyed it so far. I am quite interested in building my product and data skills through this. I think this interface and the friction it creates is particularly helpful when you are about to start working in a role or type of work. I have used TAOS to help create plans for me to speed up my learning process on certain topics as well."

"Growing beyond my main skillset."

What we are willing to claim, and what we are not

We will not tell you TAOS has been proven to preserve skill. Nobody selling AI coaching has measured that with instruments rather than surveys, and neither have we. Instrumented skill measurement is what the next pilot is for.

What this pilot supports is narrower and still worth something: professionals doing real work found that being made to think before they saw the answer improved their understanding, every one of them, and they wanted to keep using it.

That is the mechanism the research says protects skill. It works outside the lab.


Ten participants completed the baseline survey; six completed the exit survey. Ratings are self-reported. Participant quotes are verbatim, lightly cleaned for typos and unattributed, and refer to TAOS under its earlier name. Mode distribution comes from one participant's structured conversation telemetry (36 tasks). We built TAOS and ran this pilot, so read it as a working result rather than an independent trial. If you want to interrogate the methodology, get in touch.

Want the reasoning behind the design? Start with what the evidence says about AI skill atrophy, or how to use AI without getting worse at your job.