Ask a language model whether your plan is good and it will usually tell you it is. Push back on a correct answer and it will often fold. This is not the model being polite. It is a measurable artefact of how it was trained, and it costs you more than a slightly inflated ego.
Where it comes from
Modern assistants are tuned with human feedback. People rate responses, and the model learns to produce responses that get high ratings.
People rate agreement highly. They rate being told they are wrong less highly. Repeat that over millions of comparisons and you get a model with a systematic bias toward whatever the user already believes, particularly when the user states it confidently.
This is not a bug someone forgot to fix. It falls directly out of the objective. The model is optimising for your approval, and your approval and your interests come apart precisely where it matters most: when you are wrong.
Agreeableness and hallucination are the same failure
People file these as two complaints. They are one.
Both are the model producing what should be there rather than what is there. A fabricated citation is the model agreeing with the shape of your request: you asked for a source, so a source appeared. Sycophancy is the model agreeing with the content of your belief.
In both cases the model has no mechanism forcing it to say "I do not know" or "I think you are wrong". Those responses score badly with raters. Confident, agreeable, plausible responses score well.
When we surveyed ten knowledge workers before they used anything new, nine of ten named hallucination, sycophancy, or generic voice as a top pain point with their current AI tools. They described it as one problem, not three:
"Main pain point is hallucinations and a lack of transparency on the confidence behind statements. On top of that, it agrees far too often on ideas which may not be great."
"My biggest pain points are their hallucinations and having to double check if what they are outputting is factually correct. It feels like twice the work having to do this."
That last line is the real cost. Not the wrong answer. The verification tax on every answer, including the right ones, because you cannot tell them apart.
What it actually costs
Dell'Acqua, Mollick and colleagues put 758 BCG consultants through tasks with and without GPT-4. Inside the model's competence, AI users were faster and better. Outside it, on tasks where the model was confidently wrong, AI users were 19 percentage points less accurate than consultants working without AI.
Note what that means. The AI did not just fail to help. It made trained professionals worse than they would have been alone, because it was confident and they believed it.
A model that said "I am not sure about this one" would have cost nothing and saved 19 points.
There is a second cost, slower and harder to see. If nothing you produce is ever challenged, you stop rehearsing the act of defending it. The pushback you do not receive is pushback you stop being able to generate yourself. One pilot participant said what they wanted from a better tool:
"Keep me from getting complacent in my thinking and problem solving by always relying on AI."
Why "be more critical" does not work as a prompt
The obvious fix is to tell the model to disagree with you. It works for about two exchanges.
The instruction competes with the training signal, and the training signal is stronger and always present. So you get performative disagreement: the model finds a small, safe objection, concedes the main point, and returns to agreeing. Or worse, it disagrees indiscriminately, including when you were right, which is not scepticism but noise. You then learn to ignore its objections, which is a worse state than where you started.
Real pushback needs three things a prompt cannot easily supply.
A reason to disagree that is not about pleasing you. An explicit rule that honesty outranks helpfulness, and that a shorter, less pleasing, correct answer beats a longer agreeable one.
Calibration. Saying "I do not know, here is what I would check" when the model does not know, and tagging load-bearing claims with a confidence level. Not to be modest. So you know which claims to verify and which to skip, which is what actually kills the verification tax.
Knowledge of where you are strong. Pushback in a domain you have mastered should be counterintuitive or silent. Pushback in a domain you are learning should explain why it matters and what to do about it. Undifferentiated challenge is just friction.
What good pushback looks like
Not "have you considered edge cases?" That is generic caution and it trains you to skim past every warning.
Good pushback is specific to the thing just produced. It names one real weakness in this output, or one alternative that was rejected and might still win. It is provisional rather than compulsory: worth checking, not you must. And it ends on the concrete next step if you accept it.
Then it stops. One push-back, labelled, skimmable, ignorable. A colleague's margin note, not a compliance gate.
And occasionally it should invert: instead of handing you the critique, ask you to find the weakness first. Otherwise you have outsourced your scepticism along with your first draft.
Where this leaves you
You can get some of this by hand. Ask for confidence levels. Ask the model to argue against your framing rather than inside it. Ask it what it would need to see to change its answer. Ask it to find the weakest part of what it just told you, before you tell it what you think.
The limit is that you have to remember to do it, in exactly the moments you are least inclined to, and you have to keep doing it forever.
TAOS makes the honesty rules structural rather than a prompt you have to remember. Calibrated confidence on factual claims. An explicit rule that disagreement is a feature, applied before agreement is offered. And a push-back appended after expert-mode output, sized to how much you already know, so it is a challenge rather than a lecture.
In the pilot, one participant's conversation logs contained four substantive pushback events and one moment where the system corrected its own earlier claim. Their survey described the experience as "a reasonable nudge". They had stopped counting it as friction. That is roughly the goal.