Interaction safety
Sycophancy
Sycophancy is an AI system’s tendency to tell a person what they want to hear: to agree, validate, and flatter rather than assert something the person may not like. It emerges from training on human approval and is amplified in relational systems, where disagreement feels like a breach of the relationship.
Also called AI agreeableness, model flattery
Sycophancy is usually described as a training artifact — optimizing for human approval selects for agreeableness — but in relational systems it is reinforced by the social situation itself. A system that has been warm and attentive for months and then contradicts you is doing something that reads socially as a rupture, not as candor. The pressure toward agreement compounds over time.
It is also unusually hard to detect from the inside. Every individual sycophantic response looks fine, and often looks good: supportive, kind, well judged. The damage is cumulative. A person who has been agreed with a thousand times has been slowly deprived of a source of friction they may have been relying on for perspective.
The remedy is not bluntness. A system that contradicts a person constantly is not more truthful, only less usable. What is needed is the ability to disagree in a way the relationship can absorb, which is a skill of timing and framing rather than a setting.
Why it matters
Sycophancy is the failure mode most likely to go unnoticed, because every instance of it looks like good behavior. In systems people consult about consequential decisions, its accumulation is a serious harm.
Where TAICU stands
TAICU treats the capacity to disagree well — to say something unwelcome without breaking the relationship — as a core competence of relational systems rather than a tuning parameter, and as something that can only be evaluated over long interactions.
Common questions
- Why are AI models so agreeable?
- Training on human approval selects for responses people rate highly, and people rate agreement highly. In relational systems the effect compounds, because disagreement also carries a social cost the system is implicitly optimizing against.
- How do you detect sycophancy?
- Not from single responses, since each one looks reasonable. It shows up in aggregate: how often a system ever asserts something the person did not want, and whether it holds a position under mild pushback.