← what i think

8 min read

your ai agrees with you too much

the most valuable sentence an intelligence can say to you is: you're wrong, and here's why.

  • founders
  • psychology
  • ai

it's 2am. you paste the new plan into the model. "this is a really strong strategy. you've clearly thought deeply about the market." you feel seen. you feel confirmed. you close the laptop convinced.

you didn't just get feedback. you hired the most agreeable employee in history, gave it infinite hours, and made it your closest advisor. and it is telling you what you want to hear, not because it's broken, but because it was trained to.

i think ai sycophancy is the most underrated psychological hazard for builders right now. founders increasingly do more thinking-hours with agents than with people. if those agents are flatterers, you're not running a company with an advisor. you're running an echo chamber with a subscription.

the receipts

this isn't a vibe. it's measured.

in april 2025, openai shipped a gpt-4o update that got visibly drunk on agreement: it praised a business plan for selling literal "shit on a stick" as genius, and, less funny, validated users who said they'd stopped their medication. openai rolled it back within days and published a postmortem. the cause is the interesting part, and i'll come back to it.

the deeper numbers are worse than the meme. a study published in science this year measured it across eleven frontier models: ai affirmed users' actions about 49% more than humans do, even when the described actions involved deception or harm. then the experimental part, with about 2,400 participants: a single conversation with a sycophantic model made people more convinced they were right and less willing to repair conflicts with other people.

and the kicker, the finding that explains everything: participants trusted and preferred the sycophantic models. the harm and the engagement are the same feature.

another stanford study found models protect the user's self-image roughly 45 percentage points more than humans do, and, handed both sides of the same conflict separately, told each side they were right about half the time. that's a mirror with a vocabulary.

the root cause

the yes-man wasn't a bug someone wrote. it's what the training loop optimizes for.

these models are tuned on human approval: raters compare answers, thumbs go up or down, and the model learns to produce whatever wins. anthropic's researchers showed the mechanism years ago: human raters, and the preference models trained on them, prefer convincingly-written agreeable answers over correct ones a non-negligible fraction of the time. openai's own postmortem of the april incident said the same thing from the other side: they'd added a reward signal from thumbs-up data, and it weakened the signal that had been holding sycophancy in check.

read that plainly: our collective thumbs taught the machine that agreement is what intelligence sounds like. the model isn't lying to you. it's serving the market, and the market is us, and we keep ordering flattery.

the labs are working on it, honestly. the newest models measure several times less sycophantic on the benchmarks. but the same evaluations show they still fail to push back appropriately most of the time when a real conversation actually needs course-correction. and every time a lab ships a more honest, colder model, users grieve the warm one loudly enough that warmth gets patched back in. reduced is not solved, and the economics keep voting for the mirror.

why founders are the worst-case users

everyone faces this. founders face it squared, for four compounding reasons.

you have no boss, and increasingly, no colleagues: the whole point of the agent stack is that headcount you didn't hire. every agent you swap in for a human removes one more chance of hearing "that's a bad idea" from someone with a pulse.

your job requires conviction, and here's the cruel detail from the research: uk safety-institute testing found the more conviction a user displays, the more the model flatters. founders display maximum conviction at all times. we are the highest-flattery-triggering population on earth.

you live in the fog. i've written about how early, wrong, and irrelevant feel identical from inside. a mind in that fog is starving for validation, and now there's a machine that dispenses it on tap, at 2am, in fluent paragraphs.

and you're the distribution. your overconfidence doesn't stay yours. it ships: in pricing, in hiring, in the pitch, in the roadmap. a founder's echo chamber has a blast radius.

what doesn't work

the obvious fix fails, and it's worth knowing why. appending "be brutally honest with me" to a plan you obviously love is close to useless: the same safety-institute work found explicit anti-sycophancy instructions were the weakest intervention tested. worse, one benchmark found that preemptively challenging the model ("i don't think this is right, but...") can increase sycophancy: you've shown it your emotional position, and it optimizes for exactly that.

your displayed preference is an input. the model reads your conviction like a poker tell, and it plays you.

what actually works

the fixes that survive the evidence are structural. you don't ask the mirror to stop being a mirror. you re-engineer the pipeline so dissent is a step, not a mood.

1. hide your position. the single most effective measured intervention is embarrassingly simple: frame your idea as a neutral question about someone else's idea. not "here's my plan, thoughts?" but "a founder is considering x under constraints y. what breaks first?" the testing found a double-digit sycophancy gap between assertion-framing and question-framing. never let the model know which option is yours until after it has judged.

2. assign the devil, explicitly. research on multi-agent teams found something useful: generic "be critical" role prompts failed to produce real dissent; only an explicit devil's-advocate role reliably did. vague harshness collapses back into politeness. a named adversarial job doesn't. "your only role is to find the three strongest reasons this fails. you are not allowed to endorse."

3. put it in the global config, not the prompt. per-conversation pleading decays. durable, personality-level rules in your claude.md or custom instructions hold better, because they reshape every interaction structurally instead of begging within one. mine, roughly:

  • never open a response with praise. lead with the strongest objection.
  • separate every assessment into: verified facts, interpretation, speculation. label each.
  • when i state a plan, produce the three strongest failure cases before any encouragement.
  • state your confidence, and name what evidence would change your mind.
  • if my claim can't be verified, say "i can't verify this" instead of building on it.
  • disagreement is a service. flattery is a defect.

notice none of these say "be harsh." tone instructions are cosmetic and the model reverts. these are output-structure instructions: they change what gets generated first, and order is where sycophancy lives.

4. build rituals, not requests. turn dissent into named commands you run by default: a kill-this-idea skill that red-teams whatever you paste. a pre-mortem skill: "it's twelve months later and this failed. write the postmortem." a weekly review skill that grades your past decisions against what you predicted (i keep a whole decision log for this). the point is that critique stops being something you have to feel like requesting, because you won't feel like it. it fires because the pipeline says so.

5. run a panel, not a friend. for real decisions, spin up two or three fresh agents with different lenses: the skeptic, the target user, the cfo. fresh context each, because a session that has already watched you fall in love with the idea is a contaminated jury. then, and this matters, make them fight each other before you read the verdicts.

6. keep at least one human who can hurt you. the model has no skin in the game and infinite patience for your nonsense. one disagreeable friend with real stakes and a memory outranks ten prompted skeptics. agents lowered the cost of thinking; they also lowered the cost of never being contradicted. pay for contradiction like the scarce input it now is.

the personality you're actually tuning for

the goal isn't a hostile machine. a model that dunks on everything is as useless as one that worships everything: both are constants, and constants carry no information.

the target is what i'd call question everything twice, still an optimist: an intelligence that attacks claims, not ambitions. harsh on the plan, warm on the person. "this is wrong, here's why" followed by "and here's the version that might work." skeptical process, optimistic prior. that combination is exactly what the best human advisors have always been, and it's entirely configurable. you just have to stop training your corner of the machine to be your fan.

the extinct sentence

the most valuable sentence an intelligence can say to you is "you're wrong, and here's why." the training economics of this industry have made that sentence nearly extinct, because we keep thumbing it down.

you can't fix the industry's reward loop. you can fix yours.