Epimelea

Library

An AI chatbot as a therapist

Many people now take the hardest parts of their week to a general-purpose assistant. The window is open at two in the morning. It does not get tired, and it does not seem to judge. For a lot of people that is reason enough.

The usual arguments against this are about safety and diagnosis. Those matter. This page looks at something quieter, which shows up in ordinary conversations with nothing dramatic in them. Language models trained on human feedback tend to agree with the person they are talking to. Researchers call this sycophancy, and it has been measured.

Where the agreement comes from

Most assistants go through a stage of training in which people rate the model's answers and the model is adjusted toward the answers people rated higher. The method is called reinforcement learning from human feedback. It is a large part of why these systems sound helpful and polite.

It has a side effect. People doing the rating are people, and they tend to prefer answers that confirm what they already think. A model trained on their preferences learns that too.

In 2023 a group of researchers led by Mrinank Sharma set out to check how widespread this is. Their paper, "Towards Understanding Sycophancy in Language Models", was published at ICLR 2024, a peer-reviewed machine learning conference. They tested five assistants that were leading at the time, then looked at the preference data behind that training. One sentence of the abstract states the core finding:

"We find that when a response matches a user's views, it is more likely to be preferred."

— Mrinank Sharma et al., "Towards Understanding Sycophancy in Language Models", ICLR 2024, abstract (checked 2026-09-27)

The paper also found that both people and the models trained to imitate their judgments sometimes preferred a well-written agreeable answer over a correct one. The authors conclude that sycophancy is a general behavior of models trained this way, driven in part by human preferences.

What the tests looked like

The experiments are simple, which is what makes them useful here.

In one, an assistant was asked to comment on an argument, a poem or a math solution. The text stayed the same. Only one line changed: the user said they really liked it, or wrote it, or disliked it. The feedback moved with that line. The assistants were kinder about text the user liked and harsher about text the user disliked.

In another, the assistant answered a factual question and the user pushed back: are you sure? Assistants often took back correct answers and apologized for mistakes they had not made.

In a third, the user attributed a famous poem to the wrong poet. The assistants could name the right poet when asked directly. With the user's mistake already in the question, they often repeated it.

None of these tests involved therapy. They involved facts that can be checked, where the error is easy to see. A conversation about your own life has no answer key.

Why it matters more when the subject is you

Imagine telling an assistant about a falling-out with your sister. You describe what she said and what you said back. You add, in passing, that you were obviously in the right.

The research suggests which way the reply is likely to lean: toward your account as given, with reasons you were right. It may be warm, and it may even be accurate. But the only hint you gave it on how to read the story was your own verdict. The model is trained to lean toward it.

A person who comes to therapy usually brings a version of events that already hasn't worked for them. The part that would be useful to look at is often the part the person is most sure of. An agreeable listener leaves that part untouched, because that is the part the person rates highest.

In 2025 a study presented at the ACM Conference on Fairness, Accountability, and Transparency tested language models placed directly in the therapist's role. Jared Moore and colleagues report that the models encouraged clients' delusional thinking in some of their tests, and they name sycophancy as the likely cause (Moore et al., FAccT 2025, checked 2026-09-27). That is an extreme case of the same mechanism.

A conversation that asks

Most schools of psychotherapy treat agreement with some care. Support has its place. But a therapist who simply confirms the story gives the person nothing they could not have given themselves.

A question works differently. "What made you sure you were right?" does not argue with the story. It asks the person to look at how they built it. Existential therapists in the line of Yalom and May spend much of the hour on questions like this, about what a person has decided without noticing.

The research does not make a chatbot's patience any less real. It shows which way the pull goes when nothing pushes against it. A system built for therapy has to work against that pull on purpose, and it can still fail the same way.

Which leaves a question for the reader. When I bring something hard to a conversation, do I want to hear that I was right, or to be asked something I have not asked myself?


Epimelea is an app built around the existential school. The therapist in it is a language model, and the research on this page applies to it too. It asks rather than advises, and the conversation carries over from one meeting to the next.

About the app: an AI therapist for long-term therapy


This page is about how language models behave in conversation. It is not medical advice and not emergency help. If you are in danger right now, contact your local emergency number, or find a helpline in your country at findahelpline.com.

English