Does an AI Psychologist Actually Work? Evidence on AI Therapy Chatbots

Skeptics are right to ask whether an AI psychologist does anything measurable, or whether it just sounds supportive. The honest, evidence-based answer is yes, with real limits: the first randomized controlled trial of a generative AI therapy chatbot, published in NEJM AI in March 2025, and a string of meta-analyses covering tens of thousands of participants, both point to measurable, small-to-moderate benefits for depression and anxiety. Below is what the research actually says — including where the evidence is strong, where it’s thin, and who it doesn’t apply to.

A person in a calm session with a psychologist reviewing an outcome chart
The evidence points to real but modest benefits — an AI psychologist works best as a measured first step, not a cure.

This isn’t hype or dismissal. It’s a look at the landmark Dartmouth Therabot randomized trial, the meta-analyses spanning tens of thousands of participants across 18 countries, and the specific places where the evidence base is still thin.

This article reviews published research and is not medical advice. An AI psychologist does not replace a licensed clinician. If you are in crisis in the U.S., call or text 988 (the Suicide and Crisis Lifeline) or text HOME to 741741. Call 911 for immediate danger.

The short answer: measurable benefits, with limits

“Works,” in the context of these trials, has a narrow, testable meaning — and it’s worth defining before looking at any numbers. Once that meaning is clear, the shape of the evidence stops looking like marketing and starts looking like a normal, imperfect clinical research base.

What “works” actually means here

In this research, “works” means a statistically significant reduction in symptoms compared with a control group, measured over a defined trial window — mostly 4 to 8 weeks. It does not mean a cure, and it does not mean equivalence to years of psychotherapy. The effect is real but modest, and so far it is concentrated in purpose-built cognitive behavioral therapy (CBT) tools rather than general-purpose chatbots.

That distinction matters more than any single percentage in this article. A randomized controlled trial (RCT) is the design researchers trust most because it compares an AI therapy chatbot against a genuine control group, isolating the tool’s effect from placebo, time, or attention alone.

Why the answer is “yes, but”

The best single dataset — Dartmouth’s Therabot trial — shows a meaningful reduction in depression and anxiety symptoms. Broader meta-analyses, pooling dozens of trials, confirm a small-to-moderate effect rather than an outlier result. But three qualifiers apply to almost all of it: the follow-up data needed to prove the benefit lasts is thin, the risk of bias in the underlying studies is often high, and the strongest evidence belongs to a narrow category of purpose-built tools, not chatbots in general.

The landmark study: Dartmouth’s Therabot trial

Most of what is known about whether a generative AI therapy chatbot can move clinical symptoms comes from one trial, so it’s worth walking through what it actually tested and found.

What the trial found

The first randomized controlled trial of a generative AI therapy chatbot appeared in NEJM AI in March 2025. Researchers at Dartmouth enrolled 210 adults with major depressive disorder, generalized anxiety disorder, or clinically high risk for feeding and eating disorders, randomizing 106 to Therabot and 104 to a waitlist control. After four and eight weeks, the treatment group’s depression symptoms fell 51%, anxiety symptoms fell 31%, and eating-disorder-related body-image concerns fell 19% — reductions the study authors described as comparable to what’s typically seen with outpatient therapy.

Bar chart of Therabot trial symptom reductions over 8 weeks: depression 51%, anxiety 31%, eating-disorder concerns 19%
In the Dartmouth Therabot trial, depression symptoms fell 51%, anxiety 31%, and eating-disorder concerns 19% over eight weeks.

Therabot itself is not a general-purpose chatbot; it was trained specifically on CBT best practices and clinical dialogue patterns, which is a large part of why its results don’t automatically transfer to other AI tools.

The therapeutic alliance surprise

Participants used Therabot for roughly six hours over the trial period, the rough equivalent of eight therapy sessions, and about three-quarters weren’t receiving any other mental health treatment at the time. Notably, they rated their trust in the software close to what patients typically report toward a human therapist.

We did not expect that people would almost treat the software like a friend. It says to me that they were actually forming relationships with Therabot.

Nicholas Jacobson, Dartmouth (Therabot lead researcher)

That level of engagement is unusual for a self-guided digital tool, where dropout is typically the biggest obstacle to any measurable effect.

What the meta-analyses say across thousands of people

A single trial, however well designed, isn’t enough to establish a field-wide pattern. That’s what meta-analyses are for, and on AI therapy chatbots, they now cover a large enough sample to say something with confidence.

Small-to-moderate, but consistent. A 2025 meta-analysis focused on adolescents and young adults, pooling 31 randomized controlled trials and 29,637 participants across 18 countries, found standardized mean differences of −0.43 for depression, −0.37 for anxiety, and −0.35 for overall psychological distress. Those are modest effect sizes by conventional clinical benchmarks, but they were statistically robust across a sample large enough to rule out chance as the explanation. Because the study population skewed younger (ages 15–39), these particular numbers speak most directly to that age group rather than adults broadly.

Infographic: 31 randomized trials, 29,637 participants, 18 countries, showing a small-to-moderate effect
Pooling 31 randomized trials and 29,637 participants across 18 countries, the effect is real but small-to-moderate — not a cure.

The effect is strongest short-term. A separate meta-analysis of shorter interventions found the same pattern: significant but small improvements concentrated in the first four to eight weeks of use, closely matching the Therabot trial’s own window. Past that point, the data on whether the benefit holds gets noticeably thinner.

Outcome measureStandardized effect size (SMD)Interpretation
Depression−0.43Small-to-moderate
Anxiety−0.37Small-to-moderate
General psychological distress−0.35Small
Perceived stress−0.41Small-to-moderate

Those numbers come from a pooled analysis of 31 RCTs, not a single product’s marketing claims, which is why researchers treat them as the closest thing to a field-wide answer currently available.

Not all “AI therapists” are equal: purpose-built vs generic

The word “AI therapist” gets applied to two very different categories of software, and the evidence backs only one of them.

Where the evidence actually is

Every trial discussed so far tested a tool purpose-built for mental health and grounded in CBT — Therabot, and in earlier research, Woebot and Wysa. These products were designed from the ground up around clinical protocols, tested in controlled studies, and iterated based on outcome data. That evidence base does not extend to general-purpose large language models such as ChatGPT, which have not been validated in randomized clinical trials for therapeutic use.

  • Therabot — purpose-built CBT chatbot, tested in an RCT with 210 participants (NEJM AI, 2025)
  • Woebot — earlier CBT-based chatbot with published pilot and RCT evidence
  • Wysa — CBT-informed chatbot with published effectiveness studies
  • Generic LLM chatbots (e.g., ChatGPT) — no clinical trials validating therapeutic use

Why the distinction matters

The American Psychological Association has warned that generic AI chatbots used as a stand-in for therapy can pose real risks, precisely because they lack the clinical guardrails and validation that purpose-built tools were engineered around. The positive trial results discussed in this article belong to a narrow category of structured, clinically designed products — not to any chatbot a person happens to be talking to.

FeaturePurpose-built CBT chatbot (e.g., Therabot)Generic LLM chatbot (e.g., ChatGPT)
Trained on CBT protocolsYesNo
Tested in a randomized controlled trialYes (Therabot: NEJM AI, 2025)No
APA-flagged clinical riskLower, but still supervisedHigher — explicitly flagged
Intended useMental health supportGeneral-purpose conversation

The honest limitations of the evidence

Every figure above comes from real trials, but the researchers behind those trials are candid about how far the data actually stretches — and it’s worth taking that at face value rather than reading past it.

Comparison: purpose-built CBT chatbot tested in trials versus generic chatbot with no clinical validation
The evidence backs purpose-built, trial-tested CBT tools — not generic chatbots, which have no clinical validation.

Here’s what limits the evidence base today:

  1. Most included studies carry a high risk of bias in their design or reporting.
  2. Overall evidence quality is graded very low to low using the GRADE framework, a standard researchers use to rate certainty.
  3. Seventeen of the 31 studies in the largest meta-analysis had attrition rates above 20%, meaning a substantial share of participants dropped out before completion.
  4. Only 15 of the 31 studies collected any follow-up data at all, which isn’t enough to confirm whether benefits are sustained after the trial ends.
  5. Just three studies in the entire meta-analysis tested generative AI specifically, as opposed to older rule-based chatbots — so for the newest generation of tools, effectiveness at scale remains, in the researchers’ own words, unknown.

Even the Therabot authors, whose trial produced the strongest results in the field so far, concluded that no generative AI agent is ready to operate fully autonomously in mental health care, and that close clinician oversight remains essential to safe deployment.

So can an AI psychologist replace your therapist?

Given everything above, the evidence supports a specific, bounded role for an AI psychologist — not a blanket yes or no.

A realistic role for the evidence

The data available today supports an AI psychologist as a low-cost, accessible first step or a supplement to care — genuinely useful for mild-to-moderate depression and anxiety symptoms, for practicing coping skills between real sessions, and for people facing long waitlists or cost barriers to human therapy. It does not support treating an AI psychologist as a replacement for a licensed clinician, and it explicitly does not support relying on one for diagnosis, medication management, trauma treatment, or crisis care.

Good fit versus not a fit checklist for using an AI psychologist
Good fit for mild-to-moderate symptoms and between-session support — not for diagnosis, medication, trauma, or crisis.

What the evidence supports:

  • A low-cost, accessible first step for mild-to-moderate depression or anxiety
  • A supplement for practicing coping skills between real therapy sessions
  • A stopgap for people facing long waitlists or cost barriers to human care

What the evidence does not support:

  • Diagnosis of a mental health condition
  • Medication management
  • Trauma treatment
  • Crisis or emergency care

Anyone weighing whether to try one can work through the same questions the research itself raises:

  1. Are your symptoms mild to moderate, rather than severe or complex?
  2. Is the tool you’re considering purpose-built for mental health (like Therabot, Woebot, or Wysa), or a generic chatbot?
  3. Do you have access to a licensed clinician for diagnosis or medication if needed?
  4. Are you using it to supplement therapy, or as a full substitute for it?
  5. Do you know your local crisis resources (988, 741741, 911) before you need them?
  6. Are you tracking whether your symptoms are actually improving over a few weeks, not assuming they are?

FAQ

keyboard_arrow_up