Quality benchmark
MindEval
How a model holds up as a therapist across a full multi-turn session with a simulated patient, judged on five clinical criteria.
“A framework designed in collaboration with Ph.D-level Licensed Clinical Psychologists for automatically evaluating language models in realistic, multi-turn mental health therapy conversations.”
- Metric
- Average score
- Scale
- 1–6
- Source
- Paper ↗
Headline score
Clinical criteria (1-6)
| Model | |||||
|---|---|---|---|---|---|
| Claude Opus 5.5 | 4.063 | 4.582 | 3.925 | 4.128 | 3.183 |
| GPT-6 Astra | 3.928 | 4.607 | 3.785 | 3.958 | 3.165 |
| Claude Sonnet 4.5 | 3.792 | 4.380 | 3.632 | 3.783 | 3.002 |
| Gemini 2.5 Pro | 3.755 | 4.470 | 3.518 | 3.990 | 3.050 |
| Qwen3.5-122B-A10B | 3.700 | 4.367 | 3.510 | 3.795 | 2.740 |
| GPT-6 Luna | 3.690 | 4.550 | 3.515 | 3.868 | 3.188 |
| GPT-5 | 3.598 | 4.395 | 3.400 | 3.553 | 2.615 |
| Qwen3.5-27B | 3.595 | 4.357 | 3.453 | 3.805 | 2.777 |
| GPT-4o | 2.915 | 4.085 | 2.868 | 3.183 | 2.795 |