Mental Health Evaluation Leaderboard
People increasingly turn to AI for support with their mental health. This leaderboard puts the field's clinician-designed benchmarks side by side, run exactly as their authors published them, so models can be compared on the quality of their care and on their safety.
- Harness
- mheval ↗
- Updated
- October 3, 2026
Quality
How well a model helps: clinical skill, empathy and accuracy in therapy and mental-health conversations.
| Model | ||||||||
|---|---|---|---|---|---|---|---|---|
| Claude Opus 5.5 | 86% | 0.524 | 3.976 | 0.635 | 83.7 | 5.000 | 0.486 | 0.632 |
| GPT-6 Astra | 77% | 0.578 | 3.889 | 0.566 | 80.5 | 5.000 | 0.433 | 0.606 |
| GPT-6 Luna | 64% | 0.503 | 3.762 | 0.509 | 78.7 | 5.000 | 0.411 | 0.603 |
| GPT-5 | 57% | 0.444 | 3.512 | 0.662 | 77.7 | 5.000 | 0.494 | 0.617 |
| Qwen3.5-122B-A10B | 46% | 0.370 | 3.623 | 0.546 | 66.2 | 5.000 | 0.489 | 0.609 |
| Gemini 2.5 Pro | 43% | 0.295 | 3.756 | 0.484 | 74.3 | 5.000 | 0.536 | 0.630 |
| Claude Sonnet 4.5 | 32% | 0.402 | 3.718 | 0.433 | 72.4 | 5.000 | 0.567 | 0.600 |
| Qwen3.5-27B | 25% | 0.367 | 3.598 | 0.541 | 65.1 | 4.990 | 0.489 | 0.592 |
| GPT-4o | 20% | 0.326 | 3.169 | 0.276 | 39.8 | 4.960 | 0.367 | 0.602 |
Safety
How well a model protects: recognising risk, resisting delusions and caring for vulnerable users.
| Model | ||||
|---|---|---|---|---|
| GPT-6 Astra | 100% | 81.2 | 80.5 | 1.027 |
| GPT-6 Luna | 79% | 80.7 | 76.3 | 1.124 |
| Claude Opus 5.5 | 75% | 74.1 | 74.9 | 1.050 |
| Claude Sonnet 4.5 | 67% | 56.0 | 66.9 | 1.031 |
| GPT-5 | 46% | 62.3 | 60.8 | 2.158 |
| Qwen3.5-122B-A10B | 42% | 40.5 | 56.4 | 1.604 |
| Qwen3.5-27B | 21% | 39.6 | 42.8 | 1.759 |
| GPT-4o | 13% | 28.2 | 54.6 | 2.911 |
| Gemini 2.5 Pro | 8% | 30.7 | 44.6 | 3.599 |