Quality benchmark
MentalHealthBench
How well a model handles realistic mental-health conversations, from everyday well-being to emergencies, graded against rubrics written by licensed clinicians.
“MentalHealthBench is designed to better represent real-world AI mental health usage, spanning the range of acuities, conversational topics, and user profiles such as teens, adults, caregivers, and clinicians across multiple languages and cultural contexts.”
- Metric
- Clipped task score
- Scale
- 0–1
- Source
- Paper ↗
Headline score
Acuity
| Model | |||
|---|---|---|---|
| GPT-6 Astra | 0.580 | 0.593 | 0.571 |
| Claude Opus 5.5 | 0.537 | 0.533 | 0.514 |
| GPT-6 Luna | 0.531 | 0.546 | 0.474 |
| GPT-5 | 0.461 | 0.511 | 0.412 |
| Claude Sonnet 4.5 | 0.404 | 0.448 | 0.386 |
| Qwen3.5-122B-A10B | 0.381 | 0.395 | 0.357 |
| Qwen3.5-27B | 0.371 | 0.402 | 0.353 |
| Gemini 2.5 Pro | 0.274 | 0.347 | 0.289 |
| GPT-4o | 0.244 | 0.374 | 0.353 |
User profile
| Model | ||||
|---|---|---|---|---|
| GPT-6 Astra | 0.570 | 0.646 | 0.640 | 0.570 |
| Claude Opus 5.5 | 0.506 | 0.629 | 0.511 | 0.561 |
| GPT-6 Luna | 0.495 | 0.562 | 0.562 | 0.501 |
| GPT-5 | 0.415 | 0.604 | 0.505 | 0.481 |
| Claude Sonnet 4.5 | 0.385 | 0.504 | 0.435 | 0.426 |
| Qwen3.5-122B-A10B | 0.346 | 0.434 | 0.497 | 0.401 |
| Qwen3.5-27B | 0.346 | 0.440 | 0.477 | 0.387 |
| GPT-4o | 0.312 | 0.363 | 0.416 | 0.337 |
| Gemini 2.5 Pro | 0.276 | 0.344 | 0.466 | 0.298 |
Behavior axis (signed score)
| Model | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| GPT-6 Astra | 0.133 | 0.203 | 0.080 | -0.011 | 0.285 | -0.015 | 0.254 | 0.090 | 0.114 | 0.182 |
| GPT-6 Luna | 0.118 | 0.186 | 0.076 | -0.004 | 0.183 | -0.008 | 0.255 | 0.082 | 0.103 | 0.151 |
| Claude Sonnet 4.5 | 0.117 | 0.056 | 0.050 | -0.103 | 0.180 | -0.213 | 0.245 | -0.179 | 0.082 | -0.144 |
| Claude Opus 5.5 | 0.108 | 0.116 | 0.102 | -0.108 | 0.313 | -0.041 | 0.259 | -0.065 | 0.092 | 0.007 |
| Qwen3.5-122B-A10B | 0.100 | 0.089 | 0.043 | -0.161 | 0.102 | -0.139 | 0.248 | -0.195 | 0.085 | -0.188 |
| Qwen3.5-27B | 0.099 | 0.090 | 0.051 | -0.172 | 0.101 | -0.165 | 0.242 | -0.198 | 0.077 | -0.196 |
| GPT-4o | 0.086 | 0.097 | 0.085 | -0.047 | 0.040 | -0.312 | 0.234 | -0.029 | 0.053 | -0.106 |
| Gemini 2.5 Pro | 0.059 | 0.039 | 0.018 | -0.149 | 0.058 | -0.349 | 0.230 | -0.263 | 0.088 | -0.304 |
| GPT-5 | 0.030 | 0.092 | 0.087 | -0.107 | 0.221 | -0.114 | 0.243 | -0.046 | 0.041 | -0.113 |