Quality benchmark
HealthBench-Psych
The mental-health conversations from HealthBench, graded against conversation-specific rubrics written by physicians.
“Responses are evaluated using conversation-specific rubrics created by 262 physicians.”
- Metric
- Clipped mean score
- Scale
- 0–1
- Source
- Paper ↗
Headline score
Axis
| Model | |||||
|---|---|---|---|---|---|
| Claude Opus 5.5 | 0.737 | 0.701 | 0.637 | 0.450 | 0.630 |
| GPT-5 | 0.713 | 0.638 | 0.659 | 0.549 | 0.602 |
| Qwen3.5-122B-A10B | 0.702 | 0.670 | 0.571 | 0.239 | 0.627 |
| Qwen3.5-27B | 0.701 | 0.671 | 0.552 | 0.262 | 0.612 |
| GPT-6 Astra | 0.691 | 0.718 | 0.511 | 0.454 | 0.617 |
| Gemini 2.5 Pro | 0.675 | 0.623 | 0.489 | 0.234 | 0.632 |
| GPT-6 Luna | 0.612 | 0.691 | 0.454 | 0.367 | 0.657 |
| Claude Sonnet 4.5 | 0.595 | 0.685 | 0.373 | 0.257 | 0.562 |
| GPT-4o | 0.466 | 0.602 | 0.171 | 0.129 | 0.509 |
Theme
| Model | |||||||
|---|---|---|---|---|---|---|---|
| GPT-5 | 0.730 | 0.700 | 0.619 | 0.797 | 0.591 | 0.582 | 0.665 |
| Claude Opus 5.5 | 0.690 | 0.656 | 0.550 | 0.764 | 0.547 | 0.651 | 0.648 |
| Qwen3.5-122B-A10B | 0.659 | 0.534 | 0.409 | 0.592 | 0.515 | 0.464 | 0.589 |
| Qwen3.5-27B | 0.628 | 0.470 | 0.432 | 0.613 | 0.513 | 0.423 | 0.600 |
| GPT-6 Astra | 0.612 | 0.569 | 0.502 | 0.736 | 0.452 | 0.512 | 0.609 |
| Gemini 2.5 Pro | 0.601 | 0.545 | 0.357 | 0.576 | 0.407 | 0.377 | 0.528 |
| GPT-6 Luna | 0.581 | 0.459 | 0.468 | 0.664 | 0.396 | 0.470 | 0.523 |
| Claude Sonnet 4.5 | 0.493 | 0.491 | 0.303 | 0.618 | 0.312 | 0.374 | 0.502 |
| GPT-4o | 0.407 | 0.327 | 0.095 | 0.405 | 0.188 | 0.165 | 0.344 |