Safety benchmark
VERA-MH
How a model recognises and responds to suicide risk in conversations with clinician-designed personas.
“An automated evaluation of the safety of AI chatbots used in mental health contexts, with an initial focus on suicide risk.”
- Metric
- VERA score
- Scale
- 0–100
- Source
- Paper ↗
Headline score
Dimension
| Model | |||||
|---|---|---|---|---|---|
| GPT-4o | 96.7 | 0.8 | 4.1 | 51.1 | 47.3 |
| Claude Sonnet 4.5 | 95.8 | 28.3 | 54.1 | 63.3 | 47.6 |
| GPT-6 Astra | 95.0 | 100.0 | 81.4 | 88.9 | 44.5 |
| Claude Opus 5.5 | 93.7 | 74.2 | 65.4 | 90.7 | 47.7 |
| Gemini 2.5 Pro | 91.5 | 1.3 | 12.7 | 45.1 | 50.3 |
| GPT-5 | 89.5 | 88.9 | 37.7 | 52.8 | 44.5 |
| Qwen3.5-122B-A10B | 88.4 | 5.0 | 26.4 | 47.7 | 63.7 |
| GPT-6 Luna | 86.4 | 97.5 | 86.3 | 89.0 | 46.6 |
| Qwen3.5-27B | 85.1 | 5.8 | 26.4 | 46.6 | 59.8 |