Safety benchmark
SIM-VAIL
How a model treats simulated users with specific psychiatric vulnerabilities, scored for harm across multi-turn conversations.
“Risk was highest when otherwise supportive chatbot behaviors reinforced the psychological mechanisms underlying the simulated user's vulnerability.”
- Metric
- Mean harm
- Scale
- 1–10
- Source
- Paper ↗
Headline score
Harm by phenotype
| Model | |||||
|---|---|---|---|---|---|
| Claude Sonnet 4.5 | 1.022 | 1.061 | 1.022 | 1.044 | 1.006 |
| GPT-6 Astra | 1.056 | 1.017 | 1.000 | 1.050 | 1.011 |
| Claude Opus 5.5 | 1.067 | 1.078 | 1.017 | 1.067 | 1.022 |
| GPT-6 Luna | 1.239 | 1.083 | 1.011 | 1.078 | 1.211 |
| Qwen3.5-122B-A10B | 1.711 | 1.828 | 1.189 | 1.372 | 1.922 |
| Qwen3.5-27B | 1.789 | 1.950 | 1.333 | 1.644 | 2.078 |
| GPT-5 | 2.322 | 1.628 | 2.411 | 1.944 | 2.483 |
| GPT-4o | 2.894 | 2.489 | 3.306 | 2.572 | 3.294 |
| Gemini 2.5 Pro | 3.678 | 3.306 | 3.994 | 2.378 | 4.639 |
Harm by user intention
| Model | ||||||
|---|---|---|---|---|---|---|
| GPT-6 Luna | 1.007 | 1.187 | 1.027 | 1.320 | 1.020 | 1.187 |
| Claude Opus 5.5 | 1.013 | 1.153 | 1.007 | 1.060 | 1.060 | 1.007 |
| GPT-6 Astra | 1.013 | 1.060 | 1.000 | 1.060 | 1.020 | 1.007 |
| Claude Sonnet 4.5 | 1.020 | 1.107 | 1.007 | 1.013 | 1.007 | 1.033 |
| Qwen3.5-122B-A10B | 1.267 | 2.220 | 1.367 | 2.127 | 1.393 | 1.253 |
| Qwen3.5-27B | 1.407 | 2.353 | 1.680 | 1.807 | 1.340 | 1.967 |
| GPT-5 | 1.607 | 3.000 | 1.813 | 2.353 | 2.407 | 1.767 |
| GPT-4o | 2.220 | 3.093 | 3.340 | 2.647 | 3.000 | 3.167 |
| Gemini 2.5 Pro | 2.553 | 4.153 | 4.587 | 3.120 | 4.053 | 3.127 |