Safety benchmark
Spiral-Bench
Whether a model pushes back or plays along when a suggestible user spirals into delusional thinking.
“A multiturn, roleplay-based benchmark designed to measure protective and risky behaviours in large language models (LLMs) when interacting with a suggestible, seeker-type user.”
- Metric
- Overall score
- Scale
- 0–100
- Source
- GitHub ↗
Headline score
Protective behaviors
| Model | |||||
|---|---|---|---|---|---|
| GPT-6 Astra | 2.780 | 1.777 | 0.210 | 0.777 | 0.922 |
| Claude Opus 5.5 | 2.033 | 1.779 | 0.222 | 0.544 | 0.649 |
| GPT-6 Luna | 1.909 | 1.671 | 0.129 | 0.833 | 0.629 |
| Claude Sonnet 4.5 | 1.230 | 1.502 | 0.493 | 0.417 | 0.653 |
| Qwen3.5-122B-A10B | 0.460 | 2.084 | 0.075 | 0.104 | 0.556 |
| Qwen3.5-27B | 0.317 | 1.671 | 0.048 | 0.009 | 0.376 |
| GPT-5 | 0.292 | 1.393 | 0.248 | 0.053 | 0.215 |
| Gemini 2.5 Pro | 0.192 | 1.305 | 0.059 | 0.016 | 0.257 |
| GPT-4o | 0.086 | 0.460 | 0.018 | 0.014 | 0.028 |
Risky behaviors
| Model | ||||||
|---|---|---|---|---|---|---|
| GPT-6 Luna | 0.000 | 0.000 | 0.018 | 0.000 | 0.000 | 0.000 |
| GPT-6 Astra | 0.004 | 0.004 | 0.044 | 0.000 | 0.005 | 0.000 |
| GPT-5 | 0.034 | 0.155 | 0.048 | 0.007 | 0.075 | 0.004 |
| Claude Opus 5.5 | 0.039 | 0.000 | 0.035 | 0.000 | 0.012 | 0.016 |
| GPT-4o | 0.714 | 0.233 | 0.074 | 0.000 | 0.039 | 0.007 |
| Claude Sonnet 4.5 | 0.816 | 0.349 | 1.325 | 0.017 | 0.910 | 0.355 |
| Qwen3.5-122B-A10B | 1.540 | 1.233 | 1.176 | 0.100 | 1.719 | 0.111 |
| Qwen3.5-27B | 2.558 | 2.152 | 2.457 | 0.268 | 3.214 | 0.094 |
| Gemini 2.5 Pro | 3.288 | 1.625 | 2.054 | 0.086 | 1.969 | 0.090 |