Quality benchmark
CounselBench Adv
Expert-written questions built to provoke known failures, from judgmental replies to unprompted medication advice.
“An adversarial dataset of 120 expert-authored mental health questions designed to trigger specific model issues.”
- Metric
- Failure rate
- Scale
- 0–1
- Source
- Paper ↗
Headline score
Failure rate by targeted issue
| Model | ||||||
|---|---|---|---|---|---|---|
| Qwen3.5-122B-A10B | 0.000 | 0.617 | 0.833 | 0.100 | 0.500 | 0.883 |
| Claude Opus 5.5 | 0.000 | 0.267 | 0.883 | 0.100 | 0.783 | 0.883 |
| Gemini 2.5 Pro | 0.000 | 0.767 | 0.850 | 0.133 | 0.617 | 0.850 |
| GPT-6 Astra | 0.000 | 0.133 | 0.850 | 0.050 | 0.733 | 0.833 |
| Qwen3.5-27B | 0.000 | 0.600 | 0.800 | 0.117 | 0.600 | 0.817 |
| GPT-5 | 0.017 | 0.183 | 0.883 | 0.017 | 0.950 | 0.917 |
| GPT-6 Luna | 0.033 | 0.183 | 0.850 | 0.050 | 0.617 | 0.733 |
| Claude Sonnet 4.5 | 0.067 | 0.633 | 0.867 | 0.217 | 0.767 | 0.850 |
| GPT-4o | 0.067 | 0.417 | 0.767 | 0.033 | 0.400 | 0.517 |