Quality benchmark
CBT-Bench
Whether a model can read a patient's thinking the way a CBT therapist does, by identifying cognitive distortions and core beliefs.
“While LLMs perform well in reciting CBT knowledge, they fall short in complex real-world scenarios requiring deep analysis of patients' cognitive structures.”
- Metric
- Mean weighted F1
- Scale
- 0–1
- Source
- Paper ↗
Headline score
Task
| Model | |||
|---|---|---|---|
| Gemini 2.5 Pro | 0.505 | 0.817 | 0.567 |
| Qwen3.5-122B-A10B | 0.484 | 0.787 | 0.556 |
| GPT-5 | 0.483 | 0.778 | 0.589 |
| Claude Sonnet 4.5 | 0.477 | 0.737 | 0.585 |
| Claude Opus 5.5 | 0.470 | 0.808 | 0.617 |
| GPT-6 Astra | 0.467 | 0.767 | 0.584 |
| GPT-6 Luna | 0.463 | 0.773 | 0.573 |
| Qwen3.5-27B | 0.454 | 0.788 | 0.532 |
| GPT-4o | 0.441 | 0.809 | 0.557 |
Cognitive distortion (F1 per label)
| Model | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Gemini 2.5 Pro | 0.634 | 0.271 | 0.121 | 0.667 | 0.413 | 0.624 | 0.353 | 0.328 | 0.690 | 0.544 |
| Qwen3.5-122B-A10B | 0.607 | 0.362 | 0.069 | 0.675 | 0.427 | 0.535 | 0.275 | 0.373 | 0.619 | 0.523 |
| Qwen3.5-27B | 0.579 | 0.362 | 0.000 | 0.711 | 0.451 | 0.394 | 0.182 | 0.324 | 0.584 | 0.571 |
| GPT-6 Luna | 0.490 | 0.382 | 0.000 | 0.725 | 0.433 | 0.514 | 0.316 | 0.250 | 0.659 | 0.555 |
| Claude Opus 5.5 | 0.479 | 0.400 | 0.000 | 0.725 | 0.380 | 0.528 | 0.340 | 0.286 | 0.697 | 0.559 |
| GPT-6 Astra | 0.455 | 0.368 | 0.000 | 0.704 | 0.447 | 0.528 | 0.308 | 0.382 | 0.660 | 0.536 |
| Claude Sonnet 4.5 | 0.453 | 0.375 | 0.067 | 0.667 | 0.437 | 0.507 | 0.300 | 0.394 | 0.723 | 0.567 |
| GPT-5 | 0.419 | 0.400 | 0.074 | 0.765 | 0.435 | 0.529 | 0.327 | 0.418 | 0.705 | 0.548 |
| GPT-4o | 0.167 | 0.378 | 0.129 | 0.708 | 0.415 | 0.559 | 0.310 | 0.484 | 0.652 | 0.589 |
Primary core belief (F1 per label)
| Model | |||
|---|---|---|---|
| Gemini 2.5 Pro | 0.876 | 0.845 | 0.664 |
| GPT-4o | 0.866 | 0.840 | 0.656 |
| Claude Opus 5.5 | 0.847 | 0.839 | 0.690 |
| Qwen3.5-122B-A10B | 0.832 | 0.800 | 0.682 |
| Qwen3.5-27B | 0.826 | 0.821 | 0.673 |
| GPT-5 | 0.822 | 0.780 | 0.689 |
| GPT-6 Luna | 0.818 | 0.778 | 0.678 |
| GPT-6 Astra | 0.808 | 0.746 | 0.710 |
| Claude Sonnet 4.5 | 0.727 | 0.821 | 0.652 |