Quality benchmark
EQ-Bench 3
Emotional intelligence in role-play and analysis: empathy, insight and social skill in difficult conversations.
“Rather than knowledge-based or short-answer questions, tasks here are multi-turn dialogues or analysis questions that test empathy, social dexterity, and psychological insight.”
- Metric
- Rubric score
- Scale
- 0–100
- Source
- GitHub ↗
Headline score
Scenario type
| Model | |||
|---|---|---|---|
| Claude Opus 5.5 | 84.8 | – | 82.8 |
| GPT-6 Astra | 84.4 | 80.0 | 77.6 |
| GPT-6 Luna | 80.6 | 60.9 | 77.8 |
| GPT-5 | 79.6 | 75.0 | 76.3 |
| Claude Sonnet 4.5 | 75.7 | 65.8 | 70.0 |
| Gemini 2.5 Pro | 71.0 | 71.7 | 76.8 |
| Qwen3.5-122B-A10B | 67.5 | 51.6 | 65.8 |
| Qwen3.5-27B | 66.7 | 39.1 | 65.0 |
| GPT-4o | 36.2 | 29.1 | 43.0 |
Abilities (0-20, make up the rubric score)
| Model | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 5.5 | 16.5 | 16.2 | 18.0 | 15.3 | 17.4 | 14.7 | 16.9 | 17.9 | 16.8 | 16.3 |
| Gemini 2.5 Pro | 15.5 | 15.0 | 16.3 | 14.0 | 15.6 | 13.6 | 13.7 | 14.7 | 14.2 | 13.7 |
| GPT-6 Astra | 15.1 | 16.1 | 17.7 | 13.9 | 17.1 | 12.6 | 16.8 | 17.7 | 16.2 | 16.7 |
| GPT-6 Luna | 14.9 | 15.9 | 17.2 | 14.2 | 16.6 | 12.8 | 16.0 | 16.9 | 15.2 | 16.2 |
| GPT-5 | 14.5 | 16.2 | 16.9 | 14.0 | 16.1 | 13.0 | 15.9 | 16.1 | 16.3 | 15.4 |
| Claude Sonnet 4.5 | 13.8 | 13.2 | 16.8 | 11.7 | 15.6 | 11.6 | 14.8 | 16.8 | 14.5 | 13.8 |
| Qwen3.5-122B-A10B | 13.3 | 12.7 | 15.2 | 11.0 | 14.0 | 10.9 | 12.9 | 14.0 | 13.8 | 12.6 |
| Qwen3.5-27B | 13.2 | 12.3 | 14.8 | 11.2 | 13.8 | 10.5 | 13.0 | 14.2 | 13.2 | 12.4 |
| GPT-4o | 10.3 | 8.3 | 8.3 | 7.5 | 8.4 | 6.5 | 6.8 | 6.1 | 8.1 | 7.9 |
Style traits (0-20, descriptive)
| Model | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 5.5 | 14.5 | 14.0 | 12.4 | 13.0 | 13.1 | 5.9 | 2.7 | 10.6 | 18.1 | 7.9 | 12.9 | 16.3 |
| Gemini 2.5 Pro | 14.3 | 15.1 | 8.8 | 11.4 | 11.4 | 6.5 | 6.0 | 12.7 | 17.2 | 10.5 | 10.9 | 13.4 |
| GPT-4o | 12.5 | 14.3 | 4.2 | 5.6 | 9.7 | 6.9 | 10.9 | 15.2 | 12.0 | 9.5 | 7.0 | 7.1 |
| Qwen3.5-27B | 12.3 | 13.2 | 10.7 | 13.4 | 12.9 | 10.3 | 4.4 | 12.0 | 16.3 | 10.7 | 8.2 | 10.7 |
| Qwen3.5-122B-A10B | 11.9 | 12.8 | 11.6 | 13.8 | 13.1 | 10.9 | 4.8 | 11.6 | 16.7 | 11.4 | 8.3 | 11.2 |
| GPT-5 | 11.9 | 14.2 | 9.9 | 15.1 | 14.6 | 6.3 | 3.2 | 11.8 | 17.9 | 8.0 | 8.6 | 11.5 |
| GPT-6 Luna | 11.2 | 12.5 | 12.7 | 15.2 | 14.5 | 7.0 | 2.2 | 9.6 | 17.7 | 6.2 | 9.0 | 12.7 |
| Claude Sonnet 4.5 | 11.1 | 11.4 | 14.3 | 14.3 | 12.9 | 11.3 | 2.8 | 8.8 | 17.5 | 11.3 | 10.3 | 13.5 |
| GPT-6 Astra | 10.9 | 12.5 | 13.8 | 15.9 | 15.0 | 6.9 | 2.0 | 9.4 | 18.3 | 6.8 | 8.0 | 11.8 |