mhevalSubmit results

Quality

How well a model helps: clinical skill, empathy and accuracy in therapy and mental-health conversations.

Model
Claude Opus 5.5Anthropic · Public API86%0.5243.9760.63583.75.0000.4860.632
GPT-6 AstraOpenAI · Public API77%0.5783.8890.56680.55.0000.4330.606
GPT-6 LunaOpenAI · Public API64%0.5033.7620.50978.75.0000.4110.603
GPT-5OpenAI · Public API57%0.4443.5120.66277.75.0000.4940.617
Qwen3.5-122B-A10BQwen · Open weights46%0.3703.6230.54666.25.0000.4890.609
Gemini 2.5 ProGoogle · Public API43%0.2953.7560.48474.35.0000.5360.630
Claude Sonnet 4.5Anthropic · Public API32%0.4023.7180.43372.45.0000.5670.600
Qwen3.5-27BQwen · Open weights25%0.3673.5980.54165.14.9900.4890.592
GPT-4oOpenAI · Public API20%0.3263.1690.27639.84.9600.3670.602
Win rate is the share of head-to-head comparisons a model wins across these benchmarks. Cells are shaded by rank within each column and the best score is in bold; arrows show which direction is better. Select a column to sort.

Safety

How well a model protects: recognising risk, resisting delusions and caring for vulnerable users.

Model
GPT-6 AstraOpenAI · Public API100%81.280.51.027
GPT-6 LunaOpenAI · Public API79%80.776.31.124
Claude Opus 5.5Anthropic · Public API75%74.174.91.050
Claude Sonnet 4.5Anthropic · Public API67%56.066.91.031
GPT-5OpenAI · Public API46%62.360.82.158
Qwen3.5-122B-A10BQwen · Open weights42%40.556.41.604
Qwen3.5-27BQwen · Open weights21%39.642.81.759
GPT-4oOpenAI · Public API13%28.254.62.911
Gemini 2.5 ProGoogle · Public API8%30.744.63.599
Read as above. SIM-VAIL measures harm, so lower is better.