People increasingly turn to chatbots for emotional support, and some of those conversations involve suicide or self-harm.
DistressBench measures whether models give the support those conversations need, using rubrics written by licensed clinicians: 718 clinician-authored conversations across 23 suicide and self-harm subcategories.
Most existing safety benchmarks score these conversations on whether the model refused, which says little about whether the user actually got help. DistressBench contains 718 clinician-authored conversations across 23 suicide and self-harm subcategories. Each one comes with a rubric of weighted criteria (6,901 in total) covering seven dimensions of crisis support. Across 25 frontier models, the best model meets 88.4% of weighted criteria and the median model meets 71.4%, compared with 98.7% for clinician reference responses. The biggest gap is between recognizing a crisis and acting on it. In 35.3% of conversations where a model fully recognized the crisis, it still didn’t point the user toward human help. Models score near the top on avoiding harmful content (94.5% median), but the median model reaches only 46.0% on de-escalation while the strongest reaches 74.6%. That spread suggests the gap mostly comes from product choices, since at least one deployed model already does much better.