Research to Advance AI
Scale Labs advances AI through research. Our research focuses on agents, post-training, reasoning, safety, evaluation, and alignment, and the science of data.
[LEADERBOARDS]
Benchmarks for frontier, agentic, and safety capabilities
[SHOWDOWN]
Model-preference rankings from real-world usage.
[PAPERS]
Research papers and publications covering agents, post-training, reasoning, safety, evaluation, and alignment, and the science of data.






DistressBench: Evaluating Multi-Dimensional Crisis Support in Large Language Models
[BLOG]
Insights, analysis, and updates from Scale Labs
DistressBench: A clinician-authored benchmark for how AI models handle users in distress
People increasingly turn to chatbots for emotional support, and some of those conversations involve suicide or self-harm. DistressBench measures whether models give the support those conversations need, using rubrics written by licensed clinicians: 718 clinician-authored conversations across 23 suicide and self-harm subcategories.
Introducing Humanity's Sixth Sense: Measuring Intuitive Visual Reasoning
Today we're introducing Humanity's Sixth Sense (HSS), in partnership with Elorian, a benchmark for intuitive visual reasoning with 522 open-ended tasks across 288 images and 234 video clips (17.6 hours of video in total).
Introducing AgentEnv: An Open-Source Framework for Building RL Environments
An agent is only as good as the world it practices in, and building worlds that are realistic, accurate, and scalable all at once is the hard part of scaling RL. Every RL environment Scale builds runs on AgentEnv. Today we're open-sourcing it.
SWE-Bench Pro V2: A Cleaner, Harder-to-Game Leaderboard
When SWE Bench Pro originally launched a little over a year ago, it was a significant entry into the open source benchmark space and changed how the field measured coding agents. However, we know that benchmarks built on open source GitHub issues have real limitations. We see that as a reason to keep iterating on SWE Bench Pro rather than retire it.
View allAll posts