Research to Advance AI
Scale Labs advances AI through research. Our research focuses on agents, post-training, reasoning, safety, evaluation, and alignment, and the science of data.
[LEADERBOARDS]
Benchmarks for frontier, agentic, and safety capabilities
[SHOWDOWN]
Model-preference rankings from real-world usage.
[PAPERS]
Research papers and publications covering agents, post-training, reasoning, safety, evaluation, and alignment, and the science of data.






SteerDuplex: Steerable Duplex Speech Dialogue Models
[BLOG]
Insights, analysis, and updates from Scale Labs
Introducing AgentEnv: An Open-Source Framework for Building RL Environments
An agent is only as good as the world it practices in, and building worlds that are realistic, accurate, and scalable all at once is the hard part of scaling RL. Every RL environment Scale builds runs on AgentEnv. Today we're open-sourcing it.
SWE-Bench Pro V2: A Cleaner, Harder-to-Game Leaderboard
When SWE Bench Pro originally launched a little over a year ago, it was a significant entry into the open source benchmark space and changed how the field measured coding agents. However, we know that benchmarks built on open source GitHub issues have real limitations. We see that as a reason to keep iterating on SWE Bench Pro rather than retire it.
RUBRIC DROPOUT: A SIMPLE WAY TO MITIGATE REWARD HACKING IN RUBRIC-AS-REWARD RL
Rubric-based RL can quietly learn to game its reward: the training judge keeps assigning higher scores even as true quality declines. A one-line fix, inspired by neural-network dropout, mitigates the problem at virtually no additional cost.
Introducing READY: What It Takes to Deploy an AI Agent
Today we're introducing READY (Reliable Enterprise Agent Deployment), a suite of industry-specific benchmarks that measure agents the way enterprises actually use them: working alongside people, inside real workflows.
View allAll posts