Research to Advance AI
Scale Labs advances AI through research. Our research focuses on agents, post-training, reasoning, safety, evaluation, and alignment, and the science of data.
[LEADERBOARDS]
Benchmarks for frontier, agentic, and safety capabilities
[SHOWDOWN]
Model-preference rankings from real-world usage.
[PAPERS]
Research papers and publications covering agents, post-training, reasoning, safety, evaluation, and alignment, and the science of data.






SteerDuplex: Steerable Duplex Speech Dialogue Models
[BLOG]
Insights, analysis, and updates from Scale Labs
Verification in RSI Bench
Benchmarks are useful if an increase in their score reflects a real improvement in the capabilities they intend to measure. This connection can break in several ways: a proxy reward that doesn’t measure the intended capability, weaknesses in an evaluator that can be exploited, or trade off one capability for another. Therefore, a higher score can indicate real progress or merely expose a weakness in how progress is measured.
RUBRIC DROPOUT: A SIMPLE WAY TO MITIGATE REWARD HACKING IN RUBRIC-AS-REWARD RL
Rubric-based RL can quietly learn to game its reward: the training judge keeps assigning higher scores even as true quality declines. A one-line fix, inspired by neural-network dropout, mitigates the problem at virtually no additional cost.
Introducing READY: What It Takes to Deploy an AI Agent
Today we're introducing READY (Reliable Enterprise Agent Deployment), a suite of industry-specific benchmarks that measure agents the way enterprises actually use them: working alongside people, inside real workflows.
Who Grades the Graders? Rethinking Verifier Design for Computer Use Agents
As Computer Use Agents (CUA) take on complex professional tasks like processing emails, creating spreadsheets, drawing 3D diagrams, and drafting financial memos, verifiers are at the heart of providing evaluation and training signals around their capabilities. Benchmarks like OSWorld 2.0 and Agents' Last Exam use these verifiers, in the form of programmatic checks, to measure whether an agent is capable of completing production-grade work.
View allAll posts