Scale Labs
[PAPERS][BLOG][LEADERBOARDS][SHOWDOWN]
Scale Labs

Research to Advance AI

Scale Labs advances AI through research. Our research focuses on agents, post-training, reasoning, safety, evaluation, and alignment, and the science of data.

[LEADERBOARDS]

Benchmarks for frontier, agentic, and safety capabilities

DrugDiscoveryBenchSWE Atlas - RefactoringSWE Atlas - Test WritingSWE Atlas - Codebase QnAHiL-Bench (Human-in-Loop Benchmark)
View more

[SHOWDOWN]

Model-preference rankings from real-world usage.

1claude-opus-4-645.5K votes1071.45
1gpt-5.2-chat-latest55.2K votes1069.68
1claude-opus-4-7 (Thinking)6.1K votes1064.83
1claude-opus-4-75.4K votes1062.73
3gpt-5.5-2026-04-235.3K votes1053.44
View more

[PAPERS]

Research papers and publications covering agents, post-training, reasoning, safety, evaluation, and alignment, and the science of data.

Date Title Category Authors
Date Title
9/4/2026READY or Not: Reliable Enterprise Agent DeploymentAgents, EnterpriseVeronica Chatrath, Bryan Zhu, Jingxuan Fan, George Pu, Soham Dinesh Tiwari, Soham Dan, Ryan Young, Yuan (Christy) Li, Yuang Yao, Apaar Shanker, Minglai Yang , Daniel Yue Zhang, Yunzhong He, Ying Liu, Chenguang Wang, Zhijun Yin, Yuan (Emily) Xue8/20/2026CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHREvaluation and AlignmentVeronica Chatrath, Bryan Zhu, George Pu, Jingxuan Fan, Apaar Shanker, Varun Ursekar, Anahita Sharma, Jason Qin, Keqi Han, Soham Dinesh Tiwari, Soham Dan, Vijay Kalmath, Yuan (Christy) Li, Daniel Yue Zhang, Chenguang Wang, Zainab Doctor, Zhijun Yin, Nigam H. Shah, Yuan Xue8/6/2026HarnessOpt-Bench: Evaluating LLMs at Harness OptimizationAgents, Enterprise, Evaluation and AlignmentVarun Ursekar, Apaar Shanker, Yash Maurya, Shehab Yasser, Vijay Kalmath, Veronica Chatrath, Yuan (Emily) Xue6/30/2026DrugDiscoveryBench: Can Coding Agents Assist Early-Stage Drug Discovery?Agents, Enterprise, Evaluation and AlignmentAfra Feyza Akyürek, Xinming Tu, Alec Gutmanstein, Jason Qin, Divyansh Agarwal, Sofia Monasdotter, Sergey Chekhov, Brenda Hernandez Villegas, Kirill Chugunov, Judah Engel, Veronica Chatrath, Oscar Kavanagh, Geobio Boo, Ernesto Hernandez, Ying Liu, Yuan (Emily) Xue, Aakash Sabharwal, Daniel Yue Zhang, Zainab Doctor, Yuanhao Qu, Yunzhong He, Sami Hassaan6/29/2026SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding SessionsAgents, Evaluation and AlignmentMohit Raghavendra, Anisha Gunjal, Aakash Sabharwal, Yunzhong He6/19/2026ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld TasksAgents, Evaluation and AlignmentVincent Siu, Manasi Sharma, Dawn Song, Daniel Yue Zhang, Chenguang Wang
9/4/2026
READY or Not: Reliable Enterprise Agent DeploymentAgents, Enterprise
8/20/2026
CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHREvaluation and Alignment
8/6/2026
HarnessOpt-Bench: Evaluating LLMs at Harness OptimizationAgents, Enterprise, Evaluation and Alignment
6/30/2026
DrugDiscoveryBench: Can Coding Agents Assist Early-Stage Drug Discovery?Agents, Enterprise, Evaluation and Alignment
6/29/2026
SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding SessionsAgents, Evaluation and Alignment
6/19/2026
ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld TasksAgents, Evaluation and Alignment
View more
READY or Not: Reliable Enterprise Agent Deployment
CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization
DrugDiscoveryBench: Can Coding Agents Assist Early-Stage Drug Discovery?
SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions
ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks

READY or Not: Reliable Enterprise Agent Deployment

[BLOG]

Insights, analysis, and updates from Scale Labs

AgentsSep 2, 2026

Who Grades the Graders? Rethinking Verifier Design for Computer-Use Agents

As Computer Use Agents (CUA) take on complex professional tasks like processing emails, creating spreadsheets, drawing 3D diagrams, and drafting financial memos, verifiers are at the heart of providing evaluation and training signals around their capabilities. Benchmarks like OSWorld 2.0 and Agents' Last Exam use these verifiers, in the form of programmatic checks, to measure whether an agent is capable of completing production-grade work.

AgentsAug 26, 2026

CliniCARE-Bench: Clinical AI Agents Can Be Right for the Wrong Reasons

We evaluated 16 agentic systems on 25 clinical care scenarios over 750 real patient cases. Every single system committed to an answer more often than the evidence allowed, and up to one in five correct verdicts rested on an investigation the case authors had explicitly prohibited.

Physical AIAug 13, 2026

Can Robots Learn from Watching Us?

Human demonstrations allow for diverse data collection in the real world, but pose a difficult learning problem due to the large embodiment gap between robots and humans. Being able to effectively learn from high volumes of human interaction data can unlock the future of robotics. Our research demonstrates that adding a limited amount of human data to robot datasets can nearly double performance on unseen environments while only sacrificing a nominal amount of in-distribution performance.

AgentsJul 23, 2026

TERMINAL-BENCH 3.0: Harder Tasks for Better Agents

TERMINAL-BENCH 3.0 is now live with broader coverage, harder tasks, and frontier tracking. Scale contributed the most tasks of any single organization in the launch set.

View allAll posts

[Jobs]

Date Position Location
08.26.2026
Machine Learning Research Scientist, EvaluationsSan Francisco, CA; Seattle, WA; New York, NY · 08.26.2026
San Francisco, CA; Seattle, WA; New York, NY
08.26.2026
Machine Learning Research Scientist, Post-TrainingSan Francisco, CA; Seattle, WA; New York, NY · 08.26.2026
San Francisco, CA; Seattle, WA; New York, NY
08.04.2026
ML Research Engineer, ML SystemsSan Francisco, CA; Seattle, WA; New York, NY · 08.04.2026
San Francisco, CA; Seattle, WA; New York, NY
08.04.2026
Research Scientist, Agent RobustnessSan Francisco, CA; New York, NY · 08.04.2026
San Francisco, CA; New York, NY
08.04.2026
Research Scientist, AI Controls and MonitoringSan Francisco, CA; New York, NY · 08.04.2026
San Francisco, CA; New York, NY
08.04.2026
Research Scientist, Frontier Risk EvaluationsSan Francisco, CA; New York, NY · 08.04.2026
San Francisco, CA; New York, NY
View more

Copyright 2026 Scale Inc. All rights reserved.

TermsPrivacy