Scale Labs
[PAPERS][BLOG][LEADERBOARDS][SHOWDOWN]
Scale Labs

Research to Advance AI

Scale Labs advances AI through research. Our research focuses on agents, post-training, reasoning, safety, evaluation, and alignment, and the science of data.

[LEADERBOARDS]

Benchmarks for frontier, agentic, and safety capabilities

Humanity's Sixth SenseHumanity's Last Exam (Diamond)DrugDiscoveryBenchSWE Atlas - RefactoringSWE Atlas - Test Writing
View more

[SHOWDOWN]

Model-preference rankings from real-world usage.

1claude-opus-4-645.5K votes1071.52
1gpt-5.2-chat-latest55.2K votes1069.78
1claude-opus-4-7 (Thinking)6.1K votes1063.81
1claude-opus-4-75.4K votes1062.78
3gpt-5.5-2026-04-235.3K votes1052.93
View more

[PAPERS]

Research papers and publications covering agents, post-training, reasoning, safety, evaluation, and alignment, and the science of data.

Date Title Category Authors
Date Title
10/7/2026Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal ModelsMultimodalXingang Guo, Jing Gu, Brian Jang, Renxiong Wang, Utkarsh Tyagi, Daniel Quigley, Steven Li, David Yan, Daniel Yue Zhang, Darvin Yi, Forrest Huang, HiJae Kim, Tianyi Zhang, Jared Lichtarge, Jihua Huang, Le Xue, Manan Tomar, Qiuyi Richard Zhang, Ruofei Yu, Seth Neel, Yaning Hu, Marcella Valentine, Xinzhe Jiang, Daniel Evans, Chenguang Wang, Dustin Tran, Tong Zhao, Yinfei Yang, Yunzhong He9/17/2026SteerDuplex: Steerable Duplex Speech Dialogue ModelsMultimodalUtkarsh Tyagi, Ramaneswaran Selvakumar, Advait Gosai, Sonal Kumar, Nikhil Barhate, Isabell Sagar, Steven Li, Miheer Bavare, Daniel Quigley, Fabiola Tapia Carrillo, Jose M Patron E, Diego Macías Gutiérrez, Paul Song, Ramani Duraiswami, Dinesh Manocha, Yunzhong He9/4/2026READY or Not: Reliable Enterprise Agent DeploymentAgents, EnterpriseVeronica Chatrath, Bryan Zhu, Jingxuan Fan, George Pu, Soham Dinesh Tiwari, Soham Dan, Ryan Young, Yuan (Christy) Li, Yuang Yao, Apaar Shanker, Minglai Yang , Daniel Yue Zhang, Yunzhong He, Ying Liu, Chenguang Wang, Zhijun Yin, Yuan (Emily) Xue8/27/2026Spine-Branch Coordination for Multi-agent Computer UseAgentsMian Zhang, Manasi Sharma, Sheng Zhang, Minlai Yang, Kejian Shi, Ying Liu, Zhiyu Zoey Chen, Daniel Yue Zhang8/20/2026CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHREvaluation and AlignmentVeronica Chatrath, Bryan Zhu, George Pu, Jingxuan Fan, Apaar Shanker, Varun Ursekar, Anahita Sharma, Jason Qin, Keqi Han, Soham Dinesh Tiwari, Soham Dan, Vijay Kalmath, Yuan (Christy) Li, Daniel Yue Zhang, Chenguang Wang, Zainab Doctor, Zhijun Yin, Nigam H. Shah, Yuan Xue8/6/2026HarnessOpt-Bench: Evaluating LLMs at Harness OptimizationAgents, Enterprise, Evaluation and AlignmentVarun Ursekar, Apaar Shanker, Yash Maurya, Shehab Yasser, Vijay Kalmath, Veronica Chatrath, Yuan (Emily) Xue
10/7/2026
Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal ModelsMultimodal
9/17/2026
SteerDuplex: Steerable Duplex Speech Dialogue ModelsMultimodal
9/4/2026
READY or Not: Reliable Enterprise Agent DeploymentAgents, Enterprise
8/27/2026
Spine-Branch Coordination for Multi-agent Computer UseAgents
8/20/2026
CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHREvaluation and Alignment
8/6/2026
HarnessOpt-Bench: Evaluating LLMs at Harness OptimizationAgents, Enterprise, Evaluation and Alignment
View more
Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models
SteerDuplex: Steerable Duplex Speech Dialogue Models
READY or Not: Reliable Enterprise Agent Deployment
Spine-Branch Coordination for Multi-agent Computer Use
CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models

[BLOG]

Insights, analysis, and updates from Scale Labs

Oct 7, 2026

Introducing Humanity's Sixth Sense: Measuring Intuitive Visual Reasoning

Today we're introducing Humanity's Sixth Sense (HSS), in partnership with Elorian, a benchmark for intuitive visual reasoning with 522 open-ended tasks across 288 images and 234 video clips (17.6 hours of video in total).

Oct 5, 2026

Introducing AgentEnv: An Open-Source Framework for Building RL Environments

An agent is only as good as the world it practices in, and building worlds that are realistic, accurate, and scalable all at once is the hard part of scaling RL. Every RL environment Scale builds runs on AgentEnv. Today we're open-sourcing it.

Evaluation and AlignmentSep 22, 2026

SWE-Bench Pro V2: A Cleaner, Harder-to-Game Leaderboard

When SWE Bench Pro originally launched a little over a year ago, it was a significant entry into the open source benchmark space and changed how the field measured coding agents. However, we know that benchmarks built on open source GitHub issues have real limitations. We see that as a reason to keep iterating on SWE Bench Pro rather than retire it.

Evaluation and AlignmentSep 10, 2026

RUBRIC DROPOUT: A SIMPLE WAY TO MITIGATE REWARD HACKING IN RUBRIC-AS-REWARD RL

Rubric-based RL can quietly learn to game its reward: the training judge keeps assigning higher scores even as true quality declines. A one-line fix, inspired by neural-network dropout, mitigates the problem at virtually no additional cost.

View allAll posts

[Jobs]

Date Position Location
10.06.2026
COLM 2026 - General InterestSan Francisco, CA; Seattle, WA; New York, NY · 10.06.2026
San Francisco, CA; Seattle, WA; New York, NY
10.06.2026
Machine Learning Research Scientist, EvaluationsSan Francisco, CA; Seattle, WA; New York, NY · 10.06.2026
San Francisco, CA; Seattle, WA; New York, NY
10.06.2026
Machine Learning Research Scientist, Post-TrainingSan Francisco, CA; Seattle, WA; New York, NY · 10.06.2026
San Francisco, CA; Seattle, WA; New York, NY
10.06.2026
ML Research Engineer, ML SystemsSan Francisco, CA; Seattle, WA; New York, NY · 10.06.2026
San Francisco, CA; Seattle, WA; New York, NY
10.06.2026
Research Scientist, Agent RobustnessSan Francisco, CA; New York, NY · 10.06.2026
San Francisco, CA; New York, NY
10.06.2026
Research Scientist, AI Controls and MonitoringSan Francisco, CA; New York, NY · 10.06.2026
San Francisco, CA; New York, NY
View more

Copyright 2026 Scale Inc. All rights reserved.

TermsPrivacy