Today AI models demonstrate high visual IQ: they read charts, ace exams, and extract text easily. However, they struggle on intuitive visual reasoning: it is difficult for them to read a room and effortlessly pick up on nuances.
People can. A single glance tells us who holds authority in a room, whether a car will fit between two parked ones, or why a dog might get startled by birds. We do it effortlessly, without conscious reasoning. What makes this possible is human intuition: a latent understanding, built from experience, of social dynamics, spatial relationships, physical cause and effect, and intent.
For AI agents deployed in homes, vehicles, and workplaces, this skill is not a luxury, but a core necessity. Yet most existing visual benchmarks test either expert-level academic analysis or low-level perception, leaving the intuitive reasoning that people perform instinctively largely untested.
Introducing HSS
Today we're introducing Humanity's Sixth Sense (HSS), in partnership with Elorian, a benchmark for intuitive visual reasoning with 522 open-ended tasks across 288 images and 234 video clips (17.6 hours of video in total).
The main finding: Models miss what any person picks up with a glance. People score 93.1% on these visual reasoning tasks, while the strongest model, GPT-6-astra, reaches only 53.6%. The median model scores 30.9%.
Each task pairs a scene with a human-written prompt probing the implicit structure that people infer at a glance, organized into four domains and eleven subdomains:
- Temporal & Causal Dynamics: Recovering the moment behind or ahead of the one depicted. Covers retrodiction (what already happened), mechanistic causality (what is physically driving an ongoing event), and extrapolation (what happens next).
- Physical & Spatial Logic: Recovering structure the scene doesn't show. Covers hidden and invisible properties (occluded objects, mass, invisible forces), affordance and feasibility (will it fit), alien viewpoint (what the scene looks like from somewhere the camera never stood), and spatial reachability (can a person or object get there).
- Social Understanding: Recovering the mental and normative layer of a scene. Covers theory of mind (beliefs, knowledge gaps, emotions) and social role, norm, and power dynamics (unwritten rules and who defers to whom).
- Abstract & Contextual Inference: Recovering the rule that gives a scene meaning. Covers change and consequence (what follows from a change the scene doesn't contain) and patterns and pareidolia (structure in arrangements, faces in objects).
How it Works
This capacity relies on visual intuition, a concept deeply rooted in cognitive science. People categorize natural scenes within roughly 150 milliseconds while extracting physical dynamics, motives, and social structure.
HSS is built to test that kind of inference: fast, grounded in everyday experience, and difficult to articulate step by step.
- Task authoring: Trained annotators select an image or video clip, assign it to a subdomain, and write a question, a reference answer, and a rubric of atomic criteria that a correct answer must satisfy. Every task follows three rules. It must require inference beyond what is depicted (if a high-resolution crop makes the answer obvious, it's rejected). It must need no specialized expertise. And its answer must command unanimous human agreement.
- Independent review: Each task passes through three independent review rounds to reduce subjectivity, where reviewers can pass, return for revision, or drop it. Of 3,466 authored tasks, 522 survived, an acceptance rate of 15.1%. The first round alone removed two in three tasks, mostly because the answer was visible in the scene or reviewers could not agree on it.
- Evaluation: Models answer in free form, with no multiple choice to eliminate options against. An LLM judge scores each answer against the rubric, and a task counts as solved only when every criterion is met. We re-graded five models with judges from three vendors: rankings were identical under all three (ρ = 1.0) and agreement exceeded 95%.
Research Findings
We evaluated 25 multimodal models from eight vendors alongside 20 human participants, and five things stood out:
- Models reach barely half of human performance. Human participants reach 93.1% accuracy, while the strongest model, GPT-6-astra, reaches only 53.6% even at maximum reasoning effort. The median model scores 30.9%. With the image or video removed, GPT-6-astra drops to 6.6%, confirming the tasks cannot be solved from language priors.
- More thinking does not close the gap. Models spend an average of 4,046 reasoning tokens per task on questions people answer at a glance. Raising reasoning effort helps overall but hurts in some subdomains: GPT-6-astra drops 14 points on retrodiction moving from high to xhigh effort, and peak accuracy often comes at an intermediate setting rather than the maximum.
- Models fail at seeing and inferring, not at reasoning. Across 8,573 failures, 94% trace to perception or latent inference, and only 5% to faulty logic. The two largest causes are missing the decisive visual cue (21%) and misidentifying an object, person, or role (20%). These failures are systematic: on tasks that five or more models fail, a median of 80% fail for the same reason.
- Social understanding is the weakest domain. It is the lowest-scoring domain for 21 of 25 models, averaging 24.4% against 34.1% for the other three. Video is also harder than images for 23 of 25 models, by 7.3 points on average.
- Agentic tools narrow the gap but don't close it. Inside Claude Code and Codex, where models can crop, zoom, web search and re-sample media, the best setup reaches 59.3% on a 388-task subset. Closer inspection fixes missed cues and misidentified objects, but not depth: errors from reading a 2D overlap as 3D alignment were essentially unchanged (102 to 101).
What This Means for Multimodal AI
Humans infer far more from a visual scene than what is explicitly shown, and this rapid intuition underpins everyday navigation and social interaction. Humanity's Sixth Sense (HSS) measures this capability directly.
Frontier models fall well short of humans on HSS, and the gap holds even with more test-time compute or agentic tooling. Because nearly all model errors arise in perception and inference rather than deliberate reasoning, HSS establishes intuitive visual reasoning as a measurable axis of multimodal intelligence, and one that current scaling has so far left behind.
Resources:
- Paper: https://labs.scale.com/papers//humanitys-sixth-sense
- Leaderboard: https://scale.com/leaderboard/hss
- Huggingface: https://huggingface.co/datasets/ScaleAI/HSS