Humans score 93.1% where best model stops at 53.6% on new visual test
Scale Labs launched the Humanity's Sixth Sense (HSS) benchmark on October 7, 2026, measuring something existing visual tests have barely touched: the fast, intuitive reasoning humans do "at a glance." According to the company's own paper, human participants reach 93.1 percent accuracy, while the strongest multimodal model — GPT-6-astra — manages only 53.6 percent, even at maximum reasoning effort.
What HSS actually tests
According to the Scale Labs paper, HSS comprises diverse image and video inputs, organized under a structured taxonomy, with each item paired with human-written tasks. The tasks probe implicit temporal, spatial, social and abstract structure — the things people read out of an image effortlessly: what just happened, how elements relate to each other in space, what a situation is about socially, and which abstract patterns underlie it.
The name makes the point: this is an ability Scale Labs describes as a "sixth sense" — something humans do automatically, but which, according to the company, has never been made into a measurable capability axis for machines. The paper concludes that HSS "establishes intuitive visual reasoning as a measurable axis," drawing attention to a capability that scaling on existing benchmarks "has so far left behind."
Why existing benchmarks don't capture this
Scale Labs' own explanation for why this territory has gone untested is that existing visual benchmarks either measure "deliberate expert-level analysis" or "low-level perception" — not the intuitive reasoning people actually perform in everyday life. In other words, most tests target either heavy, sustained analysis of complex material, or basic recognition of objects and text. The intuition in between — understanding a situation quickly without deliberate reasoning — has, according to the company, largely not been tested at all.
That is also part of the news value here. Frontier models score very highly on many existing visual benchmarks, but according to Scale Labs' paper that does not necessarily say anything about this kind of ability — suggesting there may be capability axes that today's training and evaluation regimes simply have not captured.
The numbers, and the agentic variant
The headline result, as presented in Scale Labs' material, is thus 93.1 percent for human participants versus 53.6 percent for GPT-6-astra — the strongest model on the benchmark — even when the model runs at maximum reasoning effort. That is a substantial gap, and it is worth emphasizing that these are the company's own reported results, not independently verified.
The paper also explores an agentic variant: a setup that applies dynamic visual manipulation to the HSS tasks. According to Scale Labs, this narrows the gap between humans and models, but does not close it. That is an indication that models can compensate somewhat by interacting more actively with the visual material — but the underlying weakness in intuitive reasoning does not disappear. The details of this agentic setup are not elaborated in the available summary, however, so exactly what "dynamic visual manipulation" involves in practice, and which models benefit most from it, remains to be seen.
What we don't know yet
Several open questions are not answered by the available material from Scale Labs:
- Methodology and dataset size: The number of tasks, where they come from, and how the taxonomy categories are distributed are not specified in the summary.
- Participant pool: How many humans were tested, and who they were, is not stated.
- Per-category results: Whether temporal, spatial, social and abstract reasoning differ systematically between humans and models does not emerge from the available material.
- Independent verification: The figures, and the name "GPT-6-astra" — and how this model relates to other frontier models — have not been verified outside Scale Labs' own reporting.
These gaps are not necessarily a sign that the study is weak — they are simply details contained in the full paper, not on the summary page where it is presented. But until there are independent replications or a review of the methodology, the numbers should be read as the company's own measurement.
Why it matters
If Scale Labs' results hold up, they point to something important in the debate over how far AI has actually come: that progress on existing benchmarks does not necessarily translate into broad, human-like abilities. A model can perform at expert level on testable, well-defined tasks while falling far short on something as seemingly basic as understanding a situation "at a glance."
HSS claims to measure such an axis. Whether it actually captures a real, general capability — or whether it can be "gamed" with more training and better architectures — is a question the coming months of research and replication will have to answer.

