mit-vista-arc-agi hero

By Jordan Vale

Give a strong AI model a better pair of eyes and a perfect visual notebook, and it can ace a hard public puzzle suite without writing code. MIT researchers — including Kaiming He — describe VISTA, a visual harness that lifts Claude Opus 5.0 from the organizers’ official score of about 40.68 to a perfect 100.00 Relative Human Action Efficiency (RHAE) on all 25 public games in ARC-AGI-3, using 57.4% fewer actions than first-time humans.

What the benchmark actually is. ARC-AGI-3 is a set of interactive visual games where the agent must discover rules and goals by poking around — no instruction manual. RHAE scores how efficiently the agent finishes levels versus first-time human players. Unfinished levels get zero; later levels weigh more.

Harness, not a bigger brain alone. VISTA feeds the model raw image observations, keeps a lossless visual memory of every frame (including animation frames), and lets the model zoom, compare past frames, and read pixel values as it reasons. The same underlying model, under the organizers’ official setup without that harness, scored far lower. GPT-5.6 Sol with VISTA reaches 99.00 RHAE; open-weight GLM-5.3 Flash reaches 66.93.

The story is the scaffolding. Ablations in the preprint show the big jump when lossless visual memory and model-directed inspection are added — not when the model alone is simply asked harder. The authors frame VISTA as a general-purpose visual harness and also report gains on other visual game and puzzle benchmarks with minimal adaptation.

Hard hedges that stay. Program-synthesis systems — Tycho and other PS approaches — already hit 100 on the public set by building executable world models in code. This paper is a preprint. And the authors themselves say training-data contamination cannot be ruled out until the private set of games is tested, because the models they used were released after the public benchmarks.

Ink, not pencil: a perfect public score without program synthesis is real in their reported runs — and still not proof of private-set generalization. Treat “first vision system to hit perfect without programs” as the authors’ claim on the public suite, checked against the caveats above — not as an unchecked “first ever.”

Why regular people should care

The Monday stake for builders and buyers: sometimes the win is how you wrap the model — memory, tools, and what it can look at again — not only how many parameters you buy. If agents that see the world need a visual notebook as much as a smarter core, product roadmaps shift toward harness design, not just model chase.

What's next

Watch private-set ARC-AGI-3 results, independent replications, and whether other labs adopt the same visual-memory pattern. The useful fact today is narrow: on the public 25-game set, a visual harness pushed Opus 5.0 to a reported perfect RHAE without program synthesis — with honest asterisks still attached.

← Back to AI