MIT VISTA report broadens tests of visual AI memory
An October 1 manuscript extends VISTA testing beyond public ARC games. Its earlier perfect score still leaves private benchmark generalization unresolved.

MIT researchers have expanded their VISTA research with tests on GameWorld, AI GameStore and visual tracking tasks from BabyVision. The first arXiv manuscript was submitted on October 1.
The project already had a public introduction dated August 5, including its perfect score on 25 public ARC AGI 3 games. The broader evaluation is the important addition in the new report.
The MIT team comprises Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu and Kaiming He.
A visual memory that keeps the original evidence
VISTA is software surrounding a multimodal model. Its public repository describes a loop in which the agent observes an image, reasons about the environment and then chooses an action. Frames remain available after the action has finished.
The agent can revisit earlier frames, enlarge selected regions and inspect exact pixel colors. It also keeps durable notes about the game and temporary notes about its current work. Those tools make it possible to check an earlier observation instead of relying entirely on a compressed written recollection.
This is a practical distinction for an agent solving a visual puzzle. A small mark or the position of a piece can matter later, after the original scene is no longer on screen. Preserving the image gives the agent another chance to inspect that detail.
The report extends the evaluation
With GPT 5.6 Sol, the study evaluates 170 GameWorld tasks, ten AI GameStore games and 39 BabyVision questions. The BabyVision subset improves from 41.0 percent to an average of 63.2 percent with visual inspection. That covers two visual tracking categories rather than the full 388 question benchmark.
Another experiment uses GLM 5.3 Flash 320B with VISTA and reports 66.93 RHAE on public ARC games. These are the authors’ results.
The earlier code release is dated September 5. That repository implements the ARC environment. Its existing results and release date should stay separate from the new manuscript’s wider experiments.
A perfect public score has limits
The August introduction reports that VISTA with Claude Opus 5.0 completed every public game with a Relative Human Action Efficiency score of 100. Completion and action efficiency both matter to that metric. It is a result on the public set.
The authors also acknowledge that their models were released after the public games. They cannot rule out exposure during training and identify private games as the stronger test of generalization. That caveat remains important when interpreting the project’s headline result.
Protocol differences matter. For its ARC comparison, the paper treats official API results as external references to its subscription CLI experiments. Its internal ablations use the same access interface.
ARC Prize’s testing policy separately distinguishes public experimentation from private evaluation. It also says community submissions are not verified by default. A linked scorecard therefore should not automatically be described as an independently certified result.
What developers can take from this
The useful engineering question is whether retaining original visual evidence improves an agent’s performance on the tasks it will actually face. A production evaluation should test that choice alongside its processing cost, time limits and failure recovery. A strong game result alone does not establish reliability in an unrelated workflow.
Our coverage of GPT 6.1 Sol and ARC test setups explains why the surrounding software belongs beside the model name when comparing scores.
Featured image is an original AI generated editorial illustration of visual memory. It does not depict a real benchmark or product interface.



