New Georgia Tech Benchmark Measures Robot Memory – Top Models Fail Half the Hardest Tasks
A new benchmark from Georgia Tech measures whether vision-language models remember what they have seen in a home – and according to the reporting, the best models fail around half of the hardest tasks.
A home robot that has seen where the bedding is kept, when the children usually tidy up, and which rooms change over the course of a day needs to be able to retrieve that experience again. This is precisely the visual memory that a new benchmark from Georgia Tech puts to the test – with weak results for the best-known models, according to secondary reporting on the project.
What FINDINGDORY tests
The benchmark, called FINDINGDORY, consists of memory evaluations with 60 tasks that require continuous engagement and awareness of context and surroundings, according to reporting on the research group of Zsolt Kira, associate professor at Georgia Tech's School of Interactive Computing. Three task categories are marked as especially difficult: spatial (where things were), temporal (when things happened), and multi-goal tasks that combine several requirements. This is exactly the kind of knowledge a robot needs in order to navigate and work in a home over time.
The reporting states that the paper was presented in early September at the European Conference on Computer Vision in Sweden, and that Yadav and Ali are co-first authors. Note that the reporting refers to the conference as the "2026" edition while citing an arXiv version with a 2025 DOI (10.48550/arxiv.2506.15635); this chronological discrepancy cannot be resolved from the available sources. The underlying material is also a single secondary article – the score levels and model rankings are therefore reported figures, not independently verified.
The results are humbling for the industry: the best model scored only around 50 percent on the hardest tasks. And surprisingly, it was not one of the Gemini or GPT models at the top, according to the reporting. It was an open-source reasoning model from Qwen that won the benchmark. Karmesh Yadav, a PhD student at Georgia Tech and one of two co-first authors, is behind the study along with Kira's group. The article mentions that both Gemini 2.0-flash and GPT-4o were tested, but the reporting does not give the exact model name or size of the winning Qwen model.
Why visual memory is the bottleneck
Kira points to a fundamental limitation: most active vision-language models (VLMs), he says, can only process a few hundred images at a time before forgetting them. A robot in a home generates continuous visual input – hours of observations – that far exceeds what the models can practically retain in context.
That is why the results matter beyond the benchmark itself: the problem is not primarily reasoning, but memory. A model can be good at interpreting a single image and still fail a task because it no longer has access to what it saw earlier.
Kira suggests two possible ways around the memory dilemma, as reported:
- Converting images to text – descriptive text that language models easily understand can bypass the storage problem.
- Hierarchical pruning – teaching the model to identify irrelevant or redundant images and delete them from storage.
Neither approach is demonstrated in the reporting as fully fledged solutions – they are Kira's own suggestions for research directions.
Why the measurements are coming now
Interest in robot memory is not purely academic. Humanoid robots are already in homes, and the company Figure reported itself that Helix 2.5 completed 237 of 420 attempts (56 percent) on tasks such as bedding, towel folding and toy tidying across 30 homes the robots had not seen before, according to TechRepublic's coverage. The figures are the company's own and are not independently verified, and they say nothing directly about the benchmark – but they illustrate the gap between today's reliability and what is needed for everyday use.
1X Technologies, led by Norwegian CEO Bernt Børnich, also points to memory as a priority research direction for home robots.
"I want the robot I have at home to remember me, my kids and my family," Børnich told Newsweek.
The benchmark thus arrives just as the industry needs a common measure for a capability that has so far lacked objective testing.
The researchers' purpose
Kira was clear about why the benchmark was created, according to the reporting:
"There was no benchmark that could test these specialized capabilities we might want from a robot," Kira said. "Our purpose in creating one is to stimulate research in this area, so that others can develop methods to solve the problem."
In other words: the goal is not to crown a winner, but to make visible a specific, measurable gap that researchers can work toward.
Open questions
The reporting leaves several important uncertainties. It does not state how well results on the benchmark translate to real robots' performance in actual homes – how task scores relate to physical robotics in practice is not documented in the sources. It has also not been demonstrated that Kira's proposed solutions (text conversion and hierarchical pruning) scale to the continuous data stream a robot actually produces.
And it remains to be seen whether the Qwen model's lead stems from properties of the model itself – such as reasoning steps – or from an architecture that fits the benchmark's constraints better. Either way, FINDINGDORY gives the research community, if the numbers hold, something the reporting describes as new: a common, concrete measure of a capability that home robots in practice depend on.

