Wire
Robot-memory benchmark tops out near 50%
Georgia Tech researchers introduced the FindingDory benchmark, 60 memory tasks for vision-language models in assistive robots; the best system managed roughly 50% on high-difficulty tasks. The Georgia Tech benchmark report describes spatial, temporal and multi-goal tests over an average day, and says a smaller open-source reasoning model matched closed models such as Gemini 2.0 Flash and GPT-4o. Robot builders should make long-horizon memory its own acceptance gate—not assume a larger context window solves household state—alongside OSWorld’s evidence that agent scores can go stale.