Wire
PAST-Bench tests agent memory over 204 episodes
PAST-Bench tests whether retained experience improves personal agents across 26 scenarios and 204 episodes, spanning seven base models and four frameworks, according to the August 4 preprint. The study finds real but uneven gains and distinguishes a better later answer from evidence that the agent actually saved, retrieved, and updated the relevant experience; its Hermes+ reference adds five interventions across that loop. For teams building persistent assistants, this turns the outcomes-versus-goals divide in agent design into a testable rule: audit the memory pathway, not just the final score, before calling a system self-improving.