Wire
TimeEvo posts +8.79 points across 10 tasks
Visa Research’s TimeEvo reports an 8.79-point accuracy gain on GPT-5.6-Luna across 10 time-series QA tasks, with positive gains on all 30 task/backbone combinations. The paper turns diagnosed failures into evidence-only tools and admits them only after a paired validation gate; its baseline study also found generic self-revision changed 147 answers and broke 56 that were previously correct. Builders adding tools to analytical agents should measure fixed-versus-broken cases on held-out data before shipping self-evolving behavior, extending the benchmark-cost case for reproducible agent evaluation.