skip to content
The Weighted Average

Wire

TrajErrBench labels 486 failed agent runs

TrajErrBench opens 486 manually annotated failed trajectories from Tau2Bench and SWE-Bench Pro for studying where long-running agents first go decisively wrong. The TrajDebug preprint proposes tracing each error’s later resolution and terminal impact rather than labeling every local mistake as the root cause, though the abstract does not disclose a numeric performance margin over baselines. Builders extending session-level provenance inside agent harnesses can use the corpus to test whether observability finds the earliest consequential failure instead of merely explaining the last visible symptom.