Wire
Timeline-Bench gives agents 15 of 56 edits
Timeline-Bench gave AI agents only 15 successful video edits out of 56, with GPT-6 Astra in Codex and curated editorial guidance topping out at 26.8%. The benchmark paper tested 16 agents and found human editors preferred reference edits in 83.5% of 2,582 judgments; 562 of 771 failed runs passed delivery and brief checks but failed quality. Builders should treat video agents as compliance automators, not editors, until multimodal perception and quality review close that craft gap—an extension of Octobench’s measured harness swing.