skip to content
The Weighted Average

Wire

FinEvo-Bench gives Codex a 19.37-point gain

FinEvo-Bench found that retained experience improved four agent scaffolds by 9.33 to 19.37 points across 120 real-case-grounded tasks spanning six financial domains, with Codex posting the largest gain. The benchmark paper used one Qwen3.7-Max backbone across the scaffolds and reports that Letta reached the highest evolved score, 91.65, while rubric feedback beat reference-answer feedback on both quality and compliance; the results still depend on a Claude Code scoring agent rather than human adjudication. For operators extending AISI’s case for governed agent evaluation, the signal is to test improvement longitudinally: retained skills can compound, but compliance drift must be scored beside task quality.