Wire
Env-Rethink lifts agent tasks by 15.1%
Env-Rethink, a new 27B environment-preparation model, reports a 15.1%+ rubric pass-rate lift across nine models and 30 tasks, while its main comparison moves mean checks from 59.4% to 72.7%. The paper on evolving agent environments says Collection Maps, Event Logs, and event-driven noise expose context and provenance failures that a stronger downstream model cannot fix; on Terminal-Bench 2.1, evolution lowered success for at least three models on 32 of 55 retained tasks. For builders, same-model harness comparisons are only half the test: benchmark file provenance and environmental noise before blaming model capability.