Wire
CalibForge lifts terminal agents 24.71 points
CalibForge generated 5,431 terminal tasks whose training gains reached 24.71 percentage points on Terminal-Bench 2.0, 27.68 points on SWE-bench Pro, and 30.04 points on Doc2Repo. The authors’ new paper says the system revises executable tasks against multiple solvers until they occupy a learnable difficulty band; models trained on the complete collection scored 32.58% and 47.57% on Terminal-Bench 2.0, and the task-generation code is public. Teams building the kind of containerized agent regression loop used by Supabase should file away the distinction: a task that runs is not necessarily a task that teaches, so synthetic-data pipelines need solver-relative difficulty checks as well as executable validation.