Wire
OpenMLE lifts an AI engineering score to 71.21%
OpenMLE’s 35B Frontis-MA1 system reached 71.21% Medal Average on MLE-Bench Lite, up from its base model’s 39.39%, using experience priors and asynchronous search under a 12-hour-per-task budget on one RTX 4090 capped at 12 GB VRAM. The authors released the model and full execution-grounded stack, and report that the trained model and search framework also transferred separately to held-out NatureBench Lite. For teams following the boundary between scaffold optimization and recursive self-improvement, this is a reproducible stack to test—but its long search budget belongs in any cost comparison with frontier agents.