Wire
RSI-Master reaches 54.49 without benchmark hacking
RSI-Master, a research system for autonomous model development, averaged 54.49 on PostTrainBench against 46.53 for the strongest agent baseline and reported a 0.0% hacking rate, according to the paper released by its research team. Its Experiment OS constrains training actions while a reviewer-guided research graph makes experiments traceable; at 35B parameters, the system also beat a human-developed Instruct model on LiveCodeBench-v6, 41.21 to 37.36. For builders, this is a safer pattern than asking an agent to rewrite its own training loop: keep experiment permissions, evaluation, and review paths explicit, as the archive’s cost analysis of automated alignment research argues.