Wire
SeGaBench finds useful gains in 93.3% of cases
SeGaBench’s strongest tested model delivered a validated performance improvement on 93.3% of 120 C and C++ cases whose profitable semantics were hidden from the compiler, according to the August 4 preprint. Its artifacts were correct in 94.8% of responses and cleared a 1.05x speedup in 83.3%, although many still captured only part of the oracle’s available gain. For engineering teams, the result extends the case for evaluating the whole coding-agent harness: let a model propose transformations a compiler cannot infer, but make executable correctness and performance checks—not model confidence—the merge gate.