Wire
Q56 finds an eight-task local-model swing
Q56 found an eight-task spread across three identical GPT-OSS-20B runs—32 to 40 correct out of 56—while Gemma 4 E2B repeated 31 exactly, in a 36-run local coding benchmark with hidden tests. One 16GB GPU and a small custom corpus cannot establish a universal ranking, but the result adds a practical warning to benchmarking coding agents on representative work: repeat each configuration before treating a one-run score as a model decision.