Wire
GradCuit lifts reasoning accuracy 6.6 points
GradCuit averaged 64.5% reasoning accuracy, beating chain-of-thought prompting by 6.6 percentage points and the strongest comparison method by 2.4 points. The August 3 preprint optimized inserted latent states while keeping model parameters frozen, testing five instruction-tuned backbones, three benchmarks, and two answer formats; across seven learning rates, accuracy variation fell from 1.53 to 0.82. For teams already seeing agent harness choices swing coding outcomes, latent-state optimization adds another inference-time lever worth measuring—but only after its extra compute cost is reported against sampling and reranking baselines.