Wire
Scientific agent scaffolding cuts scores 6.7 points
Scientific coding agents scored 0.361 with full scaffolding versus 0.428 with a basic prompt, a 6.7-point drop across 40 spatial-transcriptomics alignment tasks and 120 runs per configuration. The August 10 RelSciFM workshop record says package hints increased tool exploration without improving results, while richer setup encouraged needless transformations, brittle package-first workflows, and infrastructure failures. For teams following evidence that the harness can move matched-model accuracy by double digits, the operator lesson cuts both ways: benchmark each scaffold addition against a plain baseline and inspect traces before calling more context an upgrade.