Wire
OctoLong trains on 6.2B dependency-rich tokens
OctoLong trained open models on 6.2 billion dependency-rich code tokens assembled with an AST parser, language server, and package manager, inside a roughly 50-billion-token context-extension mixture. The paper’s evaluation spans models from 600 million to 14 billion parameters and 18 open-weight baselines; replacing just 12% of conventional context-extension data improved long-range retrieval, state tracking, repository understanding, and downstream agent tasks. Builders should file the result beside GitHub’s stacked-PR response to cross-repository agent work: a larger context window is not enough if its training corpus lacks the dependency paths real software work follows.