skip to content
The Weighted Average

Wire

Logic pretraining saves 36B language tokens

Logic pre-pretraining reached 80% linguistic-task accuracy with 36 billion fewer tokens than standard initialization in a 100-billion-token study, according to the August 4 preprint. Models first trained on formal derivations also matched the dense baseline at roughly 33% pruning sparsity, which the authors connect to a lower-rank internal representation. Teams weighing the economics behind model-compression systems should test data ordering before buying a larger run: synthetic logic may improve both acquisition speed and the model’s later compressibility.