skip to content
The Weighted Average

Wire

llama.cpp lifts DeepSeek V4 decoding 1.83x

llama.cpp added DeepSeek V4 DSpark speculative decoding in release b10228, with the contributor’s nine-prompt DGX Spark test falling from 102.15 seconds to 55.95 seconds—a 1.83x speedup. The merged implementation and benchmark accepted 1,038 of 2,236 drafted tokens, while individual throughput ranged from 19.8 to 39.3 tokens per second. For operators weighing the economics of Chinese open-weight models, the result is a useful reminder that inference engines and draft heads can move realized latency almost as much as the headline model choice.