Wire
llama.cpp lifts long-context DeepSeek speed 47%
llama.cpp raised DeepSeek V4 prompt throughput at 30K context from 33.40 to 49.18 tokens per second—a 47% gain—by adding a Metal implementation of the model’s Lightning Indexer. The b10236 release benchmarks also show throughput at 20K context rising from 45.83 to 62.01 tokens per second, while generation speed moved from 8.26 to 8.68. For operators evaluating the control premium of open-weight models, the patch is another reason to benchmark the current inference stack at the context lengths production agents actually carry rather than treating a model card’s speed as fixed.