skip to content
The Weighted Average

Wire

Inference backends drive 39% of LLM variability

Inference backends accounted for roughly 39% of the out-of-the-box variability in a fully crossed study of three instruction-tuned models, five frameworks, six benchmarks, and four generation modes. The backend comparison found structural, model-dependent score changes even with greedy deterministic decoding, with larger divergences on factual than social-bias tests; framework defaults and sampling noise caused the remaining avoidable variation. Teams benchmarking local inference after llama.cpp’s DeepSeek V4 Metal speedup should pin the backend, version, and full generation configuration before attributing a score movement to model quality.