skip to content
The Weighted Average

Wire

NVIDIA keeps private inference above 96%

NVIDIA says confidential DeepSeek-R1 inference on eight B200 GPUs retained 96.1%–98.2% of non-confidential output-token throughput. In NVIDIA’s benchmark, mean time-per-output-token overhead stayed between 1.2% and 4.3% across concurrency levels, though results are controlled and vendor-run. Teams weighing private inference should benchmark their own traffic before trading data exposure for speed; Rubin’s $40M-per-megawatt infrastructure math shows why utilization and trust boundaries now belong in the same capacity plan.