Models & Open Source
A 2B Model Pretrained on Gamer GPUs for $6,891
Tsinghua's Puro-2B trained a 2B model from scratch on RTX 5090s for $6,891 of accelerator time — about $4.93 per billion tokens.
A Tsinghua-led group pretrained a 2-billion-parameter language model from random initialization on consumer RTX 5090 gaming GPUs for a measured $6,891 of accelerator time. The canonical run consumed 22,514 active GPU-hours across 1.3988 trillion tokens, which works out to a derived figure the paper does not state: $4.93 of accelerator cost per billion tokens of pretraining. The Puro-2B technical report frames the comparison bluntly — training Llama-3.2-3B costs over $1.5 million and reproducing SmolLM3-3B over $700,000, making this run roughly 217 times cheaper than the former.
The quality claim is calibrated rather than triumphant. On the report’s fixed 15-benchmark base-model aggregate, the Puro-2B model card reports 57.81 overall against 55.14 for Qwen2-1.5B and 60.73 for Qwen2.5-1.5B — beating a 2024-class baseline, trailing its successor. A cheaper checkpoint at about $4.4K already clears Qwen2-1.5B, and the fitted Puro Cost Scaling Law suggests roughly $4,400 suffices for that bar. The team released weights under Apache 2.0, the materialized pretraining dataset, the training implementation, and the data-processing code.
What $4.93 a billion tokens actually buys
The recipe is where the cost went. Phase 1 ran 439 billion tokens on 24 RTX 5090s at a median 238 TFLOP/s per GPU; Phase 2 ran 960 billion tokens on 96 GPUs at 192. Main transformer linear-layer GEMMs use blockwise E4M3 FP8, with master weights and numerically sensitive operations held in higher precision. Selected matrix weights update through MuonH with hyperball projection at a 10x learning-rate multiplier while the rest use AdamW. The final model is an equal-weight average of six checkpoints from a constant-learning-rate branch resumed at step 218,000. None of that is exotic infrastructure — it is a curriculum, an optimizer choice, and a precision decision applied to hardware anyone can buy at retail.
Read the arithmetic the other way and the hardware choice explains itself. Twenty-two thousand GPU-hours spread across 96 consumer cards is roughly ten days of wall-clock time on a cluster that costs less to assemble than a single rack of data-center accelerators, and the FP8 path is what makes those cards keep up: blockwise E4M3 on the heaviest matrix operations roughly halves the memory traffic that would otherwise strand a 32GB consumer card. The two-phase token split — a smaller, faster Phase 1 followed by a wider Phase 2 — is the same shape frontier labs use, executed at a scale where a failed run costs a few hundred dollars instead of a few million. That is the part worth borrowing even if you never train a base model: the experiment loop, not the endpoint.
The caveat is stated by the authors and deserves equal weight in any budget built on this number. The figures are accelerator-only reproduction estimates: they exclude data acquisition and preprocessing, proxy and ablation experiments, failed runs, post-training, evaluation, storage, networking, and research labor. The model card says so explicitly. Anyone reading $6,891 as the cost of getting a 2B model into production is reading a GPU invoice as a project plan. The report also notes the scaling law fits five single-run points with no uncertainty interval and does not isolate curriculum ordering, constant-LR continuation, and checkpoint averaging as independent causal gains.
Who should care, and who should not
The wrong conclusion is that pretraining is now cheap. The right one is narrower and more useful: the floor for inspectable, from-scratch pretraining at the 2B scale has dropped to a number a university lab or a single engineering team can expense, and the full pipeline — not just weights — is public. That distinction matters because most open-weight releases give you the endpoint without the process, which makes controlled studies of data curricula impossible. Puro-2B’s second contribution is exactly such a study, examining how pretraining data ordering shapes downstream performance after post-training.
There is also a supply-chain argument buried in the hardware list. A recipe that runs on retail gaming cards is a recipe that runs on hardware no export-control regime tracks the way it tracks data-center accelerators, and that is a structural fact for research groups outside the handful of countries with unconstrained access to frontier silicon. The report does not make that argument; the bill of materials makes it implicitly.
For operators, the practical question is when from-scratch pretraining beats post-training an existing open model, and for almost everyone the answer remains post-training. A $6,891 accelerator run yields a base model that trails Qwen2.5-1.5B; the same money spent adapting a stronger open checkpoint to a specific domain generally buys more capability per dollar, which is the pattern behind Harvey’s roughly $1.1M post-training run on Kimi K3 and behind the broader finding that $26.5B of investment chases only 6% of businesses actually running open weights. Pretraining earns its cost when you need architectural control, data provenance you can audit end to end, or a licensing position no vendor will grant — the same control premium that shows up whenever small open weights are priced against hosted alternatives.
What would change the verdict: a comparable open recipe demonstrating the same cost curve at 7B or larger, where consumer-GPU memory and interconnect limits bite hardest, or an independent reproduction of the $6,891 figure on rented 5090 capacity at a stated hourly rate. Until then, treat $4.93 per billion tokens as a credible floor for a specific scale on specific hardware, not a general price for intelligence. The comparison worth holding in mind is today’s lead on ChatGPT’s dollar-a-user advertising yield: both numbers are real, and both are early points on curves whose shape nobody has yet proven.