skip to content
The Weighted Average

Wire

Meta fits a 30B agent model into 24GB

Meta released Muse Glimmer, a 30-billion-parameter open-weight agent model whose 4-bit build fits a 24GB memory envelope and reaches a reported 233.4 tokens per second on an RTX 5090 with its speculative-decoding drafter. The Apache 2.0 release says DFlash raises RTX 5090 decoding speed 3.1x, while the model card reports 51.2 on SWE-Bench Pro and support for text-plus-image input. For teams following the economics of open models that fit smaller hardware, this puts a measurable local-agent baseline on one consumer GPU—though Meta’s own safety results still argue for external guardrails before granting tools.