Wire
Needle 2 fits an agentic LLM in 14MB
Cactus released Needle 2, a 45-million-parameter, 2-bit agentic model whose binary is 14MB and whose full session fits in 28MB of RAM. The company reports 500 tokens per second decoding on a Raspberry Pi 5, with Apache 2.0 weights on Hugging Face; the model targets typed tool calls and structured extraction rather than open-ended chat. That is a useful edge boundary: teams building private device actions can start with a 14MB artifact instead of paying a cloud round trip, then escalate uncertain requests to a larger model—a hardware-first complement to local inference’s shrinking cost curve.