Wire
Gimlet and Cerebras target 3,000 tokens per second
Gimlet Labs and Cerebras plan 100 megawatts of Cerebras-powered inference capacity, targeting up to 3,000 tokens per second when the first Gimlet Cloud data center comes online later this year, according to Gimlet’s announcement. The multisilicon design combines Cerebras wafer-scale systems with GPUs; both capacity and speed are forward-looking targets, not a public production benchmark. Builders of real-time agents should test end-to-end latency and cost against the archive’s inference-price comparison before treating token rate as a product moat.