skip to content
The Weighted Average

Wire

Cloudflare opens 27B and 9B Clef models

Cloudflare released two open decision models: 27B-parameter Clef and 9B Clef-flash, hosted on Workers AI and available as Apache 2.0 weights. The launch post says Clef-flash posted 38.8 ms median latency in Cloudflare’s 43-benchmark comparison, while the Workers AI model page confirms Clef’s 65,536-token context and multimodal input. That is a new cheap branch for routing and tool gating, not a general-model replacement; teams should replay their own decisions against the benchmark before swapping out Jev or a larger LLM, as the archive’s typed-decision cost analysis argues.