skip to content
The Weighted Average

Wire

Manifest retires its router after 7,000 users

Manifest is shutting down its LLM router after four months across 7,000 cloud users produced mixed results. The gateway vendor’s postmortem says cache reads are 75%–90% cheaper than uncached input, while switching models adds behavioral variance and makes prompts, evaluations, and observability harder to maintain. For builders, that is a useful counterweight to the case for routing work across cheaper model tiers: benchmark a sticky model with prompt caching before paying for a classifier that may erase its own savings.