skip to content
The Weighted Average

Agentic Engineering

Cloudflare AI Search Halves the Semantic Free Tier

AI Search goes GA with November billing and half the semantic-only free allowance. The larger decision is what its managed retrieval bill excludes.

Old books arranged on irregular wooden bookshelves
Old books arranged on irregular wooden bookshelves. Photograph by Paul Melki

Reprice your retrieval pilot before November: Cloudflare made AI Search generally available on October 1, with billing starting November 1. Compared with its preview proposal, the included allowance for a semantic-only workload falls 50%, although the underlying query rate does not rise.

That sounds more consequential than the immediate dollar difference. In Cloudflare’s own published workload example, the changed allowance adds just $0.75 to the monthly semantic-query line. The larger purchasing question is whether a managed retrieval pipeline removes enough integration work without obscuring the costs, permissions, and quality checks that remain outside it.

The free pool splits, not the unit price

The arithmetic joins two versions of the offer. Cloudflare’s preview pricing announcement supplied a shared pool of 2,000 queries across semantic and full-text search. Its GA announcement replaces that pool with 1,000 semantic queries and 1,000 full-text queries. A semantic-only workload therefore loses (2,000 − 1,000) ÷ 2,000 = 50% of its included queries. A balanced workload can still use both allowances. There is no blanket halving of every customer’s free usage.

Keep the comparison’s status intact. The old numbers were a preview pricing proposal while billing was disabled, not a previous paid invoice. Cloudflare is establishing the commercial terms of a service that had been free during beta. Describing this simply as a doubling of prices would confuse the allowance, the unchanged usage rate, and the start of billing itself.

The current pricing documentation lists $0.75 per 1,000 semantic, vector, or hybrid queries, and $0.10 per 1,000 full-text queries beyond the included quantities. Cloudflare’s preview example used 30,000 semantic queries. Under its old proposal, (30,000 − 2,000) ÷ 1,000 × $0.75 = $21; under GA, (30,000 − 1,000) ÷ 1,000 × $0.75 = $21.75. The derived increase is $0.75, or approximately 3.6% of that query line, not of the whole application bill.

This is useful precisely because it resists the dramatic headline. For that published workload, engineering time spent rebuilding retrieval solely to recover the allowance would need an unusually strong justification. Smaller semantic-only projects cross the charging threshold earlier, while mixed workloads need separate counters. The practical response is to update the forecast and meter the actual mix, not assume that every search consumes an interchangeable free credit.

Cloudflare’s broader offer bundles parsing, chunking, selected Workers AI embeddings, keyword indexing, and reranking into ingestion, storage, and query charges. Its GA post explicitly avoids instance-hour charges and monthly minimums. That can make a small retrieval service easier to reason about than a collection of separately provisioned components. It does not establish a saving over a particular existing stack; the announcement supplies a rate card, not a matched cost study.

That distinction extends yesterday’s SkillSeek analysis of retrieval cost and matched accuracy. A simpler or cheaper retrieval stage earns a trial. It earns production only when the whole agent still finds the right evidence, recognizes missing evidence, and completes work at an acceptable cost.

The bundle stops before the answer

Read the exclusions before the allowances. Cloudflare’s billing rules include Workers AI embedding and reranking calls, but leave generation, query rewriting, and external model providers on the customer’s applicable account and gateway. Those included embedding and reranking calls also stop appearing on the Workers AI bill and in AI Gateway logs. A disappearing line item is a metering change, not proof that the work stopped occurring.

The included monthly quantities are 5 million ingestion tokens and 10 GB-month of storage, alongside the split query pools. Beyond them, base ingestion costs $0.75 per million tokens, image processing adds $0.50 per million, and stored data costs $2 per GB-month. Treat each as a different meter. A stable corpus, a frequently replaced corpus, and a large scanned-document archive can have the same query traffic and different invoices.

Token accounting adds another boundary. The pricing page counts the final indexed chunks using a common tokenizer, regardless of the embedding model. Overlapping text is counted in each chunk where it appears. The chunking documentation allows overlap from zero to 30% and explains that chunk size and retrieval count also affect the context sent to generation. Raw document size is therefore not a complete ingestion or answer-cost forecast.

Do not flatten image support into a universal upgrade. The supported-model matrix distinguishes models with native image support from text-only models. The GA explanation says text-only configurations can turn an image query into a caption before searching, while multimodal models embed the image directly. These are different retrieval representations. A screenshot containing a small but decisive interface state may demand a different test from a page of ordinary prose.

File handling is similarly conditional. The data-source rules permit 10 MiB plain-text or code files and PDFs with OCR enabled. PDFs without OCR and other converted formats remain limited to 4 MiB; oversized files are not indexed and appear in error logs. OCR is disabled by default, and changing the setting triggers a full reindex. A larger upload ceiling is not evidence that every document in the migration made it into the searchable corpus.

The first buyers who should consider switching are teams already spending effort operating a modest search pipeline over supported documents, particularly when selected Workers AI models fit their requirements. Teams with specialized ranking, access controls, or ingestion formats should first establish that the managed service preserves those requirements. The benefit is fewer components to operate, not permission to stop understanding the pipeline.

Today’s Clef brief separates a cheap decision from a correct routing outcome, while Microsoft’s streaming launch separates faster speech from batch feature parity. AI Search belongs to the same purchasing pattern: pay for the capability you actually need, and inspect the boundary where the attractive component claim ends.

Hybrid by default still has a ceiling

The most important limit may not be financial. The hybrid-search documentation permits 500,000 files per paid instance with keyword search, versus 1 million for vector-only instances. New instances use hybrid search by default. A buyer importing a large collection cannot treat the larger vector-only ceiling as the capacity of the default configuration.

Hybrid search runs keyword and vector retrieval in parallel, then merges the results. Cloudflare defaults to reciprocal rank fusion and offers an alternative based on the higher normalized score. Reranking is available but disabled by default. Included pricing and enabled behavior are different facts: an operator who wants the reranking stage still needs to configure and evaluate it.

The strongest case against migration is therefore not the smaller free allowance. It is a mismatch between the service’s boundaries and the workload. A corpus can exceed an instance ceiling, contain unsupported or oversized documents, or require ranking behavior that a default configuration does not reproduce. Splitting a collection may be workable, but then cross-instance retrieval and evaluation become part of the architecture rather than an invisible implementation detail.

Website ingestion has its own clock. The website-source documentation and pricing limits distinguish crawl size from daily crawl throughput; the free plan permits 500 pages crawled per day. An index that is inexpensive to query can still be too stale for a particular application. Before adopting it for changing documentation, measure the time from a source update to a retrievable corrected passage.

Access deserves equal attention. The earlier product announcement describes public search and MCP endpoints that can search a namespace without authentication, with custom domains and Cloudflare Access available for private use. Convenience is not a permission model. Before enabling an endpoint, decide whether the indexed material is public, whether different users may see different records, and which layer enforces that distinction. Do not infer per-document authorization from the existence of a login screen.

There is also a small but concrete cleanup trap. Cloudflare’s pricing page says website-crawled pages now live in built-in storage; the dedicated R2 bucket created by the earlier arrangement is no longer used. Objects left behind can continue to incur R2 storage charges. Inventory the old bucket and its dependencies before removing anything. A managed migration can simplify the new bill while leaving an old one running.

Our Atlas Infinite analysis separated a familiar database interface from a changed operating contract. Here the contract changes around ingestion, ranking defaults, logging, and billing. Evidence that would break the migration thesis includes lower retrieval acceptance, stale indexing, unmet access requirements, or generation and repair costs that overwhelm the managed-service saving. None is answered by the word generally available.

Spend October measuring the complete path

Use the period before November billing to build a forecast from observed work. Cloudflare has specified the November 1 start date; that is a reason to make the meters legible now, not a reason to rush an irreversible production move. Record the corpus configuration and retrieval mode alongside every test result so a change in defaults cannot masquerade as a quality improvement.

Start with an inventory of supported files, rejected files, and indexed chunks. Keep original source identifiers so a missing result can be traced back to ingestion rather than blamed automatically on the embedding model. Test explicit identifiers and paraphrased questions separately: those query types exercise the keyword and semantic paths differently. Include questions whose answer is absent, because a plausible answer without evidence is a failed retrieval application, however cheap its search call.

Next, retain a stable comparison set against the existing system. Record whether the necessary passage appears, whether the downstream answer uses it correctly, and whether a person must intervene. Compare complete request time rather than just retrieval latency. Keep generation and rewriting usage in the budget, because the managed query price does not include them. Where external models are required, use their actual invoices rather than attributing their cost to an included Workers AI stage.

The commercial decision should follow that evidence. A team can rationally pay slightly more per query to retire integration work, or keep an existing system because its quality and controls already fit. Neither decision needs an invented annual labor saving. Name the maintenance tasks that would disappear, identify who currently performs them, and verify that the new service actually removes rather than relocates that responsibility.

The operating checklist is narrow:

  • Small retrieval teams should pilot when supported data and selected models match their needs. Use shadow traffic or a reversible document collection; switch only after retrieval acceptance and freshness hold.
  • Finance and platform owners should reforecast with separate semantic and full-text allowances, indexed-token ingestion, storage, and excluded model charges. The published example’s extra $0.75 is a query-line change, not a total-cost estimate.
  • Application owners should verify defaults and permissions before enabling hybrid retrieval, OCR, reranking, or a public endpoint. Each setting changes either the data path, the work performed, or who can retrieve it.
  • Procurement should watch the first billed month for unexplained reindexing, unclaimed allowances, lingering storage, and answer-generation spending. Keep the old path available until the observed operating cost and acceptance rate justify retirement.

Cloudflare has made the retrieval component easier to buy. The defensible opportunity is to replace maintenance burden with a service whose limits are explicit. A simpler retrieval bill is not the same thing as a cheaper accepted answer; October is the time to find out whether this one becomes both.

Sources