# Cloudflare AI Search
Cloudflare AI Search is [[Cloudflare]]'s managed search and [[Retrieval-Augmented Generation (RAG)|RAG]] service. You create an instance, give it data (uploaded files, an [[Cloudflare R2|R2]] bucket or a website you own), and query it in natural language from a Worker, the REST API, an MCP endpoint or a drop-in web component. Parsing, chunking, embedding, keyword indexing, reranking and syncing all happen on Cloudflare's side.
It started life as **AutoRAG** (open beta, April 7, 2025) and was renamed AI Search on September 25, 2025. It went generally available on October 1, 2026, and usage billing starts on November 1, 2026.
Why care? Because a "simple" RAG pipeline is never simple. You need storage, a parser, a chunker, an embedding model, a [[Vector Store|vector database]], a keyword index if exact terms matter, fusion logic, a reranker, a sync job that notices changed and deleted files, and something to stop it all from drifting. AI Search packs that into one product with one bill. The price is control: you get Cloudflare's defaults, Cloudflare's models and Cloudflare's data sources.
## From AutoRAG to AI Search
- **April 7, 2025**: AutoRAG open beta. It wired [[Cloudflare R2|R2]], [[Cloudflare Workers AI|Workers AI]], [[Cloudflare Vectorize|Vectorize]] and [[Cloudflare AI Gateway|AI Gateway]] together for you, but those resources lived in your account and were billed separately. R2 was the only data source
- **2025**: metadata filtering by folder and timestamp (April), streaming, file-name filters, 3-5x faster indexing (July), deleted files finally removed from the index on sync (July), max results raised from 20 to 50 (August), website data source plus [[#NLWeb|NLWeb]] (August)
- **September 25, 2025**: "AutoRAG is now AI Search", with third-party models (OpenAI, Anthropic and others) through AI Gateway. Cloudflare framed the new name as a bigger mission than RAG: search infrastructure for every developer
- **October 2025**: reranking and per-request system prompts. **November 2025**: custom HTTP headers so the crawler can reach pages behind authentication
- **March 23, 2026**: new OpenAI-compatible REST endpoints (`/search`, `/chat/completions`), public endpoints, UI snippets, the MCP endpoint and custom metadata filtering. The old AutoRAG endpoints keep working, but new features only land on the new API
- **April 16, 2026**: the big rework. [[#Search modes|Hybrid search]] (BM25 next to vectors), relevance boosting, cross-instance search, built-in storage and vector index, and new `ai_search` / `ai_search_namespaces` Workers bindings replacing `env.AI.autorag()`
- **June 2026**: every pre-April instance migrated to managed infrastructure (completed June 18). [[Wrangler]] got namespace and sync-job commands
- **August 6, 2026** (Agents Week): custom domains, [[Cloudflare Access and Zero Trust|Cloudflare Access]] in front of public endpoints, namespace-wide public endpoints, the `discover` crawl mode for sites without a complete sitemap, and a preview of the pricing model
- **October 1, 2026**: GA. Hybrid search on by default, multimodal embeddings (image queries included), OCR for scanned PDFs on every account, 10 MiB limit for text files and OCR'd PDFs
The crawler user agent followed the rename in February 2026 (`Cloudflare-AutoRAG` became `Cloudflare-AI-Search`; the old token still works in `robots.txt`).
## How it works
**Indexing** is asynchronous:
1. Ingest from built-in storage (indexed as soon as a file is uploaded), an R2 bucket or a website (both on a sync schedule)
2. Convert everything to Markdown with Workers AI's `toMarkdown()` (PDF, Office, OpenDocument, HTML, CSV, images, plus a long list of plain-text and code formats)
3. Chunk it with recursive chunking: split at paragraphs, then sentences, then by token count. You control chunk size and overlap (0 to 30%)
4. Embed each chunk. Multimodal models embed image pixels directly; text-only models embed a generated caption
5. If keyword search is on, add the chunk to a BM25 inverted index
6. Store vectors (in Vectorize, under the hood), the keyword index and the content
**Querying** is synchronous:
1. Optional query rewriting by an LLM (only with the `messages` format; it resolves follow-ups like "how do I deploy one?" using the conversation history)
2. Embed the query
3. Run vector search and, in hybrid mode, BM25 in parallel
4. Fuse both lists with reciprocal rank fusion (`rrf`, the default) or `max`
5. Optionally [[Reranking|rerank]] with a cross-encoder (`@cf/baai/bge-reranker-base`, off by default)
6. Return chunks (`/search`), or pass them to a generation model (`/chat/completions`)
There's also a similarity cache on generated answers. It matches near-duplicate prompts with MinHash and locality-sensitive hashing, at four strictness levels with playful names (`super_strict_match` up to `anything_goes`). The cache used to keep answers for a fixed 30 days; since June 2026 the default is 48 hours, configurable from 10 minutes to 6 days, with an on-demand purge.
## Search modes
- **Vector**: finds meaning even when the words differ, but loses exact strings. Cloudflare's own example is an error code like `ERR_CONNECTION_REFUSED`
- **Keyword (BM25)**: matches the exact terms. Tokenizer `porter` (stemming, for prose) or `trigram` (substrings, for code), match mode `and` or `or`
- **Hybrid**: both, fused. The default for new instances since GA. It halves the file cap on Workers Paid (500,000 files instead of 1,000,000)
On top of that: metadata filters with Vectorize-style operators (`$eq`, `$in`, `$gte`...), up to 5 custom metadata fields per instance, relevance boosting on up to 3 fields (e.g. newest first by `timestamp`), and a score threshold. Each chunk can return a `scoring_details` object with vector score, BM25 score, both ranks and the rerank score, which is really useful when you need to debug why something ranked where it did.
## Data sources
- **Built-in storage**: upload through the Items API or the dashboard. Backed by R2 and Vectorize, but you never see those resources
- **R2 bucket**: synced on a schedule (every 6 hours by default, configurable from 15 minutes to 24 hours) or on demand via API or `wrangler ai-search jobs create`, at most every 30 seconds. Custom metadata comes from R2 object headers
- **Website**: crawled with Browser Run (formerly [[Cloudflare Browser Rendering|Browser Rendering]]), either from sitemaps or with the `discover` mode that follows links. Static or rendered (headless browser) mode, CSS content selectors to skip navigation and footers, path filters, up to 5 authentication headers, metadata pulled from `<meta>` tags. The site must be a zone on the same Cloudflare account
- An instance can combine built-in storage with one external source
Instances that receive no search request for 31 days get their scheduled syncs paused automatically (they stay searchable).
## Models
- **Embedding**: Workers AI (`qwen3-embedding-0.6b` is the default, plus BGE M3, BGE large, EmbeddingGemma and the multimodal `qwen3-vl-embedding-2b`) or, through AI Gateway, OpenAI `text-embedding-3-small/large` and Google `gemini-embedding-001` / `gemini-embedding-2`. The embedding model is fixed at creation; changing it means a new instance
- **Generation and query rewriting**: Workers AI models (Llama 3.x and 4 Scout, gpt-oss, DeepSeek V4, Qwen 3.8, Kimi K2.7 Code, GLM 4.7 Flash / 5.3...) or any model reachable through AI Gateway's chat completions endpoint. You can override the generation model per request
- **Smart Default** lets Cloudflare pick and upgrade models for you over time
- The GA blog post says multimodal embeddings use [[Matryoshka Embeddings|Matryoshka Representation Learning]] to keep vectors small
## Ways to use it
- **Workers bindings**: `ai_search` binds one instance (`env.MY_SEARCH.search(...)`); `ai_search_namespaces` binds a namespace, so a Worker can create, list and delete instances at runtime (one per tenant, per customer or per agent) and search several instances in one call. The legacy `env.AI.autorag("name")` binding still works, with no planned removal
- **REST API**: `/accounts/{account_id}/ai-search/namespaces/{namespace}/instances/{id}/search` and `.../chat/completions`, in the OpenAI `messages` format, plus Instances, Items and Jobs APIs
- **Public endpoints**: unauthenticated `/search`, `/chat/completions` and `/mcp` on `<id>.search.ai.cloudflare.com` (or `ns-<id>...` for a namespace), with rate limiting, CORS, custom domains and optional Cloudflare Access
- **[[Model Context Protocol (MCP)|MCP]]**: every instance can expose a `search` tool to any MCP client. Cloudflare's own "Dev Stack MCP" (cited docs for Cloudflare, Astro, Vite, Hono and more, one instance per site) runs on it
- **UI snippets**: web components for a search bar, search modal and chat bubble
- **CLI and SDKs**: `wrangler ai-search` (create, search, stats, namespaces, jobs, with `--json`), a Python SDK, an `ai-search-provider` package for the [[Vercel AI SDK]], a retriever in `langchain-cloudflare` for [[LangChain]], and [[Cloudflare Agents SDK]] guides
### NLWeb
NLWeb is an open protocol started by [[Microsoft]] for natural-language queries on websites: an `/ask` endpoint for people and an `/mcp` endpoint for agents. AI Search can deploy an NLWeb Worker on top of a website instance in a few clicks. Cloudflare still labels it a public preview.
Cloudflare uses AI Search for search on its blog, its developer docs and cloudflare.com, which is a decent signal that it holds up on large documentation sites.
## Limits (October 2026)
| Limit | Workers Free | Workers Paid |
| --- | --- | --- |
| Instances per account | 100 | 5,000 |
| Files per instance | 100,000 | 1M (500K with hybrid) |
| Pages crawled per day | 500 | Unlimited |
| Pages per `discover` crawl | 100,000 | 100,000 |
| Max file size | 10 MiB for text, code and PDFs with OCR; 4 MiB for other formats | same |
| Custom metadata fields | 5 per instance | 5 per instance |
| Instances per cross-instance query | 10 | 10 |
Metadata is capped at 10 KiB per vector, and only the first 64 bytes of each string are filterable.
## Pricing
Billing starts November 1, 2026. Every account, Free or Paid, gets 5M ingestion tokens, 10 GB-month of storage, 1,000 semantic queries and 1,000 full-text queries per month. Beyond that:
- Ingestion: $0.75 per million tokens, plus $0.50 per million for images and OCR
- Storage: $2.00 per GB-month
- Semantic, vector and hybrid queries: $0.75 per 1,000
- Full-text queries: $0.10 per 1,000
Embedding and reranking with Workers AI models, storage, vector indexing and Browser Run crawling are included. Generation, query rewriting and third-party models are not: they're billed as Workers AI or AI Gateway usage. Ingestion tokens are counted on the final chunks with the `cl100k_base` tokenizer, so chunk overlap is billed every time it repeats. Cloudflare's own example (20,000 documents, 1,000 images, 30,000 queries a month) lands around $35 for the first month and around $21 after that, almost all of it queries.
## Compared with other managed RAG services
- **OpenAI file search / vector stores**: upload files to a vector store, get automatic chunking and embedding, hybrid search (RRF between embeddings and keywords) with a ranker. $0.10 per GB per day of storage after 1 GB free (roughly $3 per GB-month) and $2.50 per 1,000 tool calls in the Responses API, plus model tokens. Tightly bound to [[OpenAI]] models, no crawler
- **Google [[Vertex AI]] Search** (billed as "Agent Search" on Google Cloud's pricing page): $1.50 per 1,000 queries (Standard) or $4.00 (Enterprise, with website search and generated answers), index storage around $5 per GiB-month, 10,000 free trial queries a month. More enterprise ranking and structured-data features, much higher per-query price
- **Amazon Bedrock Knowledge Bases**: the classic self-managed flavor makes you bring and pay for your own vector store. The Managed Knowledge Base (June 2026) is the closest equivalent: $5 per GB of raw data per month, $1 per 1,000 hybrid Retrieve calls, free parsing, embedding and reranking with the managed models, and connectors for SharePoint, Confluence, Google Drive, OneDrive, S3 and the web
On paper, AI Search is the cheapest per query of the hosted options, and the only one where a public MCP endpoint and website search are a toggle. AWS and Google win on enterprise connectors: AI Search only reads R2, your uploads and sites you host on Cloudflare. The storage prices aren't measured the same way (index size vs raw data), so compare on your own corpus rather than on the headline numbers.
## Criticisms and caveats
- **Early AutoRAG was rough.** An April 2025 hands-on review found answers decent but retrieval weak, only two embedding models, Llama-only generation, 1.7 s query rewriting and no BM25 or reranking. Most of those gaps are closed now; the request for built-in retrieval evaluation still isn't. A Hacker News commenter in October 2025 summed up the beta mood bluntly: AutoRAG "has huge issues too"
- **The free tier shrank.** During the April 2026 beta, Workers Free got 20,000 queries a month and Workers Paid unlimited. At GA, both plans get 1,000 semantic plus 1,000 full-text queries. Anything beyond a hobby project will pay for queries
- **Two bills, not one.** Retrieval is bundled, but answer generation and query rewriting still land on your Workers AI or AI Gateway invoice
- **Lock-in by design.** Data sources are R2, uploads and websites on your own Cloudflare zones. There's no connector for Drive, Notion or Confluence, and the website source refuses domains you don't host on Cloudflare
- **Public means public.** Public endpoints have no authentication by default, CORS isn't access control, and when you add Access on a custom domain the default `search.ai.cloudflare.com` hostname keeps answering until you set `default_domain_enabled` to `false`
- **Keyword search at scale is still a work in progress.** Hybrid halves the file cap, and the GA post says Cloudflare is refactoring the keyword engine because large stores hit limits
- **Migration leftovers.** Instances that crawled a website before June 2026 left a dedicated R2 bucket in your account that AI Search no longer uses but that may still count toward R2 storage billing
- **Your own bot rules apply to the crawler.** WAF, Bot Management or Turnstile can block it until you allowlist it (Bot Detection ID `122933950`)
## My take
If your content already lives on Cloudflare, this is the shortest path from "a pile of docs" to search for your users AND an MCP tool for agents. Hybrid search on by default, scoring details per chunk and namespaces for per-tenant indexes are the right defaults; I'd still keep reranking off until I've measured that it helps on my own queries. If you need fine control over chunking, custom retrieval logic or embeddings you can swap, drop down to [[Cloudflare Vectorize]] and build the pipeline yourself (see [[RAG Pipelines]]). And watch the query bill: at $0.75 per 1,000, an agent that searches ten times per task gets expensive faster than you'd think.
## References
- Documentation: https://developers.cloudflare.com/ai-search/
- How AI Search works: https://developers.cloudflare.com/ai-search/concepts/how-ai-search-works/
- Limits and pricing: https://developers.cloudflare.com/ai-search/platform/limits-pricing/
- Release notes: https://developers.cloudflare.com/ai-search/platform/release-note/
- Supported models: https://developers.cloudflare.com/ai-search/configuration/models/supported-models/
- Hybrid search: https://developers.cloudflare.com/ai-search/configuration/indexing/hybrid-search/
- Workers binding migration from `env.AI.autorag()`: https://developers.cloudflare.com/ai-search/api/migration/workers-binding/
- MCP endpoint: https://developers.cloudflare.com/ai-search/api/search/mcp/
- NLWeb: https://developers.cloudflare.com/ai-search/how-to/nlweb/
- Blog, Introducing AutoRAG (2025-04-07): https://blog.cloudflare.com/introducing-autorag-on-cloudflare/
- Blog, AI Search: the search primitive for your agents (2026-04-16): https://blog.cloudflare.com/ai-search-agent-primitive/
- Blog, give your agents a search engine for your data (2026-08-06): https://blog.cloudflare.com/ai-search-easier/
- Blog, AI Search is now generally available (2026-10-01): https://blog.cloudflare.com/ai-search-ga/
- Changelog, AutoRAG open beta (2025-04-07): https://developers.cloudflare.com/changelog/post/2025-04-07-autorag-open-beta/
- Changelog, AutoRAG renamed AI Search (2025-09-25): https://developers.cloudflare.com/changelog/post/2025-09-25-ai-search-more-models/
- Changelog, reranking and system prompts (2025-10-28): https://developers.cloudflare.com/changelog/post/2025-10-27-ai-search-reranking-system-prompt/
- Changelog, new REST API endpoints (2026-03-23): https://developers.cloudflare.com/changelog/post/2026-03-23-ai-search-new-rest-api/
- Changelog, public endpoints, UI snippets and MCP (2026-03-23): https://developers.cloudflare.com/changelog/post/2026-03-23-ai-search-public-endpoint-and-snippets/
- Changelog, hybrid search and relevance boosting (2026-04-16): https://developers.cloudflare.com/changelog/post/2026-04-16-hybrid-search-and-relevance-boosting/
- Changelog, built-in storage and namespace bindings (2026-04-16): https://developers.cloudflare.com/changelog/post/2026-04-16-ai-search-namespace-binding/
- Changelog, similarity cache controls (2026-06-24): https://developers.cloudflare.com/changelog/post/2026-06-24-ai-search-similarity-cache-controls/
- Changelog, custom domains, Access and namespace endpoints (2026-08-06): https://developers.cloudflare.com/changelog/post/2026-08-06-public-endpoint-custom-domains-and-namespaces/
- Changelog, AI Search GA (2026-10-01): https://developers.cloudflare.com/changelog/post/2026-10-01-ai-search-generally-available/
- Cloudflare AutoRAG first impressions, Pranit Bauva (2025-04-26): https://bauva.com/blog/cloudflare-autorag-first-impressions/
- Hacker News comment on AutoRAG (2025-10-15): https://news.ycombinator.com/item?id=45587797
- OpenAI pricing (file search): https://platform.openai.com/docs/pricing
- OpenAI retrieval and vector stores: https://platform.openai.com/docs/guides/retrieval
- Google Cloud Agent Search (Vertex AI Search) pricing: https://cloud.google.com/generative-ai-app-builder/pricing
- Amazon Bedrock pricing (Knowledge Bases): https://aws.amazon.com/bedrock/pricing/
- Amazon Bedrock Managed Knowledge Base announcement (2026-06): https://aws.amazon.com/blogs/aws/introducing-amazon-bedrock-managed-knowledge-base-for-faster-more-accurate-enterprise-ai-applications/
- NLWeb project: https://github.com/nlweb-ai/NLWeb
## Related
- [[Cloudflare]]
- [[Cloudflare Vectorize]]
- [[Cloudflare Workers AI]]
- [[Cloudflare R2]]
- [[Cloudflare AI Gateway]]
- [[Cloudflare Workers]]
- [[Cloudflare Browser Rendering]]
- [[Cloudflare Agents SDK]]
- [[Cloudflare Access and Zero Trust]]
- [[Wrangler]]
- [[Retrieval-Augmented Generation (RAG)]]
- [[RAG Pipelines]]
- [[Embeddings]]
- [[Matryoshka Embeddings]]
- [[Vector Store]]
- [[Semantic Search]]
- [[Reranking]]
- [[AI Retrieval Patterns]]
- [[Model Context Protocol (MCP)]]
- [[LLM Knowledge Bases Over Unstructured Data]]
- [[Vertex AI]]
- [[OpenAI]]