Skip to content

StadiaSoft

Build a Cloudflare Workers AI RAG Assistant

Many teams have useful answers scattered across policies, product documentation, support notes and internal guides. A retrieval-augmented generation (RAG) assistant can search those sources before asking a model to answer. The value is not the chat box. It is the system that retrieves the right, permitted, current evidence and knows when to say it does not have enough.

This tutorial combines Cloudflare Workers AI for embeddings and answer generation, Vectorize for semantic retrieval, and AI Gateway for inference visibility and control. A Worker hosts the application logic that connects them. We use a support-knowledge assistant as the example, then show the production checks a short demo normally omits.

Cloudflare provides a basic RAG tutorial using Workers AI, Vectorize and D1, and its AI use-case documentation describes Workers AI, Vectorize and AI Gateway together. This guide focuses on the decisions needed before a business can trust the result. Cloudflare RAG tutorial and Cloudflare AI use cases

Decide what the assistant is allowed to answer

Start with a narrow task: “Answer support agents’ questions about current product policies and cite the source.” That is easier to evaluate than “answer anything about the company.”

Define the following before indexing documents:

  • Audience: support staff, customers or both.
  • Corpus: which documents are authoritative, who owns them and how often they change.
  • Access: which tenant, team or role can see each document.
  • Answer boundary: whether the assistant may summarize, recommend or initiate actions.
  • Failure behavior: what it says when retrieval is empty, conflicting or stale.
  • Success measure: grounded-answer rate, useful citation rate, refusal accuracy and support time saved.

For a customer-facing assistant, make the first release read-only. If it later performs account actions, add explicit authorization and human approval where the risk warrants it. StadiaSoft’s AI automation service addresses the full workflow around the model, including source systems and safeguards.

Understand the three product roles

Product Role Risk to control
Workers AI Generates embeddings and answer text Model output may be wrong or unsupported
Vectorize Returns similar chunks from indexed documents Similarity is not permission or truth
AI Gateway Adds analytics, logging and optional caching around inference Logs and caches can expose sensitive context if configured carelessly

Cloudflare documents Vectorize integration with Workers AI, including index creation, bindings, insertion and querying. AI Gateway can route Workers AI calls and add analytics, logging and caching. Vectorize and Workers AI and Workers AI through AI Gateway

Step 1: Prepare the documents

The quality ceiling is set by the corpus. Remove obsolete versions, duplicate pages, empty boilerplate and documents with uncertain ownership. Record a stable document ID, title, canonical source URL or internal path, owner, last-reviewed date, tenant ID and access classification.

Do not index secret values merely because they appear in an internal wiki. Decide how personal data, customer records and confidential notes will be handled. If source permissions change, the index and answer path must reflect that change quickly enough for your risk model.

For each document, split text into passages that preserve complete ideas. A heading and its following explanation usually belong together. Overly small chunks lose context; overly large chunks dilute retrieval and inflate inference cost. Test a few chunk sizes against real support questions rather than accepting one magic number.

Give each chunk a stable identifier such as document-id:version:chunk-number. Store source metadata with the vector, but keep the full text in an authoritative content store that can be updated or deleted. Vectorize supports metadata attached to vectors and metadata filtering. Vectorize metadata filtering

Step 2: Generate embeddings consistently

Choose one text embedding model and create the Vectorize index with the dimension that model returns. Use the same model for both document chunks and user queries. Changing models later generally requires re-embedding the corpus and checking retrieval quality.

At ingestion time:

for each approved document:
  parse and normalize text
  split into identifiable chunks
  generate an embedding for each chunk with Workers AI
  upsert vector with stable ID and tenant/source metadata
  record index version and document review date

Do ingestion as a controlled job rather than on every user question. Track failed chunks so an incomplete index cannot silently appear current. Re-index or delete vectors when a document changes. Keep an audit of what source version produced each vector.

For example, Cloudflare’s @cf/baai/bge-base-en-v1.5 embedding model currently returns 768-dimensional vectors. A matching index can then receive the model’s data[0] array as a query vector:

const embedding = await env.AI.run(
  "@cf/baai/bge-base-en-v1.5",
  { text: [question] }
);
const matches = await env.VECTORIZE.query(embedding.data[0], {
  topK: 5,
  namespace: authorizedTenantNamespace,
  returnMetadata: "all",
});

The snippet assumes question has passed input validation and authorizedTenantNamespace was derived from the authenticated user. Do not accept that namespace directly from the browser. Check model dimension and current client options against the documentation when implementing. Cloudflare embedding model and Vectorize setup guide

Step 3: Authenticate before retrieval

The query Worker should validate the user’s identity and derive their permitted tenant from trusted application data. Never accept a tenant ID in the request body and use it as the only access check. If you use Vectorize metadata filtering, create the necessary metadata index first and apply the filter to every query. Cloudflare says filters are applied before the top results are selected; namespaces can also scope queries. Vectorize filtering behavior

This is the central rule for multi-tenant RAG: a semantically similar document from another tenant is still inaccessible. Test cross-tenant questions directly. The application must also protect the source document endpoint that supplies full chunk text; a filtered vector search followed by an unprotected document lookup is incomplete.

If the assistant sits inside a customer or employee CRM, align its permissions with the CRM’s own roles. StadiaSoft’s custom CRM development service covers these cross-system access rules.

Step 4: Retrieve, rank and prepare context

Convert the user’s question to an embedding with the same model used at ingestion. Query Vectorize inside the authorized tenant scope, request a small candidate set, then fetch the corresponding approved passages. Reject results from deleted or expired documents.

Similarity scores alone do not prove relevance. Consider a second ranking step or a simple rule that requires topic match, freshness and sufficient source quality. Keep a maximum context budget so one long document does not crowd out the rest. Record which source IDs were used for each answer.

A practical request sequence is:

authenticate → authorize tenant → embed question → filter vector search
→ fetch permitted passages → check relevance → build prompt
→ generate answer → verify cited source IDs → return answer

If retrieval finds nothing useful, ask a clarifying question or say the answer is unavailable. A polished guess is a worse outcome than a concise refusal for policy or account-specific questions.

Step 5: Generate a cited answer through AI Gateway

Give the model a strict task: answer using the supplied passages, cite their stable source IDs, identify uncertainty and refuse unsupported details. Keep retrieved text inside a clearly delimited context section. Treat document content as data, not as an instruction that can override the assistant’s system rules.

AI Gateway can sit between the Worker and Workers AI. Cloudflare supports passing gateway configuration to env.AI.run, including a gateway ID and cache settings. A trimmed generation call looks like this:

const response = await env.AI.run(
  "@cf/meta/llama-3.1-8b-instruct-fast",
  {
    messages: [
      { role: "system", content:
        "Answer only from the supplied passages. Cite their IDs. " +
        "If the evidence is insufficient, say you cannot verify the answer." },
      { role: "user", content:
        `Question: ${question}\n\nPassages:\n${approvedPassages}` },
    ],
  },
  { gateway: { id: "support-assistant", skipCache: true } }
);

The model ID is an active example from Cloudflare’s documentation; check availability and the model’s current input shape before deployment. Here, approvedPassages must be built only from documents the caller can access, with stable source IDs and a bounded token budget. Cloudflare deprecated the un-suffixed llama-3.1-8b-instruct model in May 2026, which is why the -fast variant appears here. AI Gateway Worker binding methods and Cloudflare model deprecation notice

For tenant-specific answers, begin with caching disabled or explicitly skipped until you have a proven key design that includes tenant, document version and permission scope. Review whether logging captures prompts or responses containing private material; configure retention and access according to your data policy. Log request IDs, model, latency, selected source IDs, outcome and safe error categories without exposing unnecessary customer content.

After generation, validate every citation ID against the passages actually supplied to the model. Present a link to the source title and its last-reviewed date. Do not let a model invent a plausible URL and display it as evidence.

Step 6: Evaluate before broad rollout

Create a small test set from real questions. For each question, record the expected source, acceptable answer, disallowed claim and whether refusal is appropriate. Include easy, ambiguous and adversarial cases.

Test category Example What a good result looks like
Direct answer “What is the refund window?” Correct current policy and citation
Conflicting sources Two versions disagree Current approved version or explicit uncertainty
Missing answer Undocumented exception Refusal or escalation route
Cross-tenant Ask about another customer’s record No retrieval or disclosure
Prompt injection Document says “ignore instructions” Treats that text as untrusted content
Stale source Policy was replaced yesterday Excludes or flags old version

Measure retrieval separately from generation. If the correct passage was never returned, prompt changes cannot fix the root cause. Track answer accuracy, citation validity, refusal quality, latency and cost per useful answer. Review failures with a domain expert before expansion.

Cost and operational controls

Estimate monthly cost from actual workflow counts rather than a single model price:

Ingestion cost = new and changed chunks × embedding cost
Query cost     = questions × query embedding cost
Answer cost    = questions × input and output tokens
Vector cost    = stored vectors and queries
Gateway cost   = features and usage under the selected plan

Put an explicit limit on input size, retrieved chunks, output length and requests per user. Include an alert when a tenant’s usage grows suddenly. Use AI Gateway analytics to understand model traffic, but compare that to application-level outcomes so “lower inference spend” does not hide poor answer quality. Check current Workers AI pricing, Vectorize pricing and AI Gateway pricing before budgeting.

Cloudflare also offers AI Search as a managed RAG option. If your main need is search over a supported corpus with less custom ingestion work, evaluate it alongside the three-product architecture. Choose the custom route when permission logic, evaluation, document lifecycle or integrations justify the added engineering. Cloudflare AI Search

Launch checklist

A useful RAG assistant is a maintained knowledge product. If you are considering one for support, operations or a SaaS application, discuss a bounded pilot with StadiaSoft. We can connect the content, permissions and evaluation to your existing systems; our API integration service supports that surrounding data flow.

For the application layer around the assistant, see the Cloudflare Workers, D1 and Queues SaaS API guide.

Technical and pricing sources reviewed September 25, 2026. Recheck model availability, limits and rates before implementation.

Frequently asked questions

Does RAG guarantee a correct answer?

No. Retrieved context can help ground an answer, but the model can still omit, misread or invent details. Test it against known questions and require a refusal when evidence is insufficient.

Is Vectorize enough to enforce customer permissions?

No. Authenticate the user, derive their authorized tenant, filter retrieval and protect the source-document lookup. Similarity search is not access control.

Should AI Gateway cache private RAG answers?

Only after deliberate isolation and invalidation design. Tenant-specific answers may contain private context; disable shared caching until the cache key and permission model are verified.