Guide · October 1, 2026
RAG From Pilot to Production: Retrieval, Evaluation, and Operations
Take retrieval-augmented generation from demo to production: chunking strategy, retrieval evaluation, grounding, access control, and ongoing operations.
Read as MarkdownMost retrieval-augmented generation (RAG) pilots start with a promising demo: a few documents loaded into a vector store, a handful of questions answered correctly, and a stakeholder who is ready to believe the hard part is done. Then the system meets the real corpus — tens of thousands of documents, contradictory content, PDFs with scanned tables, permissions that matter — and answer quality collapses. The pilot stalls not because the language model is weak, but because retrieval, evaluation, and operations were never treated as engineering problems.
Production RAG is primarily a data and retrieval problem. The model is the last mile. This guide walks through the decisions that separate a working demo from a dependable system: how documents are prepared, how retrieval is tuned and measured, how answers are kept grounded, how access control is enforced, and how the system is kept healthy after launch.
Chunking strategy
Chunking is the single highest-leverage decision in a RAG pipeline. The retriever can only return what the index contains, and the index only contains the chunks it was given. Bad chunks — fragments that cut mid-sentence, context-free paragraphs, tables flattened into noise — cap retrieval quality permanently, no matter how good the embedding model is.
Match chunking to document type
There is no universal chunk size. Start by classifying the corpus:
- Structured prose (policies, manuals, knowledge base articles): fixed-size chunks of roughly 300–800 tokens usually work. The unit of meaning is the paragraph or section, so chunks in this range keep enough context to be useful on their own.
- Long-form or dense documents (regulatory filings, contracts, research reports): prefer larger chunks (800–1,500 tokens) so that a retrieved passage still contains the full argument, not a severed fragment.
- Tabular and structured data (spreadsheets, price lists, comparison tables): avoid naive text extraction. Serialize tables in a format that preserves row and column relationships, such as markdown tables or row-level records with header metadata attached, so a chunk carries both the cell value and its meaning.
A common mistake is chunking everything the same way. A corpus that mixes policy PDFs, meeting notes, and spreadsheet exports needs different strategies per source, configured in the ingestion pipeline rather than applied globally.
Overlap, and when it helps
Overlapping chunks by 10–20 percent reduces the chance that a critical sentence lands exactly on a chunk boundary and gets split across two unretrievable halves. Overlap is cheap insurance at ingestion time, but it is not a substitute for semantic chunking. If documents are dominated by short, self-contained sections — FAQs, tickets, release notes — overlap adds storage and index cost for little benefit. Measure before adopting it as a default.
Structure-aware chunking
Naive character-based splitting ignores the document’s own structure. Wherever possible, parse the structure first and chunk along it: sections under headings, individual table rows or groups, list items as units. This preserves the semantic boundaries the author already created. PDF ingestion in particular benefits from structure-aware extraction — headings, captions, and page metadata retained at parse time become the retrieval context later. A chunk that knows which document, section, and page it came from is dramatically more useful than a floating paragraph.
Metadata is part of the chunk
Every chunk should carry metadata at write time: source document, document type, author or owner, last-updated date, version, and any classification labels (internal, confidential, region). Metadata serves two purposes. First, it enables filtered retrieval — answers restricted to documents newer than a date, or from an approved source. Second, it is the foundation for access control and auditability later. Reconstructing provenance after indexing is painful; capture it during ingestion, when it is free.
For a broader view of how data preparation fits into AI delivery work, see our capabilities.
Retrieval quality
Once chunks exist, the job is returning the right ones. A single vector search over embeddings is a reasonable starting point, but it rarely survives contact with a large, heterogeneous corpus. Production retrieval is layered.
Hybrid search
Dense (vector) search is strong on semantic similarity — it finds passages that mean the same thing as the query even when the wording differs. Sparse (keyword) search, such as BM25, is strong on exact matches — part numbers, error codes, proper nouns, acronyms. Enterprise corpora are full of exact-match queries, so hybrid retrieval that combines both scores consistently outperforms either alone. The combination can be as simple as a weighted score blend, or reciprocal rank fusion across two result lists. Either way, test both modalities separately first so you know which queries each one serves.
Reranking
Retrieval usually runs in two stages: a fast first pass that pulls a broad candidate set (say, the top 50–100 chunks), followed by a slower, more capable reranker that scores each candidate against the query and selects the top handful to pass to the model. Cross-encoder rerankers are the standard tool here — they read the query and chunk together rather than relying on embedding distance alone, and they meaningfully improve precision for the chunks that actually reach the model. The cost is latency, so rerank only the candidate set, not the whole index, and budget the added milliseconds explicitly.
Query rewriting
Real user queries are underspecified: “what’s the policy on that”, “the December version”, “did we change this”. Query rewriting transforms the raw input before retrieval — expanding acronyms, resolving pronouns against conversation history, adding domain synonyms. It can be rule-based or model-based, but it must be observable: log the rewritten query alongside the original so failures can be traced. A rewriting layer that silently invents terms is worse than none at all.
One more practical point: retrieval quality is bounded by the corpus. Before tuning the retriever further, check whether the answer even exists in the indexed documents. A surprisingly large share of “RAG failures” are actually coverage failures.
Evaluation
You cannot improve what you do not measure, and RAG systems degrade quietly. Evaluation needs to be built in from the start, with datasets that reflect real usage — not the five questions the team asked during the demo.
Build a golden dataset
A golden dataset is a set of representative questions, each paired with the correct answer and — critically — the specific document chunks that contain the evidence. Assemble it from real user queries, support tickets, and help-desk logs, not from questions written to be easy. Cover the hard cases: questions whose answers span multiple documents, questions with time-sensitive answers, questions that should not be answered at all (no relevant documents, or the documents are outside the user’s permissions). Aim for breadth over size initially; a few dozen well-chosen cases beat hundreds of near-duplicates. The dataset is a living artifact — add every production failure to it as a regression case.
Retrieval metrics
Before evaluating answers, evaluate retrieval, because retrieval is where most quality is won or lost. Standard metrics:
- Recall@k: of the evidence chunks in the golden set, what fraction appears in the top-k retrieved results. This measures coverage — if the right chunk is never retrieved, the model cannot answer correctly.
- Precision@k: of the retrieved chunks, what fraction is actually relevant. This measures noise — irrelevant chunks in context encourage hallucinations and waste context budget.
- MRR (mean reciprocal rank): how high the first relevant result ranks. Matters for systems that pass only a few chunks to the model.
Track these per query category, not just as global averages. A system with 90 percent average recall can still fail every query about the newest product line — and averages will hide it.
Answer-level metrics
Retrieval metrics tell you the model had the evidence; answer metrics tell you whether it used it correctly:
- Faithfulness: is every claim in the answer supported by the retrieved context. This is the core anti-hallucination metric.
- Answer relevance: does the response actually address the question asked, rather than reciting related-but-off-topic context.
- Citation accuracy: do the cited sources support the specific claims attached to them.
Human review on a sampled subset remains the gold standard for answer quality, especially early on. LLM-as-judge scoring can scale the volume, but validate the judge against human labels first — an uncalibrated judge optimizes for confident-sounding text, which is exactly the wrong target. For more on structuring evaluation into delivery, see our insights.
Grounding and citations
A production RAG system must attribute its claims. Citations serve the user — who needs to verify — and the organization — which needs auditability.
Answer attribution
Require the model to cite the specific chunks behind each substantive claim, and surface those citations in the interface: document title, section, and ideally a link or excerpt. This changes user behavior in a healthy way — users trust the system more and verify faster. Enforce attribution in the prompt and validate it programmatically where feasible: claims without citations, or citations that point to chunks never retrieved, should be flagged.
“I don’t know” behavior
The most important behavior in a production assistant is knowing when not to answer. Define it explicitly:
- When retrieval returns nothing relevant above a confidence threshold, the system should say so rather than answer from model memory.
- When retrieved documents conflict, the system should surface the conflict instead of picking a winner silently.
- When the question falls outside the indexed corpus, the system should decline rather than generalize.
These behaviors need to be in the golden dataset as test cases — including questions designed to elicit refusals — and measured like any other requirement. A system that never says “I don’t know” will eventually answer confidently from nothing, and that is the failure that ends pilots.
Access control and data leakage
In enterprise settings, the corpus is not uniformly visible. Documents carry owners, classifications, and permission sets, and the assistant must respect them — a wrong answer is a quality problem, but a leaked confidential document is a security incident.
Permissions-aware retrieval
The safest pattern is to filter at retrieval time: apply the user’s permission set as a filter on the search so unauthorized documents are never retrieved, never placed in context, and never seen by the model. This is preferable to retrieving everything and asking the model to withhold — models are not access-control mechanisms, and a determined or careless prompt can defeat instruction-level filtering. Permissions-aware retrieval requires the metadata captured during ingestion (owner, classification, group membership) to be current, which means the index must reflect permission changes with minimal lag.
PII handling
Enterprise corpora contain personal data — names, contact details, employee identifiers, customer records. Decide the policy before indexing: which fields are redacted at ingestion, which are masked at display time, and which are excluded entirely. Redaction at ingestion is the strongest protection but loses information; masking at retrieval preserves utility while protecting display. Either way, the policy must be explicit, documented, and tested — including adversarial tests where a user asks directly for personal data that exists in the corpus.
Logging and audit
Log what was retrieved for each query and which documents were cited in each answer. This supports incident investigation, access-pattern review, and compliance reporting. Treat retrieval logs with the same care as the documents themselves — they can reveal who accessed what.
Operating RAG in production
Shipping the system is the midpoint, not the finish line. RAG systems decay: documents change, new content arrives, queries drift, embeddings age. Operations is what keeps answers correct in month twelve.
Index refresh
Define a refresh strategy per source. Fast-changing sources (tickets, release notes, policy updates) need near-real-time or scheduled reindexing; stable sources (archived reports, historical records) can be indexed once. Handle updates as updates, not full rebuilds where possible — reindex only what changed, keyed by document version and last-modified date from ingestion metadata. And when documents are superseded, decide whether old versions stay searchable: for time-sensitive questions, version awareness matters, and the retriever needs metadata (effective dates, superseded flags) to prefer current content.
Monitoring
Monitor the full stack, not just uptime:
- Retrieval health: distribution of retrieval scores, rate of zero-result queries, reranker latency. A rising zero-result rate often means the corpus no longer matches what users ask.
- Answer quality: sampled human review on an ongoing basis, faithfulness scores, refusal rate. Watch for drift, not just incidents.
- User signals: which answers get corrected, which citations get clicked, which queries get rephrased. These are free labels.
- Cost and latency: tokens consumed per query (retrieval candidates plus generation), reranker inference cost, embedding cost at index time. RAG cost is dominated by context — every unnecessary chunk in the prompt is money spent and latency added.
Set thresholds and alerts on the metrics that predict user-visible quality, and feed production failures back into the golden dataset as regression cases.
Cost
RAG economics are straightforward to model once the pipeline is instrumented: embedding cost at index and reindex time, reranker inference per query, and generation cost driven by the number and size of chunks in context. The cheapest lever is usually context discipline — retrieving fewer, better chunks rather than stuffing the prompt. Evaluate the tradeoff explicitly: a reranker costs latency and compute, but if it halves the chunks needed in context while improving quality, it usually pays for itself.
Putting it into practice
The pattern across all of these decisions is the same: RAG quality comes from deliberate engineering at every stage — ingestion, retrieval, evaluation, grounding, access control, and operations — not from a single component. Organizations that treat the pilot as a prototype to be rebuilt with these disciplines in place tend to reach production on their second iteration; organizations that keep tuning the demo tend to stall.
If you are moving a retrieval-augmented system from pilot to production and want an outside review of the architecture, the evaluation setup, or the operating model, our engagements describe how we typically work with teams at this stage — and you can reach us through the contact page to start the conversation.