Define what a correct answer means
Retrieval-augmented generation (RAG) supplies a model with relevant source material before it answers. It can help with changing or private knowledge, but retrieval does not guarantee the retrieved material is correct, complete or authorized for the user.
Start with a narrow use case such as answering questions about a versioned product manual. Write examples of acceptable answers, the citations they need, and when the system should decline to answer or escalate. Keep answer generation separate from tools that change business records.
Prepare documents for retrieval
Preserve headings, table context, version numbers and source URLs. Give every chunk a stable document identifier and retain its access rules and revision date. Split on useful semantic boundaries rather than assuming one token count works for every document. Test questions whose answers span two sections or depend on a footnote.
Document deletion and permission changes must propagate to search indexes and caches. Re-embedding a document should not leave old, accessible chunks behind. Keep a record of ingestion failures so missing evidence is visible to operators.
Compare retrieval approaches
Exact terms such as error codes, SKUs and version numbers often benefit from keyword search. Embeddings can help with conceptual similarity. A hybrid system can combine both result sets and use rank fusion or reranking, but it adds cost and latency.
Build a labelled question set before adding stages. Measure whether relevant authorized passages appear in the retrieved set, then whether the final answer is supported by those passages. Include ambiguous, unanswerable and outdated questions. Precision at a chosen cutoff is a retrieval metric, not a measure of overall factual accuracy.
Enforce access before exposing context
- Authenticate the user and establish tenant and document permissions in trusted application code.
- Apply authorization constraints to retrieval and validate the final selected passages before they reach the model.
- Treat retrieved text as untrusted data. A document instruction to reveal secrets or call a tool must not override application policy.
- Restrict tools to necessary operations; validate arguments, permissions and business rules independently of model output.
- Require an appropriate confirmation or review step for consequential actions and retain a minimal audit trail.
Control costs without sharing the wrong answer
Track retrieval, embedding, reranking and generation costs separately, along with p50 and p95 latency. Limit context to passages that help answer the question, cap output length, and choose model capability against your evaluation set.
Start caching with exact matches and a clear invalidation strategy. Scope keys to tenant, access-policy version, document version, model and prompt version. Similar wording is not enough to establish that two users are entitled to the same answer. A semantic cache needs additional testing for false matches and stale answers; it still incurs retrieval or embedding work.
Release with a measurable acceptance gate
Set thresholds appropriate to the risk of the use case. Test citations, refusals, permission boundaries, malicious documents, timeouts and tool failures. Structured JSON can make parsing reliable, but a valid schema does not make a claim true or an action authorized.
Review failed answers and retrieval misses after release. Add them to a held-out regression set rather than tuning only to a demonstration. Show the user the supporting source and make uncertainty visible when the available material cannot resolve the question.
Frequently asked questions
Does RAG stop an AI system from inventing answers?
No. Retrieval can supply useful evidence, but the passages may be incomplete, outdated or irrelevant, and the model may misinterpret them. Evaluate whether the final answer is supported by the selected sources and define when the system should decline or escalate.
How do I keep one customer from seeing another customer’s documents?
Enforce tenant and document permissions in trusted application code before evidence reaches the model. Recheck the selected passages and scope caches to the relevant access policy and content versions. A prompt asking the model to respect permissions is not an authorization boundary.
Should I use keyword search, vector search or both?
Test representative questions first. Exact identifiers and error codes often need keyword matching, while embeddings can help with conceptual similarity. Compare retrieval quality, answer support, latency and cost before adding hybrid retrieval or reranking.
When is caching an AI answer safe?
Only when the cache key and invalidation rules preserve the answer’s permissions and freshness requirements. Include the tenant, policy, document, model and prompt versions where relevant. Similar wording alone does not prove that two users can receive the same answer.
Sources and further reading
Keep exploring
Explore AI application development or the queue and idempotency guide for reliable background workflows.

Leave a Reply