Books → Learning Ai From Scratch → Introduction → Chapter 12


Chapter 12

Chapter 12

RAG Architecture

Retrieval-Augmented Generation for Practical, Grounded AI Applications

 

Learning Objectives

  • Understand what Retrieval-Augmented Generation (RAG) is and why it is used in modern LLM applications.
  • Learn the difference between standalone LLM answers and grounded answers generated using private or external knowledge.
  • Understand the complete RAG indexing pipeline and retrieval pipeline from documents to final answer.
  • Learn the role of document loading, chunking, embeddings, vector storage, retrieval, re-ranking, context building, prompt construction, source citation, answer validation, and feedback loops.
  • Compare Basic RAG, Advanced RAG, Agentic RAG, Multi-query RAG, Hybrid RAG, Graph RAG, Corrective RAG, and Conversational RAG.
  • Study practical examples such as HR policy chatbot, banking knowledge assistant, and technical support chatbot.
  • Prepare a production checklist for designing reliable and secure RAG systems.

Estimated reading time: 2 to 3 hours. Recommended practice time: 4 to 6 hours.

 

 

Chapter Index

1. 12.1 What is Retrieval-Augmented Generation?

2. 12.2 Why RAG is Needed

3. 12.3 Problems with Standalone LLMs

4. 12.4 Knowledge Cutoff and Fresh Information

5. 12.5 Hallucination Reduction

6. 12.6 Private Data Usage

7. 12.7 Complete RAG Architecture Diagram

8. 12.8 RAG Indexing Pipeline

9. 12.9 RAG Retrieval Pipeline

10. 12.10 Core Components: Loading, Chunking, Embedding, Vector Storage, Retrieval, Re-ranking, Context Building, Prompt Construction, LLM Generation, Citation, Validation, and Feedback

11. 12.11 RAG Patterns: Basic, Advanced, Agentic, Multi-query, Hybrid, Graph, Corrective, and Conversational RAG

12. 12.12 Example 1: HR Policy Chatbot

13. 12.13 Example 2: Banking Knowledge Assistant

14. 12.14 Example 3: Technical Support Chatbot

15. 12.15 Common RAG Mistakes

16. 12.16 RAG Production Checklist

17. 12.17 Mini Project: Build a Small RAG Application

18. 12.18 Chapter Summary, Key Terms, and Exercises

 

 

12.1 What is Retrieval-Augmented Generation?

Retrieval-Augmented Generation, usually called RAG, is an architecture pattern used to make Large Language Model applications answer questions using relevant information retrieved from a trusted knowledge source. The word retrieval means searching and fetching useful information. The word augmented means adding that retrieved information to the prompt. The word generation means the LLM generates the final answer using both the user question and the retrieved context.

A simple way to understand RAG is this: instead of asking an LLM to answer only from its memory, we first give it the right pages, records, or knowledge snippets, and then ask it to answer based on those materials. This makes the answer more relevant to your company, your documents, your latest data, and your specific business process.

For example, suppose an employee asks: “How many days of casual leave do I have according to our company policy?” A standalone LLM may know general leave policy concepts, but it does not automatically know your company policy. A RAG system first searches the company HR policy document, retrieves the relevant leave policy paragraphs, adds them to the prompt, and then asks the LLM to answer using only that context.

User Question
     |
     v
Retrieve relevant information from trusted data
     |
     v
Add retrieved information to the LLM prompt
     |
     v
LLM generates grounded answer
     |
     v
Answer with source reference

RAG is not a model by itself. It is an application architecture. It combines search, data engineering, embeddings, vector databases, prompt engineering, and LLM generation into one workflow.

Simple Daily Life Analogy

Imagine a student is asked to answer a question in an exam. There are two ways the student can answer. First, the student can answer only from memory. Second, the student can open a textbook, find the exact page, read the relevant paragraph, and then answer. A standalone LLM is similar to answering from memory. RAG is similar to answering after opening the relevant textbook page.

The important point is that RAG does not blindly copy the document. It retrieves the relevant material and uses the LLM to convert it into a clear, natural language answer.

12.2 Why RAG is Needed

RAG is needed because business AI applications usually require accurate, current, private, and explainable answers. General-purpose LLMs are trained on large public and licensed datasets, but your enterprise knowledge is usually stored in internal files, databases, policies, tickets, manuals, emails, wikis, PDFs, and business systems. The LLM needs a safe way to use this information without retraining the model every time the data changes.

RAG gives this capability. It allows the system to connect an LLM with external knowledge sources at answer time. This is why RAG is widely used for document question answering, policy assistants, customer support bots, technical support assistants, banking knowledge assistants, legal document assistants, healthcare document assistants, and enterprise search systems.

Business Need

Why Standalone LLM May Fail

How RAG Helps

Use company-specific policy

The model may not know internal policies.

Retrieves policy documents and answers from them.

Use latest information

Training data may be old.

Retrieves current documents or database records.

Show sources

LLM may not cite reliable evidence.

Returns answer with retrieved source references.

Reduce hallucination

Model may invent missing details.

Restricts answer to retrieved context.

Avoid model retraining

Fine-tuning is expensive and slow for frequent updates.

Update the knowledge index instead of retraining the model.

Protect private data

Private data should not be baked into a public model.

Keeps data in controlled storage and passes only needed context.

 

RAG as a Practical Bridge

RAG is a practical bridge between traditional enterprise systems and modern LLMs. Your existing documents, databases, and knowledge bases remain the source of truth. The LLM becomes the reasoning and communication layer that explains retrieved information in human language. This is why RAG is often the first serious architecture used when companies move from simple chatbot experiments to real AI applications.

12.3 Problems with Standalone LLMs

A standalone LLM can be very useful for general explanation, writing assistance, summarization, brainstorming, and coding help. But when it is used directly for company-specific decisions, it has several limitations.

Problem

Explanation

Example

No access to private data

The model does not automatically know your internal documents or databases.

It cannot know the latest employee reimbursement policy stored in your HR portal.

Knowledge cutoff

The model may not know information created after its training or update period.

It may not know a new product policy released this month.

Hallucination

The model may produce confident but incorrect information when context is missing.

It may invent an approval limit for loan processing.

No source reference

The model may answer without showing where the answer came from.

A compliance officer cannot verify the answer.

Weak enterprise control

It is difficult to force answers to use only approved sources.

A bot may mix public knowledge with internal policy.

Expensive retraining

Updating the model for every document change is not practical.

An HR policy changes weekly; retraining weekly is not realistic.

 

RAG does not solve every problem automatically, but it gives a strong foundation for controlling where answers come from. A well-designed RAG system tells the LLM: “Here is the relevant information. Answer using this information. If the answer is not present, say that you do not know.”

12.4 Knowledge Cutoff and Fresh Information

Knowledge cutoff means the model may not know events, documents, rules, or facts that were created after its training data was collected. Even when a model has current web or tool access, an enterprise application should not depend only on the model’s internal memory for business-critical facts.

In a company, information changes constantly. New HR policies are published, banking interest rates are updated, product manuals are revised, security processes change, and support articles are corrected. A standalone model cannot be expected to remember every internal update. RAG solves this by retrieving from sources that can be refreshed daily, hourly, or even near real time.

Example: Latest Banking Policy

Suppose a bank changes its home loan prepayment policy. A customer asks the AI assistant about prepayment charges. A standalone LLM may answer based on old or general banking knowledge. A RAG-based assistant can retrieve the latest policy document or database record and answer according to the current bank-approved rule.

12.5 Hallucination Reduction

Hallucination means the AI gives an answer that sounds correct but is not actually supported by facts. Hallucination can happen when the model is asked about missing, ambiguous, or very specific information. In business applications, hallucination can create serious risk, especially in banking, healthcare, legal, HR, and compliance workflows.

RAG reduces hallucination by grounding the model in retrieved evidence. The model is not simply asked, “What is the policy?” Instead, it is given the relevant policy text and instructed to answer from that text. This does not guarantee perfection, but it greatly improves control and verifiability.

Without RAG:
User asks -> LLM answers from memory -> Possible invented detail

With RAG:
User asks -> Search approved sources -> Provide evidence to LLM -> Grounded answer with source

A good RAG system also includes guardrails. For example, if retrieval does not find enough evidence, the assistant should say: “I could not find this information in the available documents,” instead of inventing an answer.

12.6 Private Data Usage

Private enterprise data is one of the main reasons RAG is important. Companies have knowledge stored in internal systems that cannot be placed into public training datasets. Examples include customer records, contracts, employee policies, product support manuals, internal wiki pages, service tickets, audit documents, and banking procedures.

RAG allows the application to use private data at query time. The private data remains in the company-controlled storage layer, such as a database, document store, data lake, or vector database. Only the relevant snippets are sent to the LLM as context, and only for the user request being processed.

However, using private data in RAG does not automatically make the system secure. The architecture must enforce authentication, authorization, data masking, access control, logging, and policy-based filtering. For example, an employee should not be able to retrieve payroll records belonging to other employees just because the vector database contains them.

Important Security Rule

In RAG systems, retrieval must respect user permissions. The retriever should only search documents that the current user is allowed to access. This is sometimes called permission-aware retrieval or access-controlled retrieval.

12.7 Complete RAG Architecture Diagram

The following text diagram shows a complete RAG architecture. In real projects, the exact components may vary, but the general flow remains similar.

                         COMPLETE RAG ARCHITECTURE

                 +---------------------------------------+
                 |            User Interface             |
                 |  Web app / Mobile app / Chat window   |
                 +-------------------+-------------------+
                                     |
                                     v
                 +---------------------------------------+
                 |        Backend API / Orchestrator      |
                 | Auth, logging, request validation      |
                 +-------------------+-------------------+
                                     |
          +--------------------------+--------------------------+
          |                                                     |
          v                                                     v
+-----------------------+                         +-------------------------+
| Conversation Memory   |                         | Prompt Template Manager |
| chat history summary  |                         | instructions, format    |
+-----------------------+                         +-------------------------+
          |                                                     |
          +--------------------------+--------------------------+
                                     |
                                     v
                       +-----------------------------+
                       | Query Processing Layer      |
                       | rewrite, classify, embed    |
                       +--------------+--------------+
                                      |
                                      v
                       +-----------------------------+
                       | Retriever                   |
                       | vector search + filters     |
                       +--------------+--------------+
                                      |
                                      v
                       +-----------------------------+
                       | Vector Database / Search    |
                       | chunks, embeddings, metadata|
                       +--------------+--------------+
                                      |
                                      v
                       +-----------------------------+
                       | Re-ranker / Context Builder |
                       | select best evidence        |
                       +--------------+--------------+
                                      |
                                      v
                       +-----------------------------+
                       | Final Prompt Construction   |
                       | question + context + rules  |
                       +--------------+--------------+
                                      |
                                      v
                       +-----------------------------+
                       | Large Language Model        |
                       | answer generation           |
                       +--------------+--------------+
                                      |
                                      v
                       +-----------------------------+
                       | Output Validation           |
                       | citations, safety, format   |
                       +--------------+--------------+
                                      |
                                      v
                       +-----------------------------+
                       | Final Answer to User        |
                       | answer + source citations   |
                       +-----------------------------+


             BACKGROUND INDEXING PIPELINE

Documents / DB / Web / Wiki / PDFs / Tickets
        |
        v
Document Loader -> Text Extraction -> Cleaning -> Chunking
        |
        v
Embedding Model -> Vector DB -> Metadata + Access Permissions

 

Notice that RAG has two major sides. The first side is the background indexing pipeline, where documents are prepared, chunked, embedded, and stored. The second side is the live retrieval pipeline, where a user question is processed, relevant chunks are retrieved, and the LLM generates an answer.

12.8 RAG Indexing Pipeline

The indexing pipeline prepares knowledge for future retrieval. It usually runs before users ask questions. It may run once for static documents, daily for updated documents, or continuously for systems where data changes frequently.

High-Level Indexing Flow

Source Documents
   |
   v
Document Loading
   |
   v
Text Extraction and Cleaning
   |
   v
Chunking
   |
   v
Metadata Attachment
   |
   v
Embedding Generation
   |
   v
Vector Storage and Indexing
   |
   v
Ready for Retrieval

Step 1: Document Loading

Document loading means collecting data from different sources. These sources may include PDF files, DOCX files, HTML pages, CSV files, database tables, SharePoint folders, cloud storage buckets, ticketing tools, wikis, CRM systems, or knowledge base articles.

The loader must understand how to read each source. A PDF loader extracts text from PDF files. A database loader reads rows from tables. A web loader reads HTML pages and removes unnecessary navigation content. A ticket loader reads fields such as ticket title, description, category, resolution, and date.

Step 2: Text Extraction and Cleaning

After loading documents, the system extracts clean text. Real documents often contain headers, footers, page numbers, repeated menus, broken line breaks, tables, images, signatures, and formatting noise. If noisy text is stored in the vector database, retrieval quality becomes poor.

Cleaning may include removing repeated headers, normalizing spaces, fixing broken paragraphs, converting tables into readable text, removing irrelevant legal footer text, and preserving important structure such as section headings and bullet points.

Step 3: Chunking

Chunking means splitting a large document into smaller pieces. LLM and embedding systems work better when each stored unit contains a focused topic. If a full 100-page HR policy is stored as one item, search may be too broad. If every sentence is stored separately, context may be incomplete. Good chunking balances precision and context.

Chunking Strategy

Description

When Useful

Fixed-size chunks

Split text by a fixed token or character count.

Simple documents and quick prototypes.

Paragraph-based chunks

Split by paragraphs or sections.

Policies, manuals, learning content.

Heading-based chunks

Use headings and subheadings to preserve structure.

Technical documentation and books.

Sliding window chunks

Use overlapping chunks so context is not lost at boundaries.

Long documents where important content spans sections.

Semantic chunks

Split by meaning, topic, or concept boundary.

High-quality enterprise RAG systems.

 

Step 4: Metadata Attachment

Metadata is information about the chunk. It may include document name, source URL, page number, section title, department, date, document version, access permission, language, product name, customer segment, or confidentiality level. Metadata is extremely important because it allows filtering and source citation.

Example metadata for one document chunk:
{
  "document_id": "HR-POLICY-2026",
  "title": "Employee Leave Policy",
  "section": "Casual Leave",
  "page": 7,
  "department": "HR",
  "version": "2026.1",
  "access_level": "employee",
  "source_url": "internal/hr/leave-policy"
}

Step 5: Embedding Generation

The embedding model converts each chunk into a vector, which is a list of numbers representing the meaning of the text. Similar texts have similar vectors. For example, a chunk about “annual leave” and a query about “vacation policy” may be close in vector space even if they do not use the exact same words.

Step 6: Vector Storage and Indexing

The chunk text, embedding vector, and metadata are stored in a vector database or vector index. The vector database builds an index so that similar chunks can be found quickly during user queries. This is what enables semantic retrieval.

A common indexing record contains: chunk ID, source document ID, chunk text, vector embedding, metadata, created date, updated date, and permission fields.

12.9 RAG Retrieval Pipeline

The retrieval pipeline runs when a user asks a question. Its goal is to find the most relevant evidence and pass it to the LLM in a controlled way.

High-Level Retrieval Flow

User Question
   |
   v
Validate user and request
   |
   v
Rewrite or clarify query if needed
   |
   v
Create query embedding
   |
   v
Search vector database and/or keyword index
   |
   v
Apply metadata and permission filters
   |
   v
Retrieve top candidate chunks
   |
   v
Re-rank and select best chunks
   |
   v
Build context
   |
   v
Construct final prompt
   |
   v
Call LLM
   |
   v
Validate answer and attach sources
   |
   v
Return final response

Step 1: User Query

The user query may be a simple question, a long instruction, or a follow-up in a conversation. For example: “What is the reimbursement limit for hotel stay?” or “Can you explain the policy that applies to my travel request?” The system must understand whether the query requires retrieval, calculation, classification, summarization, or escalation.

Step 2: Query Embedding

The system sends the user question to an embedding model and receives a query vector. This vector represents the meaning of the question. The vector database then compares this query vector with stored chunk vectors to find semantically similar chunks.

Step 3: Retrieval

Retrieval means searching the vector database and returning candidate chunks. The simplest method is top-K retrieval, where the system returns the K most similar chunks. If K is 5, the system returns the five most relevant chunks according to similarity score.

However, top-K retrieval alone is not enough in enterprise systems. The retriever should also apply filters. For example, if the user is from the HR department, the system may search HR policies. If the user asks about product X, the system should filter by product. If the document has access restrictions, the system must filter by permissions.

Step 4: Re-ranking

The first retrieval step may return many candidates. A re-ranker is used to reorder these chunks based on deeper relevance. Vector search is fast, but it may not always rank the best evidence first. A re-ranker reads the query and each candidate chunk more carefully and assigns a better relevance score.

Step 5: Context Building

Context building means selecting and arranging the final pieces of evidence that will be placed into the prompt. The context builder must respect token limits. It should include the most useful chunks, preserve source information, remove duplicates, and avoid irrelevant text.

Step 6: Prompt Construction

The final prompt usually contains instructions, the user question, retrieved context, citation rules, safety instructions, and output format rules. The prompt should clearly tell the model to answer only from the provided context and say when the information is not available.

Example RAG prompt template:
System instruction:
You are a company policy assistant. Answer using only the provided context.
If the answer is not found in the context, say: "I could not find this in the available policy documents."
Provide short source references after the answer.

User question:
{user_question}

Retrieved context:
[Source: {document_title}, Page {page_number}]
{chunk_text}

Answer format:
- Direct answer
- Important conditions
- Source reference

Step 7: LLM Generation

The LLM reads the final prompt and generates a natural language answer. A good answer should be clear, concise, grounded in the retrieved evidence, and formatted according to the required output style.

Step 8: Source Citation

Source citation means showing where the answer came from. Citations improve trust and allow users to verify the answer. In internal business systems, citations may point to document names, page numbers, policy sections, ticket IDs, or knowledge base articles.

Step 9: Answer Validation

Before returning the answer, the system may validate the response. Validation may check whether the answer uses sources, whether it contains prohibited content, whether it leaks private information, whether it follows the required JSON schema, and whether it admits uncertainty when the evidence is missing.

Step 10: Feedback Loop

A feedback loop collects user ratings, corrections, failed queries, missing documents, and low-confidence answers. This feedback is used to improve document quality, chunking, prompts, retrieval settings, re-ranking, and evaluation datasets.

12.10 Core Components in Detail

Document Loading

Document loading is the entry point of the knowledge pipeline. A loader connects to a source system and extracts content. In a small project, loading may mean reading files from a local folder. In an enterprise project, loading may require connectors for cloud storage, SharePoint, Confluence, ServiceNow, Salesforce, databases, email archives, ticketing systems, and internal APIs.

The loader should also capture metadata. Without metadata, the system may retrieve content but fail to tell the user where it came from. Metadata is also required for filtering, permissions, audit, and version control.

Chunking

Chunking is one of the most important design decisions in RAG. Bad chunking can destroy retrieval quality even when the LLM is very powerful. If chunks are too large, the retriever may return text that contains too many unrelated topics. If chunks are too small, the LLM may not receive enough context to answer correctly.

A practical beginner strategy is to start with heading-based chunks for documents and use overlap for long paragraphs. For example, chunk size can be around 500 to 1,000 tokens with 50 to 150 tokens of overlap. The exact value depends on the document type, embedding model, vector database, and use case.

Problem

Symptom

Possible Fix

Chunks too large

Retrieved context contains many unrelated details.

Reduce chunk size or split by headings.

Chunks too small

Answer misses conditions or exceptions.

Increase chunk size or add overlap.

No overlap

Important sentence is split between chunks.

Use sliding overlap.

Tables broken

Financial or policy data becomes unreadable.

Convert tables into structured text.

No section metadata

Answer cannot cite the right section.

Store heading, page, and document title metadata.

 

Embedding

Embedding converts text into numbers so that semantic similarity can be calculated. In RAG, embeddings are used twice: once during indexing for document chunks, and once during retrieval for the user query. The quality of the embedding model strongly affects retrieval quality.

For multilingual applications, choose an embedding model that supports the required languages. For domain-specific content, test whether the model understands domain terms. For example, banking terms such as NPA, KYC, EMI, and foreclosure may need careful evaluation.

Vector Storage

Vector storage holds embeddings and allows fast similarity search. The stored records should include the vector, chunk text, and metadata. In production, vector storage must also support backup, deletion, filtering, access control, indexing strategy, and refresh operations.

Retrieval

Retrieval is the process of finding candidate chunks. Retrieval may use vector search, keyword search, metadata filtering, or a combination. Good retrieval is more important than many beginners realize. If the retriever sends wrong context, the LLM may produce a wrong answer even if the model itself is strong.

Re-ranking

Re-ranking improves the order of retrieved chunks. It is useful when the first-stage retriever returns many potentially useful results. A common approach is to retrieve top 20 or top 50 chunks using a fast vector search, then re-rank them and pass only the best 3 to 8 chunks to the LLM.

Context Building

Context building decides what finally goes into the LLM prompt. The context builder should remove duplicates, group related chunks, include source information, respect token limits, and preserve important ordering. For example, in a policy document, it may be helpful to include the definition section before the rule section.

Prompt Construction

Prompt construction combines the system instruction, user question, retrieved context, output format, and safety rules. A weak prompt may allow the model to use unsupported assumptions. A strong RAG prompt clearly separates instructions from evidence and asks the model to answer only from the evidence.

LLM Generation

The LLM generates the answer. The model can summarize, explain, compare, extract, classify, or transform the retrieved context. For factual enterprise answers, the model temperature is often kept low so that answers are more stable and less creative.

Source Citation

Citations help the user trust and verify the answer. In a RAG chatbot, citations can be shown as document name, section title, page number, URL, ticket ID, policy version, or database record ID. Citations are also useful for auditing and troubleshooting.

Answer Validation

Answer validation checks whether the response follows rules. It may include schema validation for JSON, citation validation, PII detection, forbidden-topic checks, hallucination checks, and confidence scoring. For high-risk use cases, a human may review the answer before it is shown or acted upon.

Feedback Loop

The feedback loop improves the system over time. Users may mark an answer as helpful or incorrect. Analysts can review failed questions and identify whether the problem was missing data, poor chunking, poor retrieval, bad prompt design, or LLM reasoning error.

12.11 Different RAG Patterns

RAG is not one fixed design. Different use cases require different RAG patterns. A simple FAQ bot may use Basic RAG. A banking compliance assistant may require Advanced RAG with permissions, re-ranking, citations, evaluation, and human review. A research assistant may use Agentic RAG to decide which tools and searches to perform.

RAG Pattern

Main Idea

Best For

Complexity

Basic RAG

Retrieve top chunks and send them to the LLM.

Simple document Q&A and prototypes.

Low

Advanced RAG

Adds better chunking, filters, query rewriting, re-ranking, validation, and citations.

Enterprise knowledge assistants.

Medium

Agentic RAG

An agent plans retrieval steps and may call multiple tools.

Research, troubleshooting, complex workflows.

High

Multi-query RAG

Generates multiple search queries for one user question.

Ambiguous or broad questions.

Medium

Hybrid RAG

Combines vector search with keyword or structured search.

Domains where exact terms matter.

Medium

Graph RAG

Uses entity relationships and knowledge graphs with retrieval.

Connected knowledge, compliance, investigations.

High

Corrective RAG

Checks whether retrieved context is good; repairs retrieval if needed.

High-accuracy systems.

High

Conversational RAG

Uses chat history and follow-up handling.

Chatbots and assistants.

Medium

 

Basic RAG

Basic RAG is the simplest version. Documents are chunked, embedded, stored in a vector database, and retrieved when the user asks a question. The top chunks are inserted into the prompt and the LLM generates the answer.

Basic RAG:
User Question -> Query Embedding -> Vector Search -> Top Chunks -> Prompt -> LLM -> Answer

Basic RAG is excellent for learning and prototypes. It helps beginners understand the core idea. However, Basic RAG may fail when documents are large, queries are ambiguous, exact terms matter, permissions are important, or citations must be strict.

Advanced RAG

Advanced RAG improves reliability by adding more processing around retrieval and generation. It may include query rewriting, metadata filtering, hybrid search, re-ranking, context compression, citation validation, and answer evaluation.

Advanced RAG:
User Question
  -> Query Understanding
  -> Query Rewrite / Expansion
  -> Hybrid Retrieval
  -> Metadata and Permission Filters
  -> Re-ranking
  -> Context Builder
  -> Grounded Prompt
  -> LLM Answer
  -> Citation and Safety Validation

Advanced RAG is common in production enterprise applications. It requires more design effort, but the answer quality is usually much better than Basic RAG.

Agentic RAG

Agentic RAG uses an AI agent to decide how to retrieve information. Instead of one fixed retrieval step, the agent can plan multiple steps. It may search a knowledge base, call a database, inspect retrieved evidence, decide that more information is needed, and then search again.

For example, in a technical support chatbot, the agent may first identify the product, then retrieve troubleshooting steps, then call a system status API, and finally generate a response. Agentic RAG is powerful but more complex and riskier because the agent has more freedom.

Multi-query RAG

Multi-query RAG creates multiple versions of the user question to improve retrieval. This is useful because users may phrase questions differently from the document. For example, a user asks “Can I take leave for family emergency?” The system may create search queries such as “emergency leave policy,” “family emergency leave,” and “special leave conditions.”

User Question
   |
   +--> Query 1: original wording
   +--> Query 2: formal policy wording
   +--> Query 3: synonym-rich wording
   |
   v
Retrieve from all queries -> merge results -> re-rank -> answer

Hybrid RAG

Hybrid RAG combines semantic vector search with keyword search. This is useful when exact terms, codes, IDs, product names, legal clauses, error messages, or banking acronyms are important. Vector search understands meaning, but keyword search is often better for exact matches.

For example, if a user searches for error code “ORA-12154,” keyword search is very important. Vector search may understand database connection errors, but exact error-code matching is necessary for correct retrieval.

Graph RAG

Graph RAG uses relationships between entities. A knowledge graph stores entities such as customers, accounts, products, policies, departments, systems, tickets, and their relationships. Graph RAG can answer questions that require following connections rather than just finding similar text.

For example, in banking, a graph may connect customer accounts, loan products, KYC documents, risk categories, and policy rules. A Graph RAG system can retrieve relationship-based context and then ask the LLM to explain it.

Corrective RAG

Corrective RAG checks whether the retrieved context is actually useful. If retrieval quality is poor, the system may rewrite the query, search again, use a web or database tool, or ask the user for clarification. This reduces the chance of answering from weak evidence.

Conversational RAG

Conversational RAG handles multi-turn conversations. Users often ask follow-up questions such as “What about for managers?” or “Is this also applicable in probation?” The system must use conversation history carefully, rewrite the follow-up question into a complete standalone question, and retrieve the right context.

Conversation:
User: What is the travel reimbursement limit?
Assistant: ...
User: What about international travel?

Conversational RAG rewrite:
"What is the international travel reimbursement limit in the company travel policy?"

The challenge is to use relevant conversation history without overflowing the context window or accidentally using old assumptions.

12.12 Example 1: HR Policy Chatbot

Problem Statement

A company wants an AI chatbot that employees can use to ask questions about HR policies. Employees should be able to ask about leave, travel reimbursement, work from home, probation, resignation, code of conduct, and benefits. The chatbot should answer only from approved HR documents and show sources.

Architecture

HR PDFs / DOCX / Intranet Pages
       |
       v
Document Loader -> Text Cleaning -> Heading-based Chunking
       |
       v
Embedding Model -> Vector Database
       |
       v
Employee Question -> Auth Check -> Query Embedding
       |
       v
Permission Filter + Vector Search -> Re-ranker
       |
       v
Context Builder -> HR Prompt Template -> LLM
       |
       v
Answer + Policy Section + Page Reference

Sample User Questions

  • How many casual leaves are allowed in a year?
  • Can I carry forward unused leave?
  • What is the hotel reimbursement limit for domestic travel?
  • What is the notice period during probation?
  • Can I work from home for medical reasons?

Sample Retrieved Context

Source: Employee Leave Policy, Section 3.2, Page 6
Employees are eligible for 12 casual leave days per calendar year. Casual leave cannot be carried forward to the next calendar year. Unused casual leave expires at year-end.

Sample Answer

You are eligible for 12 casual leave days per calendar year. Casual leave cannot be carried forward. If you do not use it before the end of the calendar year, it expires. Source: Employee Leave Policy, Section 3.2, Page 6.

Important Design Points

  • Use only approved HR documents as sources.
  • Store document version and policy effective date as metadata.
  • Restrict sensitive HR documents using access control filters.
  • Use citations so employees can open the original policy.
  • Escalate to HR when the answer is missing or ambiguous.

12.13 Example 2: Banking Knowledge Assistant

Problem Statement

A bank wants an internal AI assistant for employees. Relationship managers and support staff need quick answers about products, KYC rules, account opening documents, loan eligibility, dispute handling, and compliance steps. The assistant must be accurate, cite sources, and avoid giving unsupported financial advice.

Architecture

Banking Policies / Product Manuals / Compliance SOPs / FAQs
       |
       v
Secure Loader -> OCR if needed -> Cleaning -> Chunking by Product and Rule
       |
       v
Embeddings + Metadata: product, segment, branch, effective date, access role
       |
       v
Vector DB + Keyword Index + Permission Filters
       |
       v
Employee Query -> Intent Detection -> Hybrid Retrieval -> Re-ranker
       |
       v
Context Builder -> Compliance Prompt -> LLM
       |
       v
Answer + Source + Disclaimer + Escalation Path

Why Banking Needs Advanced RAG

Banking content contains exact product names, interest-rate references, rule numbers, legal conditions, KYC requirements, and compliance restrictions. Semantic search alone may not be enough. Hybrid retrieval is usually better because exact terms such as “KYC,” “NPA,” “Form 60,” “PAN,” “foreclosure,” and product codes must be matched carefully.

Sample User Question

“What documents are required to open a savings account for a minor?”

Possible System Flow

1. Authenticate the employee and identify their access role.

2. Classify the query as account-opening policy.

3. Generate query embedding and search vector database.

4. Also run keyword search for “minor savings account documents.”

5. Filter by latest effective policy version.

6. Re-rank retrieved policy chunks.

7. Build context with required documents, conditions, and exceptions.

8. Generate answer with citation and compliance caution.

9. Log the interaction for audit.

Risk Controls

  • The assistant should not make promises to customers.
  • The assistant should cite official policy documents.
  • The assistant should warn users to verify final rules in core banking or compliance systems when required.
  • Sensitive customer data should not be exposed through retrieval.
  • All queries and retrieved sources should be audit logged.

12.14 Example 3: Technical Support Chatbot

Problem Statement

An IT support team wants a chatbot that can answer technical troubleshooting questions using product manuals, previous support tickets, known error codes, runbooks, and knowledge base articles. The bot should reduce repetitive support tickets and help support engineers resolve issues faster.

Architecture

Product Manuals / KB Articles / Resolved Tickets / Runbooks
       |
       v
Loader -> Clean -> Extract Error Codes -> Chunk by Procedure
       |
       v
Embeddings + Metadata: product, version, error code, OS, component
       |
       v
Vector DB + Keyword Search for exact error codes
       |
       v
User Query -> Product Detection -> Hybrid Retrieval
       |
       v
Re-ranker -> Context Builder -> Troubleshooting Prompt
       |
       v
LLM Answer -> Step-by-step Solution -> Ask for Logs if Needed

Example Query

“Oracle connection is failing with ORA-12154 from the VB6 application. What should I check?”

Why Hybrid Retrieval Helps

The exact error code ORA-12154 is very important. A pure semantic search may retrieve general Oracle connection issues, but keyword search ensures that documents containing the exact error code are prioritized. A good support RAG system stores error codes as metadata and supports exact matching.

Sample Answer Structure

  • Meaning of the error in simple terms.
  • Most likely causes.
  • Step-by-step checks.
  • Commands or configuration files to inspect.
  • When to escalate.
  • Source knowledge base links.

12.15 Common RAG Mistakes

Many RAG projects fail not because the idea is bad, but because the implementation ignores data quality, retrieval quality, security, or evaluation. The following mistakes are common in beginner and enterprise projects.

Mistake

What Happens

How to Avoid It

Using poor-quality documents

The system retrieves outdated, duplicated, or confusing content.

Clean documents and remove obsolete versions.

Bad chunking

Answers miss context or retrieve irrelevant sections.

Test chunk sizes and use heading-based splitting.

No metadata

Filtering and citations become weak.

Store title, page, section, version, department, and permissions.

No permission filtering

Users may retrieve restricted information.

Apply access control before retrieval.

Too many retrieved chunks

Prompt becomes noisy and expensive.

Use re-ranking and context selection.

Too few retrieved chunks

LLM lacks enough evidence.

Tune top-K and chunk size.

No evaluation dataset

Team cannot measure improvement.

Create golden questions and expected answers.

Ignoring missing answers

LLM invents information.

Prompt model to say when information is not found.

No feedback loop

Same failures repeat.

Collect user feedback and failed queries.

No source citation

Users cannot verify answers.

Return document references with answers.

Overusing RAG for everything

System becomes complex unnecessarily.

Use direct prompting or structured queries when simpler.

Ignoring cost

Token and vector DB costs grow quickly.

Cache, limit context, batch index, and monitor usage.

 

Mistake: Treating RAG as Magic

RAG is not magic. It is a carefully designed retrieval and generation system. If the documents are wrong, the chunks are bad, the retriever is weak, or the prompt is unclear, the output will be unreliable. A powerful LLM cannot fully compensate for poor retrieval.

Mistake: Not Testing Real User Questions

Teams often test RAG with ideal questions written by developers. Real users ask incomplete, messy, informal, and ambiguous questions. A good RAG system must be tested with real user language, abbreviations, spelling mistakes, and follow-up questions.

12.16 RAG Production Checklist

Before deploying a RAG system to production, use a checklist. The checklist should cover data, retrieval, prompts, security, evaluation, monitoring, cost, and operations.

Data Checklist

  • Have all source documents been approved by the business owner?
  • Are outdated documents removed or marked as inactive?
  • Are document versions and effective dates stored?
  • Are tables, forms, and scanned documents processed correctly?
  • Is metadata captured for title, source, page, section, owner, date, and permissions?
  • Is sensitive information masked or protected?

Retrieval Checklist

  • Is the embedding model suitable for the language and domain?
  • Is chunk size tested with real questions?
  • Is top-K tuned properly?
  • Is metadata filtering implemented?
  • Is permission-aware retrieval implemented?
  • Is hybrid search required for exact terms?
  • Is re-ranking used for important use cases?
  • Are duplicate chunks removed before prompt construction?

Prompt and Generation Checklist

  • Does the system prompt clearly say to answer only from provided context?
  • Does the prompt define what to do when the answer is missing?
  • Does the answer format match user needs?
  • Are citations required and validated?
  • Is temperature controlled for factual answers?
  • Are hallucination-risk instructions included?
  • Are output parsers used where structured output is required?

Security Checklist

  • Is user authentication implemented?
  • Are document permissions enforced before retrieval?
  • Are prompts protected against prompt injection?
  • Is sensitive data masked or blocked where needed?
  • Are logs protected and audited?
  • Are secrets stored securely?
  • Is there a policy for data sent to external LLM APIs?
  • Is human review required for high-risk outputs?

Evaluation Checklist

  • Is there a golden dataset of common questions?
  • Are expected answers and source documents defined?
  • Is retrieval quality measured separately from answer quality?
  • Are hallucinations tracked?
  • Are missing-answer cases tested?
  • Are regression tests run after document or prompt changes?
  • Is user feedback reviewed regularly?

Operations Checklist

  • Is there a document refresh schedule?
  • Is indexing monitored for failures?
  • Are vector database backups configured?
  • Can deleted documents be removed from the index?
  • Are latency and cost monitored?
  • Is there a fallback when the LLM or vector DB is unavailable?
  • Is there an escalation path for unanswered questions?

12.17 Mini Project: Build a Small RAG Application

This mini project helps you understand RAG practically. You can implement it later using Python, a document loader, an embedding model, a local vector database such as Chroma or FAISS, and an LLM API or local LLM. The goal is not to build a perfect production system, but to understand the full flow.

Project Goal

Build a small RAG chatbot that answers questions from 5 to 10 PDF or text documents. The documents can be HR policies, course notes, product FAQs, or your own learning notes.

Suggested Steps

1. Create a folder named data and place 5 to 10 documents inside it.

2. Load the documents and extract text.

3. Clean the text by removing extra spaces and irrelevant headers.

4. Split the documents into chunks.

5. Create embeddings for each chunk.

6. Store chunks, embeddings, and metadata in a vector database.

7. Accept a user question from the console or web UI.

8. Create an embedding for the question.

9. Retrieve top matching chunks.

10. Build a prompt using the retrieved chunks.

11. Call the LLM to generate an answer.

12. Display answer with document name and chunk reference.

13. Test with at least 20 questions.

14. Improve chunking, prompts, and top-K based on failures.

Simple Pseudo-code

# Indexing phase
for document in documents:
    text = load_and_clean(document)
    chunks = split_into_chunks(text)
    for chunk in chunks:
        vector = embedding_model.embed(chunk.text)
        vector_db.insert(
            id=chunk.id,
            vector=vector,
            text=chunk.text,
            metadata={"source": document.name, "page": chunk.page}
        )

# Retrieval phase
question = input("Ask a question: ")
query_vector = embedding_model.embed(question)
results = vector_db.search(query_vector, top_k=5)
context = build_context(results)
prompt = create_prompt(question, context)
answer = llm.generate(prompt)
print(answer)
print_sources(results)

Testing Questions

  • Ask direct questions where the answer is clearly present.
  • Ask questions using different words than the document.
  • Ask questions where the answer is not present.
  • Ask follow-up questions if you support chat history.
  • Ask ambiguous questions and check whether the system asks for clarification.
  • Ask questions that require exact numbers or conditions.

Expected Learning Outcome

After this mini project, you should understand how documents become chunks, how chunks become embeddings, how embeddings are searched, how retrieved context is added to the prompt, and how the LLM generates an answer. You will also understand that most RAG improvement work happens in data preparation, chunking, retrieval, prompt design, and evaluation.

12.18 Chapter Summary

RAG is one of the most important architecture patterns in modern Generative AI applications. It helps LLMs answer using trusted, private, and current information. Instead of relying only on model memory, a RAG system retrieves relevant context from documents or databases and passes that context to the LLM.

A complete RAG system has two major pipelines. The indexing pipeline loads documents, cleans text, chunks content, creates embeddings, and stores vectors with metadata. The retrieval pipeline takes a user question, creates a query embedding, retrieves relevant chunks, re-ranks them, builds context, constructs a prompt, calls the LLM, validates the answer, returns citations, and collects feedback.

RAG has many patterns. Basic RAG is useful for learning and prototypes. Advanced RAG adds better retrieval, filtering, re-ranking, validation, and citations. Agentic RAG allows an AI agent to plan retrieval steps. Multi-query RAG improves recall. Hybrid RAG combines vector and keyword search. Graph RAG uses relationships. Corrective RAG repairs bad retrieval. Conversational RAG handles follow-up questions.

Production RAG requires careful attention to data quality, security, permissions, evaluation, monitoring, and cost. A strong RAG system is not only an LLM call. It is a full AI application architecture.

Key Terms

Term

Meaning

RAG

Retrieval-Augmented Generation; an architecture that retrieves relevant information and uses an LLM to generate an answer.

Indexing pipeline

The background process that prepares documents for retrieval.

Retrieval pipeline

The live process that finds relevant context for a user question.

Document loader

A component that reads content from files, databases, websites, or business systems.

Chunking

Splitting large documents into smaller meaningful pieces.

Embedding

A numerical representation of text meaning.

Vector database

A database that stores embeddings and supports similarity search.

Retriever

A component that fetches candidate chunks for a query.

Re-ranker

A component that reorders retrieved chunks based on deeper relevance.

Context builder

A component that selects and formats evidence for the LLM prompt.

Prompt construction

Combining instructions, question, context, and output rules into the final prompt.

Source citation

A reference showing where the answer came from.

Hallucination

A confident but unsupported or incorrect model output.

Hybrid search

Combining semantic vector search with keyword search.

Graph RAG

RAG that uses relationships and knowledge graphs as part of retrieval.

Conversational RAG

RAG that uses conversation history and follow-up question handling.

 

Practice Exercises

Exercise 1: Explain RAG in Simple Words

Write a 10-line explanation of RAG for a non-technical manager. Use the textbook analogy or company document chatbot analogy.

Exercise 2: Identify RAG Components

Take the HR policy chatbot example and list each component: document source, loader, chunking strategy, metadata, embedding model, vector database, retriever, prompt template, LLM, citation method, and feedback method.

Exercise 3: Design a Chunking Strategy

Assume you have a 50-page employee handbook. Design a chunking strategy. Mention chunk size, overlap, metadata, and how you will handle tables.

Exercise 4: Create 20 Test Questions

Create 20 realistic questions for a RAG chatbot. Include direct questions, indirect questions, missing-answer questions, ambiguous questions, and follow-up questions.

Exercise 5: Compare RAG Patterns

Create a table comparing Basic RAG, Advanced RAG, Hybrid RAG, and Conversational RAG for a banking knowledge assistant.

Exercise 6: Production Risk Analysis

List 10 risks of deploying a RAG chatbot in a bank or healthcare company. For each risk, write one mitigation strategy.

Mini Assessment

1. What problem does RAG solve that a standalone LLM cannot solve reliably?

2. Why is chunking important in RAG?

3. What is the difference between indexing pipeline and retrieval pipeline?

4. Why should metadata be stored with each chunk?

5. What is re-ranking and when is it useful?

6. Why is hybrid search useful for technical support systems?

7. What should a RAG system do if it cannot find the answer in the retrieved context?

8. What is the difference between Basic RAG and Agentic RAG?

9. Why is permission-aware retrieval important?

10. Name five items from a production RAG checklist.

Suggested Practical Assignment

Choose one area from your work or study: HR policies, IT tickets, training notes, product FAQs, banking policies, or school notes. Collect a small set of documents. Design a RAG architecture on paper. Your design should include data sources, indexing pipeline, retrieval pipeline, prompt template, source citation method, evaluation questions, and production risks. Do not start with coding. First make the architecture clear. After that, implement a small prototype.

Transition to the Next Chapter

In this chapter, you learned how LLMs can be connected with external knowledge using RAG. The next chapters will build on this foundation by exploring LLM application patterns, agents, evaluation, monitoring, guardrails, cost, security, and complete example projects.