Books → Learning Ai From Scratch → Introduction → Chapter 15


Chapter 15

Chapter 15

Fine-tuning vs RAG vs Prompting

Choosing the right approach to build reliable, cost-effective, and maintainable LLM applications

 

Chapter Goal

By this point in the book, you have learned about LLMs, prompting, embeddings, semantic search, vector databases, RAG, and agents. This chapter teaches one of the most important practical decisions in GenAI projects: should you solve the problem using better prompts, RAG, fine-tuning, or a combination of all three?

 

Learning Objectives

  • Understand prompting, RAG, and fine-tuning in simple practical terms.
  • Know when prompting is enough and when it is not enough.
  • Know when RAG is the best choice for private, changing, or document-based knowledge.
  • Know when fine-tuning is useful for behavior, style, classification, or repeated task patterns.
  • Compare cost, complexity, data requirements, maintenance, security, and hallucination risk.
  • Use a decision tree to select the right approach for a real AI project.
  • Apply the decision process to customer support, legal assistant, brand-writing, and classification examples.

15.1 Introduction: The Most Common GenAI Design Confusion

When teams start building LLM applications, they often ask: “Should we fine-tune a model?” This is a natural question because traditional machine learning projects often involved collecting data, training a model, and deploying it. However, modern GenAI application development is different. Many practical business problems can be solved without training a new model at all.

In many cases, a well-designed prompt is enough. In other cases, the model needs access to company-specific documents, policies, product manuals, tickets, contracts, or database records. In that situation, RAG is usually better than fine-tuning. Fine-tuning becomes useful when you want the model to consistently follow a style, format, label taxonomy, tone, or task behavior across many examples.

A beginner should understand this clearly: prompting, RAG, and fine-tuning are not competitors in every situation. They are different tools for different problems. A production system may use all three together. For example, a customer support assistant may use prompt templates for structure, RAG for company knowledge, and fine-tuning for consistent ticket classification.

Simple Rule

Use prompting to control the instruction. Use RAG to provide external knowledge. Use fine-tuning to teach repeated behavior or style. Do not use fine-tuning just to add company documents.

 

15.2 What Is Prompting?

Prompting means giving instructions to an LLM so that it produces the type of answer you want. A prompt may be a simple question, a detailed task instruction, a role definition, an output format, examples, constraints, and context. Prompting is the fastest and cheapest way to control an LLM application.

A prompt does not change the model permanently. It only guides the model during the current request. If you send a different prompt, the model may behave differently. That is why prompt design is important in production systems. The prompt should be clear, stable, testable, and reusable.

Bad prompt:
Summarize this.

Better prompt:
You are a business analyst. Summarize the following meeting notes for a project manager.
Return the answer in this structure:
1. Key decisions
2. Open risks
3. Action items with owner and due date
Keep the summary under 250 words.

What Prompting Is Good For

  • Changing tone, format, or level of detail.
  • Asking the model to classify, summarize, extract, rewrite, translate, or explain.
  • Creating structured output such as JSON, tables, or bullet-point reports.
  • Testing an idea quickly before building a larger architecture.
  • Controlling the role of the model, such as tutor, reviewer, architect, analyst, or support agent.

Limitations of Prompting

Prompting cannot reliably give the model knowledge that is not in the prompt or not already known by the model. If your company policy is not included in the prompt, the model may guess. If the policy changes every month, a static prompt will become outdated. Prompting also has context-window limits. You cannot paste an entire knowledge base into every request.

Prompting Area

Good Use

Poor Use

Tone control

Write politely for customer email

Teach a new company knowledge base permanently

Output format

Return JSON with category and confidence

Store millions of documents inside the prompt

Quick prototype

Test a support reply assistant

Replace a full retrieval pipeline for large documents

Instruction control

Act as a Python tutor

Guarantee factual answers from private documents not provided

 

15.3 What Is RAG?

RAG stands for Retrieval-Augmented Generation. It is an architecture where the application retrieves relevant information from external sources and provides that information to the LLM as context before asking it to answer. The LLM does not have to remember everything. Instead, the system fetches the right information at runtime.

RAG is especially useful when answers must be based on private documents, frequently changing information, large knowledge bases, compliance documents, technical manuals, product data, or internal policies. Instead of fine-tuning the model on all documents, the system stores document chunks in a searchable index, retrieves relevant chunks for each question, and asks the LLM to answer using only those chunks.

Simple RAG flow:
User question
   -> Convert question to embedding
   -> Search vector database
   -> Retrieve relevant document chunks
   -> Build prompt with question + retrieved context
   -> LLM generates grounded answer
   -> Return answer with sources

RAG in One Line

RAG gives the LLM the right reference material at the time of answering.

 

What RAG Is Good For

  • Question-answering over company documents.
  • Chatbots for HR policies, banking products, insurance terms, product manuals, legal documents, and IT knowledge bases.
  • Reducing hallucination by grounding answers in retrieved sources.
  • Using private data without retraining the model.
  • Handling frequently changing content through document refresh and re-indexing.
  • Providing source citations so users can verify answers.

Limitations of RAG

RAG quality depends heavily on retrieval quality. If the retriever fetches the wrong chunks, the LLM may produce a weak answer. If documents are poorly chunked, metadata is missing, or the vector index is stale, the system will struggle. RAG also adds architectural complexity: document ingestion, chunking, embeddings, vector database, re-ranking, prompt construction, evaluation, and monitoring.

RAG Strength

RAG Challenge

Uses updated private knowledge

Needs indexing and refresh pipeline

Can cite sources

Source selection may be wrong if retrieval is poor

Good for large document collections

Chunking and metadata design become important

Usually cheaper than fine-tuning for knowledge use cases

Adds vector database and retrieval evaluation

Can reduce hallucination

Cannot eliminate hallucination completely

 

15.4 What Is Fine-tuning?

Fine-tuning means taking a pre-trained model and training it further on a smaller, task-specific dataset. The goal is not usually to teach the model a large knowledge base. The goal is to adjust the model behavior so it consistently performs a task, follows a format, uses a certain style, or maps inputs to outputs in a specialized way.

For example, suppose a company receives thousands of support tickets and has a fixed set of routing labels: Network, Hardware, Security, HR Payroll, Database, Application Support, and Access Management. A fine-tuned model can learn from many labelled examples and classify new tickets in the same style. Another example is a brand-writing assistant that consistently writes in a company-approved tone.

Fine-tuning dataset example for classification:
Input: "I cannot connect to VPN after password reset."
Output: {"category": "Network", "priority": "Medium"}

Input: "Salary slip is missing my reimbursement amount."
Output: {"category": "HR Payroll", "priority": "High"}

Fine-tuning in One Line

Fine-tuning changes model behavior using training examples; it is not the best first choice for storing changing documents.

 

What Fine-tuning Is Good For

  • Consistent classification across many examples.
  • Strict style or tone adaptation, such as brand voice.
  • Repeated structured output where prompting alone is unstable.
  • Domain-specific language patterns when examples are available.
  • Reducing prompt length by teaching recurring patterns into the model.
  • Improving small model performance for a narrow task.

Limitations of Fine-tuning

Fine-tuning requires good training data. Poor data creates poor behavior. Fine-tuning also requires evaluation, versioning, monitoring, and retraining when requirements change. It may increase cost and operational complexity. Most importantly, fine-tuning is not the best way to keep a model updated with a large set of changing company documents. For that, RAG is usually better.

15.5 How Prompting, RAG, and Fine-tuning Are Related

These three approaches operate at different layers of the AI system. Prompting is the instruction layer. RAG is the knowledge retrieval layer. Fine-tuning is the model behavior adaptation layer. In a complete application, they can work together.

Complete LLM application view:

User Request
   |
   v
Prompt Template  -> controls task, role, format, rules
   |
   +--> RAG Retrieval -> adds relevant private/current knowledge
   |
   v
LLM or Fine-tuned LLM -> generates response using learned behavior + provided context
   |
   v
Output Parser / Guardrails / Evaluation
   |
   v
Final Answer or Action

Approach

Main Purpose

Where It Fits

Prompting

Tell the model what to do now

Application/prompt layer

RAG

Provide relevant external knowledge at runtime

Retrieval/context layer

Fine-tuning

Teach repeated behavior or style

Model adaptation layer

 

15.6 When to Use Prompting

Use prompting first when the task can be solved by clear instructions and limited context. Prompting is ideal during early prototyping because it allows fast learning. Before investing in RAG or fine-tuning, try to solve the problem using a strong prompt and a few examples.

Use Prompting When...

Example

The task is simple and does not need private knowledge

Explain cloud computing to a beginner

You need a specific format

Return a JSON object with category, confidence, and explanation

You want tone control

Write a polite customer apology email

You are testing a prototype

Try a summarizer before building a full system

The context is small enough to fit in the prompt

Summarize one support ticket or one short email

 

Beginner Recommendation

Always start with prompting. If prompting fails because the model lacks knowledge, move to RAG. If prompting fails because behavior is inconsistent across many examples, consider fine-tuning.

 

15.7 When to Use RAG

Use RAG when the application must answer using information outside the model. This is common in enterprise systems because most useful business knowledge exists in documents, databases, tickets, PDFs, manuals, emails, policies, and reports. RAG lets the system retrieve relevant information without retraining the LLM.

Use RAG When...

Example

Knowledge is private

Internal HR policy chatbot

Knowledge changes frequently

Product pricing, circulars, release notes, service policies

The answer must cite sources

Legal, compliance, banking, insurance, medical-policy support

The document collection is large

Thousands of PDFs, manuals, tickets, articles

You need access control by user role

Employee sees HR docs; manager sees team policy docs

You want lower maintenance than retraining for knowledge updates

Re-index updated documents instead of fine-tuning

 

RAG decision example:
Question: "What is the maternity leave policy for employees in India?"
Required knowledge: Company HR policy PDF
Best approach: RAG
Reason: The answer must come from a private, possibly changing document.

15.8 When to Use Fine-tuning

Use fine-tuning when you have many good examples and want the model to learn a repeated pattern. Fine-tuning is useful when prompting becomes too long, too expensive, or too inconsistent. It is also useful when you need a smaller model to perform a narrow task reliably.

Use Fine-tuning When...

Example

You have high-quality labelled examples

10,000 support tickets mapped to categories

You need consistent output style

Brand-specific marketing copy

You need repeated structured behavior

Extract invoice fields in the same schema

Prompting is unstable despite strong examples

Same input type produces inconsistent labels

You want to reduce long prompts

Teach classification rules into the model

You are optimizing a narrow production task

Small model for fast ticket routing

 

Common Mistake

Do not fine-tune a model just because it answered one company-policy question incorrectly. If the problem is missing knowledge, use RAG. If the problem is repeated behavior or style, consider fine-tuning.

 

15.9 When to Combine Prompting, RAG, and Fine-tuning

Many mature systems combine all three. Prompting gives task instructions. RAG supplies external knowledge. Fine-tuning improves repeated behavior. The decision is not always “either-or.” The best architecture depends on the business requirement, risk, cost, and data availability.

Example combined customer support architecture:
1. Prompting: Define support-agent role, tone, escalation rules, and JSON output.
2. RAG: Retrieve relevant product manuals, policy documents, and known issues.
3. Fine-tuning: Improve classification of tickets into support categories.
4. Guardrails: Validate that refund or cancellation actions require human approval.

Combination

When Useful

Example

Prompting + RAG

Most document Q&A and knowledge assistants

HR policy chatbot with answer format rules

Prompting + Fine-tuning

Behavior/style tasks with structured output

Brand-writing assistant with strict format

RAG + Fine-tuning

Knowledge retrieval plus specialized behavior

Legal assistant that retrieves cases and writes in firm style

Prompting + RAG + Fine-tuning

Advanced production systems

Customer support system with retrieval, classification, and controlled response generation

 

15.10 Table Comparing Prompting, RAG, and Fine-tuning

Criteria

Prompting

RAG

Fine-tuning

Primary purpose

Control task instruction and output style

Add external/private/current knowledge

Adapt model behavior using examples

Best first step?

Yes, always start here

Use when knowledge is missing or large

Use after enough examples and clear need

Data needed

Small prompt/context/examples

Documents, chunks, metadata, embeddings

High-quality labelled training examples

Changes model weights?

No

No

Yes

Handles changing knowledge?

Poorly if knowledge is large

Very well with refresh pipeline

Poorly unless retrained

Source citations

Only if provided manually

Strong support through retrieved documents

Not natural unless combined with RAG

Setup complexity

Low

Medium to high

Medium to high

Maintenance

Prompt versioning and testing

Document refresh, index quality, retrieval evaluation

Dataset management, retraining, model versioning

Cost profile

Lowest to start

Moderate: LLM + embeddings + vector DB

Training cost + inference cost + evaluation

Security focus

Prompt injection and data leakage

Access control, document permissions, retrieval security

Training data privacy and model governance

Hallucination reduction

Limited

Good if retrieval is strong

Can improve behavior but does not guarantee factuality

Best beginner choice

Yes

Yes after basics

Later, after mastering prompting and RAG

 

15.11 Cost Comparison

Cost should be evaluated across development, runtime, maintenance, and failure risk. Prompting is cheap to start but may become costly if every request includes long examples. RAG adds embedding and vector database cost, but avoids repeated retraining. Fine-tuning can reduce prompt length and improve consistency, but requires training cost, data preparation, evaluation, and model lifecycle management.

Cost Type

Prompting

RAG

Fine-tuning

Initial development

Low

Medium

Medium to high

Runtime token cost

Can be low or high depending on prompt length

Prompt includes retrieved context, so moderate

Can be lower if shorter prompts are needed

Infrastructure

Minimal

Vector DB, ingestion pipeline, embedding jobs

Training environment, model registry, deployment

Data preparation

Low

Medium for cleaning/chunking/metadata

High for labelled examples

Maintenance

Prompt testing

Index refresh and retrieval monitoring

Retraining and model governance

 

Cost Advice

For beginners and small teams, avoid fine-tuning until you have measured prompting and RAG performance. Many business use cases are solved with prompt templates plus RAG.

 

15.12 Complexity Comparison

Prompting is the simplest approach technically, but it still needs discipline. RAG is more complex because it introduces an indexing pipeline and retrieval pipeline. Fine-tuning is complex because it introduces training data management, model evaluation, model versioning, deployment, rollback, and monitoring.

Complexity Area

Prompting

RAG

Fine-tuning

Architecture

Prompt template + LLM

Prompt + LLM + retriever + vector DB + ingestion

Training pipeline + model deployment + inference

Testing

Prompt tests and output checks

Retrieval tests plus answer quality tests

Model evaluation plus regression tests

Debugging

Inspect prompt and response

Inspect retrieved chunks and answer

Inspect training data, model output, and drift

Team skills

Prompt design and app development

Data engineering, embeddings, search, app development

ML engineering, data labelling, model operations

 

15.13 Data Requirement Comparison

Data requirements are very different. Prompting may require only examples inside the prompt. RAG requires documents and metadata. Fine-tuning requires carefully prepared input-output examples. Fine-tuning data must be consistent because the model learns from patterns. If labels are inconsistent, the fine-tuned model will also be inconsistent.

Question

Prompting

RAG

Fine-tuning

Do I need labelled data?

No, optional examples help

No labelled data required for basic RAG

Yes, high-quality examples are important

Do I need documents?

Only if included in prompt

Yes, documents are central

Optional, but documents must be converted into examples if used

Can I update data easily?

Update prompt manually

Update documents and re-index

Retrain or continue training

What data quality matters most?

Clear instructions and examples

Clean chunks, metadata, freshness, permissions

Consistent labels, outputs, style, and coverage

 

15.14 Maintenance Comparison

Maintenance is often ignored during prototypes. In production, it becomes critical. A prompt must be versioned. A RAG index must be refreshed. A fine-tuned model must be evaluated and possibly retrained when business rules change.

  • Prompting maintenance: Store prompts in version control, test changes, monitor output quality, and avoid accidental prompt drift.
  • RAG maintenance: Refresh documents, delete outdated chunks, validate metadata, monitor retrieval quality, and handle access permissions.
  • Fine-tuning maintenance: Maintain training datasets, track model versions, evaluate new versions, monitor drift, and define rollback strategy.

15.15 Security Comparison

Security requirements differ across the three approaches. Prompting risks include prompt injection and accidental leakage of secrets inside prompts. RAG risks include retrieving documents that the user is not allowed to see. Fine-tuning risks include training on sensitive data that may become difficult to remove later.

Security Area

Prompting

RAG

Fine-tuning

Main risk

Prompt injection, secret leakage

Unauthorized document retrieval

Sensitive data embedded in training data

Access control

Application-level controls

Document-level and chunk-level permissions

Training data governance

Data deletion

Remove from prompt templates

Delete documents/chunks and re-index

May require retraining or model replacement

Auditability

Log prompts and outputs carefully

Log retrieved sources and user permissions

Track dataset and model versions

Best control

Input/output validation

Permission-aware retrieval

Data review before training

 

Security Warning

Never put API keys, passwords, personal data, or confidential secrets directly into prompts or fine-tuning datasets. For RAG, always enforce permissions before retrieved content is sent to the LLM.

 

15.16 Hallucination Comparison

Hallucination means the model produces an answer that sounds confident but is false, unsupported, or invented. Prompting can reduce hallucination by instructing the model to say “I do not know,” but prompting alone cannot guarantee factuality. RAG reduces hallucination by grounding answers in retrieved documents. Fine-tuning can improve format and behavior, but it does not automatically make the model factual about changing business knowledge.

Approach

Hallucination Impact

Best Practice

Prompting

Can reduce but not eliminate hallucination

Tell the model to answer only from provided context and admit uncertainty

RAG

Strongly helps when retrieval is accurate

Show citations, validate sources, evaluate retrieval quality

Fine-tuning

Can improve task behavior but may still invent facts

Use for behavior; combine with RAG for factual knowledge

Combined

Best for production when designed carefully

Prompt rules + retrieved sources + evaluation + guardrails

 

15.17 Example Decision Tree

The following decision tree can be used by beginners, developers, and architects when choosing an approach.

Start
 |
 |-- Is the task mainly about format, tone, explanation, rewriting, summarization, or extraction?
 |       |
 |       +-- Yes --> Start with Prompting
 |       |
 |       +-- No --> Continue
 |
 |-- Does the answer require private, current, large, or changing knowledge?
 |       |
 |       +-- Yes --> Use RAG
 |       |
 |       +-- No --> Continue
 |
 |-- Do you have many high-quality examples of desired input/output behavior?
 |       |
 |       +-- Yes --> Consider Fine-tuning
 |       |
 |       +-- No --> Improve prompting, collect examples, or redesign workflow
 |
 |-- Do you need both private knowledge and consistent behavior?
 |       |
 |       +-- Yes --> Combine Prompting + RAG, then consider Fine-tuning if needed
 |
End

Decision Shortcut

If the problem is “the model does not know our documents,” choose RAG. If the problem is “the model does not follow our style or labels consistently,” consider fine-tuning. If the problem is “the instruction is unclear,” improve prompting.

 

15.18 Example 1: Customer Support Bot

A customer support bot answers product questions, troubleshoots issues, classifies tickets, and drafts replies. This is a common enterprise GenAI use case because it combines knowledge, tone, workflow, and classification.

Requirement

Recommended Approach

Reason

Answer product questions from manuals

RAG

Manuals are external knowledge and may change

Use polite support tone

Prompting

Tone can be controlled by instruction

Classify ticket category

Prompting first, fine-tuning later

Fine-tuning helps if there are many labelled tickets

Suggest troubleshooting steps

RAG + prompting

Retrieve relevant known issue and format response

Escalate high-risk complaints

Prompting + rules/guardrails

Business workflow needs deterministic policy controls

 

Example prompt for support bot with RAG:
You are a customer support assistant. Use only the retrieved product documentation.
If the answer is not in the documentation, say that escalation is required.
Return:
- Short answer
- Step-by-step solution
- Source section
- Escalation required: Yes/No

Customer issue: "My router shows red light after firmware update."
Retrieved context: <relevant manual chunks>

For this system, the recommended architecture is Prompting + RAG. Fine-tuning can be added later if the company has thousands of historical tickets and wants better category prediction or consistent response style.

15.19 Example 2: Legal Document Assistant

A legal document assistant helps lawyers or compliance teams search contracts, summarize clauses, compare terms, and draft notes. This type of system requires strong grounding, citations, and careful risk controls.

Requirement

Recommended Approach

Reason

Answer questions from contracts

RAG

Answers must come from specific legal documents

Cite clause references

RAG

Retrieved chunks can provide source references

Summarize contract risks

Prompting + RAG

Prompt defines risk categories; RAG supplies contract content

Write in firm-approved legal style

Fine-tuning optional

Useful if many examples of preferred style exist

Avoid legal hallucination

RAG + guardrails + human review

Legal answers must be verified by professionals

 

Legal Assistant Warning

A legal assistant should not be treated as an autonomous lawyer. It should provide citations, uncertainty handling, and human review workflow. RAG is usually more important than fine-tuning for legal knowledge.

 

15.20 Example 3: Brand-style Writing Assistant

A brand-style writing assistant creates social posts, product descriptions, email copy, and website content in a company-specific tone. The core challenge is not factual retrieval but consistent voice, phrasing, and structure.

Requirement

Recommended Approach

Reason

Follow brand tone

Prompting first, fine-tuning later

Prompt can define tone; fine-tuning improves consistency

Use product facts

RAG

Product facts should come from current product database/docs

Generate campaign copy

Prompting

Creative generation controlled by instructions

Match old approved examples

Fine-tuning

Many approved samples can teach style patterns

Avoid unsupported claims

RAG + guardrails

Claims should be grounded in approved product data

 

Example brand prompt:
You are a copywriter for a premium ceramic tile company.
Tone: elegant, trustworthy, simple, and aspirational.
Avoid exaggerated claims. Use only the approved product facts below.
Write a 120-word product description for homeowners.

Approved facts: <retrieved product facts>
Product: Marble-look glossy wall tile

For a brand-writing assistant, fine-tuning becomes attractive when the company has many approved writing samples and wants the model to repeatedly produce similar style with less prompt engineering.

15.21 Example 4: Classification Model

Classification is a task where the model assigns a label to input text. Examples include ticket routing, email priority, sentiment classification, fraud alert type, document type detection, or complaint category prediction.

Requirement

Recommended Approach

Reason

Classify a small number of examples

Prompting

Quick and simple

Classify thousands of tickets consistently

Fine-tuning or embedding similarity

Label consistency matters at scale

Explain why a label was selected

Prompting

Prompt can ask for explanation

Use policy rules while classifying

Prompting + RAG

Retrieve rules, then classify

Low-latency high-volume classification

Fine-tuned smaller model

May reduce cost and latency

 

Prompting-based classification:
Classify the ticket into one category:
Network, Hardware, Security, HR Payroll, Database, Application Support, Access Management.
Return JSON only:
{"category": "...", "confidence": 0.0-1.0, "reason": "..."}

Ticket: "I clicked a suspicious email link and entered my password."

Expected response:
{"category": "Security", "confidence": 0.93, "reason": "The ticket involves phishing and credential exposure."}

For beginners, start classification with prompting. If accuracy is not stable, collect labelled examples. Then compare three options: better prompt, embedding similarity with known examples, and fine-tuning. Fine-tuning should be based on measurement, not assumption.

15.22 Practical Recommendation for Beginners

A beginner should not jump directly to fine-tuning. The best path is incremental. Build the simplest version first, measure the weakness, then add the next layer.

1. Start with a clear prompt and a strong output format.

2. Test the prompt with 20 to 50 realistic examples.

3. If answers are wrong because the model lacks knowledge, add RAG.

4. If retrieval is weak, improve chunking, metadata, hybrid search, and re-ranking.

5. If output behavior is inconsistent despite good prompting and good context, collect examples.

6. Try fine-tuning only when you have enough high-quality examples and a measurable improvement target.

7. Always evaluate using test cases before going to production.

Beginner Project Path

Project 1: Prompt-based summarizer. Project 2: RAG-based document assistant. Project 3: Prompt-based ticket classifier. Project 4: Fine-tuning experiment only after collecting labelled examples.

 

15.23 Common Mistakes

Mistake

Why It Is a Problem

Better Approach

Fine-tuning to add documents

Model may still not cite or stay updated

Use RAG for documents

Using only prompting for huge knowledge bases

Context window and accuracy limitations

Use RAG

Building RAG without metadata

Poor filtering and weak retrieval

Store source, date, department, access level

Fine-tuning on messy labels

Model learns inconsistent behavior

Clean and standardize training data

Ignoring evaluation

No proof of improvement

Create test dataset and track metrics

No access control in RAG

Users may see unauthorized content

Apply permission-aware retrieval

Overengineering early prototype

Wastes time and money

Start with prompt, then add RAG/fine-tuning only when needed

 

15.24 Mini-Project: Choose the Right Approach

In this mini-project, you will practice selecting prompting, RAG, fine-tuning, or a combination for different business cases.

Scenario

Your Decision

Reason

Summarize one meeting transcript

Prompting

The full content is already available in the prompt

Answer questions from 500 HR policy PDFs

RAG

Large private document collection

Classify 100,000 support tickets into fixed labels

Fine-tuning after baseline prompting

Repeated labelled task at scale

Write product descriptions in company tone using current product specs

Prompting + RAG; fine-tuning optional

Needs style and current facts

Legal clause search with citations

RAG + human review

Needs source-grounded answers and verification

 

15.25 Decision Checklist

  • Can the task be solved with clear instructions and small context? Start with prompting.
  • Does the task require private, current, large, or changing knowledge? Use RAG.
  • Does the task require consistent repeated behavior across many examples? Consider fine-tuning.
  • Do you have enough clean labelled examples? If not, do not fine-tune yet.
  • Do users need source citations? Prefer RAG.
  • Does data change often? Prefer RAG over fine-tuning.
  • Do you need strict brand style or classification consistency? Fine-tuning may help.
  • Have you created evaluation test cases? Do this before production.
  • Have you checked cost, latency, security, and maintenance impact?
  • Can a simpler solution work? Avoid unnecessary complexity.

15.26 Chapter Summary

Prompting, RAG, and fine-tuning are three important approaches for building LLM applications. Prompting controls the instruction and output format. RAG provides external knowledge at runtime. Fine-tuning adapts the model behavior using examples. A strong AI developer or architect knows how to choose between them instead of using one approach for every problem.

For most beginners and many real enterprise projects, the best first path is prompting, then RAG if private or current knowledge is needed. Fine-tuning should be considered when there is enough high-quality training data and a clear repeated behavior problem. In production, these approaches are often combined with evaluation, monitoring, security, and guardrails.

Key Terms

Term

Meaning

Prompting

Giving instructions and context to an LLM for the current request.

Prompt template

A reusable prompt structure with variables such as user question or context.

RAG

Retrieval-Augmented Generation, where relevant external information is retrieved and given to the LLM.

Fine-tuning

Training a pre-trained model further on task-specific examples to adapt behavior.

Labelled data

Training or test examples with known correct outputs.

Hallucination

A confident but unsupported or false AI-generated answer.

Context grounding

Providing source material so the answer is based on evidence.

Model behavior

How the model responds, formats output, follows style, or performs a task.

Evaluation dataset

A set of test examples used to measure quality before and after changes.

Model versioning

Tracking different trained model versions and their performance.

 

Practice Exercises

1. Write a prompt for a meeting summarizer that returns decisions, risks, and action items.

2. Choose whether prompting, RAG, or fine-tuning is best for a chatbot that answers from school admission policy PDFs. Explain your decision.

3. Create a decision table for an IT ticket routing system. Mention which parts need prompting, RAG, and fine-tuning.

4. List five risks of fine-tuning on poor-quality data.

5. Design a RAG + prompting architecture for a company policy assistant.

6. Create 10 labelled examples for a support-ticket classification fine-tuning dataset.

7. Explain why RAG is usually better than fine-tuning for frequently changing product prices.

8. Compare the cost impact of a long prompt versus a fine-tuned model for a repeated classification task.

9. Write a security checklist for a RAG system that retrieves confidential documents.

10. Take any one real business process from your workplace and decide the best GenAI approach using the decision tree.

End of Chapter 15

Next: evaluation, monitoring, and guardrails for safer production AI systems.