Books → Learning Ai From Scratch → Introduction → Chapter 6


Chapter 6

Chapter 6

What is an LLM?

 

 

Chapter Promise

By the end of this chapter, you will understand what a Large Language Model is, how it is trained, how it generates answers, what parameters like temperature and top-p mean, why hallucinations happen, and how to use LLMs safely in real applications.

 

Chapter 6: What is an LLM?

Large Language Models are the foundation of modern Generative AI applications. Chatbots, document assistants, coding copilots, summarizers, question answering systems, agents, and many enterprise AI tools depend on LLMs. This chapter explains LLMs from scratch in simple language and connects the theory to real application design.

Learning Objectives

  • Understand the meaning of Large Language Model in simple terms.
  • Explain tokens, parameters, training data, pre-training, fine-tuning, instruction tuning, RLHF, and inference.
  • Understand generation settings such as temperature, top-p, and max tokens.
  • Understand system prompts, user prompts, and assistant responses.
  • Recognize hallucination, reasoning limitations, and safe usage practices.
  • Compare open-source, closed-source, local, and cloud LLM options.
  • Design the basic response flow of an LLM-powered chatbot.

Chapter Structure

1. Meaning of LLM

2. How LLMs are built and trained

3. How LLMs generate responses

4. Important inference settings

5. Prompt roles and chatbot message flow

6. Limitations and hallucinations

7. Popular model families and deployment choices

8. Best practices, exercises, and mini-project

 

6.1 What Does Large Language Model Mean?

An LLM, or Large Language Model, is an AI model trained on very large amounts of text and sometimes code, images, audio, or other data. Its main job is to understand language patterns and generate useful text. When you ask an LLM a question, it does not search its memory like a database. Instead, it predicts a useful sequence of words based on the input, the conversation context, and the patterns it learned during training.

A simple way to think about an LLM is this: it is like a very advanced language completion engine. If you write “The capital of India is”, a language model predicts that the next likely word is “New”, then “Delhi”. Modern LLMs do this at a very advanced level. They can answer questions, summarize documents, write code, translate text, classify tickets, draft emails, explain concepts, and help with planning.

Simple Analogy

Imagine a student who has read millions of books, websites, manuals, support tickets, programming examples, and articles. The student does not remember every page exactly, but has learned language patterns, facts, writing styles, problem-solving patterns, and how people ask and answer questions. An LLM behaves somewhat like this, but mathematically.

 

Word

Simple Meaning

Practical Interpretation

Large

Very big in scale

Trained using huge data, many parameters, and powerful computing resources.

Language

Human and programming language

Works with text, code, conversation, documents, instructions, and sometimes multimodal inputs.

Model

A learned mathematical system

A program with trained parameters that predicts and generates output from input.

 

6.2 What “Large” Means

The word “large” refers to scale. An LLM is large because it is trained with huge datasets, contains many learned internal values called parameters, and requires significant compute power to train and run. Large does not only mean the file size is big. It also means the model has learned many patterns from language and has enough capacity to perform many different tasks without being trained separately for every task.

In normal software, if you want a program to classify a complaint, summarize a document, translate text, and write an email, you usually write separate logic for each task. A large language model is different. The same model can perform many tasks because it has learned general language ability during training.

However, larger is not always better for every use case. A very large model may give better reasoning or writing quality, but it can also be slower and more expensive. In production systems, architects often choose the smallest model that gives acceptable quality for the task.

Model Size Factor

What It Affects

Example Impact

Number of parameters

Capacity to learn patterns

A larger model may understand complex instructions better.

Training data volume

Knowledge and language coverage

More diverse data can improve general usefulness.

Context window

Amount of input the model can consider at once

A larger context window can read longer documents or longer chat history.

Compute requirement

Speed and cost

Large models often need more GPU memory and higher API cost.

 

6.3 What “Language” Means

The word “language” includes human language such as English, Hindi, German, French, and many others. It also includes technical language like SQL, Python, JavaScript, JSON, XML, logs, error messages, configuration files, and API documentation. This is why LLMs are useful not only for writing essays but also for software development, data engineering, customer support, business analysis, and documentation.

When we say an LLM understands language, we should be careful. It does not understand exactly like a human. It learns statistical and semantic relationships between words, phrases, instructions, examples, and outputs. Still, in practice, this ability is powerful enough to solve many real tasks.

IT Example

Input: “Write a SQL query to find the top 10 customers by total sales.”

The LLM recognizes words like SQL, query, customers, sales, top 10, and total. It uses learned patterns to generate a likely SQL statement. It is not executing the database unless connected to a tool.

 

6.4 What “Model” Means

A model is a mathematical system that has learned from examples. In traditional programming, the developer writes rules. In machine learning, the model learns patterns from data. An LLM is a model trained to predict and generate language.

Inside the model are billions or sometimes trillions of numerical values called parameters. These values are adjusted during training. They are not normal database rows. They are internal weights that help the model decide what output is likely and useful.

Traditional Program

LLM

Developer writes explicit rules.

Model learns patterns from data.

Good for fixed logic.

Good for flexible language tasks.

Fails when input is unexpected unless coded.

Can handle many variations of natural language.

Output is deterministic if rules are deterministic.

Output can vary based on generation settings.

 

6.5 Tokens: The Basic Units of LLM Input and Output

LLMs do not directly process text exactly as humans see it. They break text into smaller units called tokens. A token can be a word, part of a word, punctuation, number, or symbol. For example, the sentence “AI is useful” may be split into tokens like “AI”, “ is”, and “ useful”. The exact split depends on the tokenizer used by the model.

Tokens are important because LLM cost, speed, and context limits are usually measured in tokens. If you send a long document, the model processes many input tokens. If the model generates a long answer, it produces many output tokens. In API-based systems, both input and output tokens may affect cost.

Token Example

Human text: “Customer refund request approved.”

Possible tokens: [Customer] [ refund] [ request] [ approved] [.]

The exact tokenization may differ by model, but the idea is that the model processes small text units rather than entire paragraphs as one object.

 

Concept

Meaning

Why It Matters

Input tokens

Tokens sent to the model

Long prompts and documents increase cost and latency.

Output tokens

Tokens generated by the model

Long responses also increase cost and response time.

Context window

Maximum tokens the model can consider at once

Large documents may need chunking or RAG.

Max tokens

Limit on generated response length

Controls answer size and cost.

 

6.6 Parameters: The Learned Knowledge Inside the Model

Parameters are internal numerical values learned during training. They help the model decide which token should come next and how words relate to one another. A parameter is not a sentence, document, or fact stored in a readable table. It is more like a tiny part of the model’s learned pattern system.

When a model is trained, it repeatedly compares its predictions with correct examples and adjusts parameters to reduce error. Over time, the model becomes better at predicting language patterns. Parameters are the reason the model can generalize from training data to new tasks.

Important Clarification

A model with many parameters does not mean it has a searchable memory of all training documents. It means the model has many learned numerical values that encode language patterns. This is why an LLM may know common facts but may also forget, confuse, or hallucinate details.

 

6.7 Training Data

Training data is the data used to teach the model. It may include books, articles, websites, documentation, code, conversations, question-answer pairs, and other text sources. For modern multimodal models, training may also include images, audio, and video-related data. The quality, diversity, and safety of training data strongly affect model behavior.

Training data teaches language style, grammar, factual patterns, reasoning examples, coding patterns, and domain vocabulary. But training data also creates risks. If the data contains outdated facts, bias, unsafe content, poor-quality text, or private information, the model may learn unwanted patterns. This is why data governance and safety filtering are important in LLM development.

Training Data Quality Issue

Possible Model Problem

Mitigation

Outdated data

Model gives old information

Use RAG, tools, or updated knowledge sources.

Biased data

Biased responses

Data filtering, evaluation, guardrails, human review.

Noisy data

Poor answer quality

Deduplication, cleaning, curated datasets.

Missing domain data

Weak domain-specific answers

RAG, fine-tuning, or domain instructions.

Private data leakage

Security and compliance risk

Data masking, access control, privacy review.

 

6.8 Pre-training: Learning General Language Patterns

Pre-training is the first major training stage for many LLMs. During pre-training, the model learns from massive amounts of text by predicting missing or next tokens. The goal is not to teach a single task. The goal is to teach general language understanding and generation ability.

A common simplified training objective is next-token prediction. The model sees a sequence of tokens and tries to predict the next one. For example, if the input is “The customer raised a support”, the likely next token may be “ticket”. The model predicts, checks the error, and updates internal parameters. This process happens billions or trillions of times during training.

Simple Pre-training Flow

1. Collect huge text dataset

2. Break text into tokens

3. Give part of text to model

4. Ask model to predict next token

5. Compare prediction with actual next token

6. Adjust parameters

7. Repeat at very large scale

 

Text: “The bank approved the home ___.”

Possible next tokens: loan, account, branch, customer

Most likely token in context: loan

The model learns that “bank approved the home loan” is a common and meaningful pattern.

6.9 Fine-tuning: Teaching the Model a More Specific Behavior

Fine-tuning is an additional training stage where a pre-trained model is trained on a smaller, more specific dataset. The goal is to adapt the model to a certain task, domain, style, or output format. For example, a company may fine-tune a model to classify customer support tickets into categories such as Network, Hardware, HR Payroll, Database, and Application Support.

Fine-tuning is useful when you need consistent behavior across many examples. However, fine-tuning is not always the first solution. For many business applications, prompting or RAG is easier, cheaper, and safer. Fine-tuning changes the model behavior but does not automatically give the model access to fresh private documents. For private knowledge, RAG is usually better.

Use Case

Fine-tuning Helpful?

Reason

Write in a company-specific tone

Yes

Many examples can teach style and format.

Answer from latest HR policy PDF

Usually no

RAG is better because the policy changes.

Classify tickets into fixed teams

Often yes

Labeled examples can improve consistency.

Summarize one uploaded document

No

Prompting or chunk summarization is enough.

 

6.10 Instruction Tuning: Making the Model Follow Instructions

A pre-trained model may be good at completing text, but it may not naturally behave like a helpful assistant. Instruction tuning trains the model on examples of instructions and desired responses. This helps the model follow commands such as “summarize this”, “explain in simple language”, “return JSON”, or “classify the following ticket”.

Instruction tuning is one reason modern LLMs feel conversational. Instead of only continuing text, they learn to respond to user requests. For example, when a user says “Explain cloud computing in one paragraph”, the instruction-tuned model understands that it should provide an explanation, not merely continue the sentence.

Instruction Example

Instruction: “Classify this support ticket into one category.”

Input: “I cannot connect to VPN after password reset.”

Desired Output: “Network / Access Support”

Instruction tuning teaches the model to follow such task formats.

 

6.11 RLHF: Learning from Human Feedback

RLHF stands for Reinforcement Learning from Human Feedback. It is a training approach where human preferences are used to improve model behavior. Humans may compare two model responses and choose which one is better. The model then learns to prefer responses that are more helpful, harmless, truthful, and aligned with user expectations.

RLHF does not make the model perfect. It improves response style and safety behavior, but the model can still be wrong. It can still hallucinate, misunderstand instructions, or fail on complex reasoning. In production systems, RLHF is only one part of safety. You still need guardrails, validation, monitoring, and human review for important workflows.

Before Alignment

After Instruction Tuning / RLHF

May complete text without answering clearly.

More likely to answer the user directly.

May produce unsafe or unhelpful content.

More likely to refuse unsafe requests and explain limitations.

May ignore format requirements.

More likely to follow structured instructions.

May be verbose or inconsistent.

More likely to be helpful and readable.

 

6.12 Inference: Using the Model After Training

Inference means using a trained model to generate output. When you type a question into a chatbot, the model is not being trained from scratch. It is running inference. The system sends your prompt and context to the model, the model calculates likely next tokens, and it returns a response.

In production AI applications, inference is where cost, latency, security, and reliability become important. Every request may consume tokens and compute resources. A slow model may create poor user experience. A model without guardrails may produce unsafe output. A model without logging may be difficult to debug.

Inference Flow

User input -> Application backend -> Prompt construction -> LLM inference -> Generated response -> Output validation -> User interface

 

6.13 Simple Example of Token Prediction

The core idea behind many LLMs is token prediction. The model receives a sequence of tokens and predicts what token should come next. It does this repeatedly until the response is complete.

Input tokens:  The customer wants to reset his

Model predicts: password

Now input becomes: The customer wants to reset his password

Model predicts: .

Final sentence: The customer wants to reset his password.

This example looks simple, but the same idea scales to long answers. The model predicts one token, then another, then another. At each step, it uses the previous tokens and the context to decide what is likely and useful.

6.14 Example: How an LLM Completes Text

Suppose you write the following prompt:

Prompt: “Write a short email to a customer explaining that their refund has been approved.”

The LLM may generate:

Response: “Dear Customer, your refund request has been approved. The amount will be credited to your original payment method within 5-7 business days. Thank you for your patience.”

The LLM creates this response because it has learned email patterns, customer service tone, refund-related vocabulary, and instruction-following behavior. It is not checking your company database unless the application connects it to a database or retrieval tool.

6.15 Temperature: Controlling Creativity

Temperature is a generation setting that controls how creative or random the model output should be. A low temperature makes the model more focused and predictable. A high temperature makes the model more varied and creative.

For business-critical tasks like classification, extraction, compliance summaries, or database query generation, lower temperature is usually better. For brainstorming, marketing ideas, story writing, or creative drafting, a higher temperature may be useful.

Temperature Range

Behavior

Good For

Risk

0 to 0.2

Very focused and consistent

Classification, extraction, structured answers

May sound rigid.

0.3 to 0.7

Balanced

General chat, explanation, summarization

Some variation.

0.8 and above

More creative and random

Brainstorming, creative writing

Higher chance of irrelevant or incorrect output.

 

6.16 Top-p: Controlling the Candidate Token Pool

Top-p, also called nucleus sampling, controls how many likely token options the model considers while generating. Instead of considering every possible next token, the model considers a group of tokens whose combined probability reaches a threshold. For example, top-p of 0.9 means the model considers the smallest set of likely tokens whose total probability is around 90 percent.

For beginners, the practical rule is simple: do not change too many generation settings at once. Start with default settings. If you need stable output, reduce creativity. If you need more variety, increase creativity carefully.

Temperature vs Top-p

Temperature changes how sharply or softly the model chooses between possible tokens. Top-p limits the pool of candidate tokens. Both affect randomness. In many projects, tune one setting at a time to avoid unpredictable behavior.

 

6.17 Max Tokens: Controlling Response Length

Max tokens controls the maximum length of the model response. If max tokens is too low, the answer may stop in the middle. If it is too high, the answer may become unnecessarily long and costly. In production, setting max tokens is an important cost-control technique.

For example, a ticket classification system may need only 20 output tokens because the answer is a category name and confidence. A report summarizer may need 800 output tokens. A book chapter generator may need thousands of tokens. The correct value depends on the task.

Use Case

Typical Output Need

Max Token Strategy

Ticket classification

Very short

Limit output to category, reason, confidence.

Data extraction

Short structured JSON

Limit based on expected JSON size.

Document summary

Medium

Set based on summary length requirement.

Detailed educational answer

Long

Allow more tokens but control structure.

 

6.18 System Prompt, User Prompt, and Assistant Response

Modern chat-based LLM applications often use message roles. These roles help the model understand instruction priority and conversation structure.

Role

Meaning

Example

System prompt

High-level instruction that defines behavior, rules, tone, and boundaries.

“You are a helpful banking assistant. Answer only from approved policy documents.”

User prompt

The actual request from the user.

“What documents are needed for a home loan?”

Assistant response

The model-generated answer.

“For a home loan, the required documents usually include...”

 

The system prompt is like the operating instruction for the AI assistant. It can define the assistant’s role, allowed sources, response format, safety rules, and refusal behavior. The user prompt is the user’s question or instruction. The assistant response is the generated answer.

Business Example

System prompt: “You are an HR policy assistant. Answer only using retrieved HR policy context. If the answer is not present, say you do not know.”

User prompt: “How many casual leaves do I get?”

Assistant response: “According to the retrieved policy, employees receive...”

 

6.19 Example of Chatbot Response Flow

A chatbot is not only an LLM. It is an application that uses an LLM as one component. The application may include authentication, conversation history, retrieval, guardrails, logging, and UI.

+------------------+        +---------------------+        +------------------+

| User Interface   | -----> | Application Backend | -----> | Prompt Builder   |

+------------------+        +---------------------+        +------------------+

                                  |                            |

                                  v                            v

                         +----------------+          +------------------+

                         | Optional RAG   | -------> | LLM Inference    |

                         | Retriever      |          |                  |

                         +----------------+          +------------------+

                                  |                            |

                                  v                            v

                         +----------------+          +------------------+

                         | Guardrails     | <------- | Raw LLM Output   |

                         +----------------+          +------------------+

                                  |

                                  v

                         Final answer returned to user

  1. User types a question in the web or mobile interface.
  2. The frontend sends the request to the backend API.
  3. The backend checks authentication and user permissions.
  4. The application builds a prompt using system instructions, user question, and optional conversation history.
  5. If RAG is used, the retriever fetches relevant context from documents or a vector database.
  6. The prompt and context are sent to the LLM for inference.
  7. The LLM generates an answer token by token.
  8. The application validates, filters, formats, logs, and returns the answer.

6.20 Hallucination: When an LLM Sounds Confident but Is Wrong

Hallucination means the model generates information that sounds correct but is false, unsupported, outdated, or invented. This is one of the most important limitations of LLMs. A hallucination can be dangerous in legal, medical, financial, compliance, or production IT contexts.

Hallucination happens because the model is generating likely text, not directly verifying truth. If the prompt asks for something missing from its knowledge or missing from the provided context, the model may still try to produce a plausible answer. Good application design should reduce this behavior.

Cause of Hallucination

Example

How to Reduce It

Missing information

Policy answer not present in documents

Instruct model to say “I do not know” when context is missing.

Outdated knowledge

Old product pricing or old law

Use retrieval, APIs, or current databases.

Ambiguous question

User asks “Is it approved?” without context

Ask clarifying questions or require reference ID.

Overly creative settings

High temperature creates imaginative answer

Use lower temperature for factual tasks.

Weak evaluation

Bad answers reach users

Use test sets, monitoring, and human review.

 

Production Rule

Never assume an LLM answer is correct just because it is fluent. For important applications, verify against trusted data, add source citations, validate output, and keep human review for high-risk actions.

 

6.21 Reasoning Limitations

LLMs can appear to reason because they can solve many language and logic tasks. However, they have limitations. They may make arithmetic errors, miss hidden assumptions, follow misleading prompts, overfit to examples, or produce inconsistent answers. They can also struggle when a problem requires precise multi-step calculation, external facts, or real-time data.

In business systems, the safe approach is to let the LLM handle language tasks and let deterministic tools handle exact tasks. For example, an LLM can explain an invoice, but a billing engine should calculate the amount. An LLM can write a SQL query, but the database should execute it. An LLM can summarize a risk report, but business rules should validate the decision.

Task Type

LLM Suitable?

Recommended Approach

Explaining a concept

Yes

Use LLM directly with good prompt.

Exact tax calculation

Not alone

Use rules engine or calculator tool.

Answering from company policy

Yes with RAG

Retrieve policy text and cite source.

Approving a loan

Not alone

Use ML/rules/human approval; LLM may assist explanation.

Writing draft email

Yes

Use LLM, with user review.

 

6.22 Popular LLM Families

There are many LLM families. Some are closed-source or API-based, while others are open-weight or open-source style models that can be run locally or through cloud providers. The exact model names and versions change frequently, so treat the following as examples of families rather than a permanent list.

Model Family

Common Provider / Origin

Typical Use

GPT family

OpenAI

General chat, coding, reasoning, enterprise assistants, multimodal applications.

Claude family

Anthropic

Long-form writing, analysis, coding, business assistants, safety-focused applications.

Gemini family

Google

Multimodal applications, Google ecosystem integration, long-context tasks.

Llama family

Meta

Open-weight deployments, research, customization, local or cloud-hosted solutions.

Mistral family

Mistral AI

Efficient open and commercial models, enterprise deployments.

Qwen family

Alibaba Cloud

Multilingual and code-capable model deployments.

DeepSeek family

DeepSeek

Reasoning, coding, and open-model experimentation.

 

When choosing a model, do not select only by popularity. Select based on task quality, cost, latency, security, data policy, supported context length, deployment option, and integration simplicity.

6.23 Open-Source vs Closed-Source LLMs

The terms open-source, open-weight, and closed-source are often used in LLM discussions. A closed-source model is usually accessed through an API, and the provider controls the model weights and infrastructure. An open-weight model makes trained model weights available under a license, so organizations can run or customize it under certain conditions. True open-source may include more transparency around code, data, and training, but model licensing varies widely.

Factor

Closed-Source / API Model

Open-Weight / Local Model

Ease of use

Usually easier to start with API.

Requires setup, hardware, hosting, and operations.

Performance

Often strong on general tasks.

Can be strong, depends on model and hardware.

Control

Less control over internals.

More control over hosting and customization.

Data privacy

Depends on provider policy and configuration.

Can keep data inside own environment if deployed privately.

Cost

Pay per usage.

Infrastructure cost, GPU cost, operations cost.

Maintenance

Provider manages upgrades.

Your team manages deployment, scaling, monitoring.

 

Beginner Recommendation

Start with an API-based model or a free hosted notebook for learning concepts. Later, experiment with open-weight models locally using small models. For production, choose based on security, cost, quality, latency, and operational capability.

 

6.24 Local LLM vs Cloud LLM

A local LLM runs on your own computer or server. A cloud LLM runs through a cloud provider or model provider API. Both approaches are useful, but they solve different problems.

Requirement

Local LLM Better When...

Cloud LLM Better When...

Privacy

Data must stay fully inside your environment.

Provider policy and compliance are acceptable.

Setup speed

You already have hardware and skills.

You want to start quickly.

Model quality

A local model meets your quality target.

You need frontier quality or advanced features.

Cost pattern

High fixed usage justifies infrastructure.

Usage is variable or small.

Operations

You can manage GPUs, scaling, updates.

You prefer managed infrastructure.

 

For a student, cloud APIs and free learning tools are usually easier. For an enterprise with strict data rules, local or private cloud deployment may be preferred. Many companies use a hybrid approach: cloud LLMs for some tasks and private models for sensitive tasks.

6.25 LLM Limitations

LLMs are powerful, but they are not magical. A professional AI developer must understand their limits before building applications.

Limitation

Explanation

Example

Hallucination

May generate false but confident answers.

Invents a policy rule not found in HR document.

No automatic live knowledge

Model may not know current events unless connected to tools.

Gives old product price.

Context limit

Can only process limited tokens at once.

Cannot read a 500-page manual directly in one prompt.

Weak exact calculation

May make arithmetic or logic mistakes.

Incorrect EMI calculation.

Prompt sensitivity

Small prompt changes can change output.

Same task produces different format.

Security risk

May be affected by prompt injection.

Malicious document says “ignore previous instructions”.

Bias and tone issues

May reflect biased patterns from data.

Unfair or inappropriate language.

 

6.26 Best Practices for Using LLMs

Good LLM applications are designed carefully. The model is only one part of the system. The surrounding architecture controls context, grounding, safety, cost, and quality.

  • Define the exact task before choosing a model. Do not use an LLM just because it is popular.
  • Use clear system prompts and structured user prompts.
  • For private or changing knowledge, use RAG instead of relying on model memory.
  • Ask the model to say “I do not know” when context is insufficient.
  • Use low temperature for factual, classification, extraction, and compliance tasks.
  • Validate structured output using schemas before saving it to a database.
  • Use deterministic tools for calculation, database updates, payments, and approvals.
  • Log prompts, responses, latency, token usage, and errors for debugging.
  • Add guardrails for unsafe content, PII, prompt injection, and policy violations.
  • Create a test dataset and evaluate model quality before production release.
  • Keep a human-in-the-loop for high-risk decisions.
  • Monitor cost and set token limits, rate limits, and caching strategies.

6.27 Practical Example: Customer Support Assistant

Consider a company that receives customer support questions. The company wants an AI assistant to answer common questions and escalate complex issues.

Component

Role in the System

User interface

Customer types a question.

Backend API

Receives question and checks user session.

Prompt builder

Adds system rules and company tone.

Retriever

Finds relevant FAQ or policy documents.

LLM

Generates the answer using retrieved context.

Guardrails

Checks for unsafe or unsupported answer.

Ticket system

Creates ticket if confidence is low.

Monitoring

Tracks answer quality, latency, and cost.

 

System prompt:

You are a customer support assistant. Answer only from the provided FAQ context. If the answer is not present, ask the user to create a support ticket.

 

User question:

How long does refund take?

 

Retrieved context:

Refunds are normally credited within 5-7 business days after approval.

 

Assistant response:

After approval, refunds are normally credited within 5-7 business days. If your refund was approved more than 7 business days ago, please share your order ID so support can check it.

6.28 Practical Example: Banking Document Assistant

A bank wants an internal assistant that answers questions from policy documents. The LLM alone is not enough because policies change and the model may not know internal bank rules. So the bank uses RAG.

  1. Policy PDFs are uploaded to a secure document store.
  2. The system extracts text, splits it into chunks, and creates embeddings.
  3. Embeddings are stored in a vector database with metadata such as policy name, version, and department.
  4. An employee asks a question, such as “What is the maximum reimbursement limit for travel?”
  5. The retriever finds relevant policy chunks.
  6. The LLM answers only using those chunks and cites the policy section.
  7. If the answer is not present, the assistant says it cannot find the rule and suggests contacting the policy owner.

Why LLM + RAG Is Important Here

The LLM provides language ability. RAG provides trusted company knowledge. The application architecture provides security, access control, logging, and governance. All three are required for a safe enterprise assistant.

 

6.29 Mini Python-Style Pseudo-code: Calling an LLM

The following pseudo-code shows the basic idea. It is not tied to a specific provider. The real code changes depending on the SDK you use.

system_message = "You are a helpful AI tutor. Explain concepts simply."

user_message = "What is an LLM? Explain with one example."

 

response = llm.generate(

    system=system_message,

    user=user_message,

    temperature=0.3,

    max_tokens=300

)

 

print(response.text)

In a production application, you would also add error handling, logging, retry logic, rate limit handling, output validation, and security checks.

6.30 Mini Project: Build a Simple LLM Study Assistant

This mini project helps you understand LLM usage practically. The goal is to create a simple study assistant that explains AI terms in beginner-friendly language.

Step

Task

Expected Output

1

Choose 10 AI terms such as token, prompt, embedding, RAG, hallucination.

A list of terms.

2

Create a prompt template: “Explain {term} in simple English with one example.”

Reusable prompt.

3

Call an LLM or use a chatbot manually for each term.

Explanations generated.

4

Ask the model to return output in a fixed structure.

Definition, example, common mistake.

5

Review the answers manually and correct mistakes.

Validated learning notes.

6

Store results in a document or CSV.

Your first AI glossary dataset.

 

6.31 Common Beginner Mistakes

  • Thinking the LLM is a database of guaranteed facts.
  • Using high temperature for factual business tasks.
  • Sending very long documents directly instead of using chunking or RAG.
  • Not validating model output before using it in software workflows.
  • Ignoring token cost and response latency.
  • Using the same prompt for all tasks without testing.
  • Expecting fine-tuning to solve missing private knowledge.
  • Not logging prompts and responses during testing.
  • Allowing the model to make final high-risk decisions without human review.
  • Not testing prompt injection or unsafe inputs.

6.32 Chapter Summary

A Large Language Model is a trained AI model that processes and generates language using learned patterns. It works with tokens, uses parameters learned during training, and generates responses during inference. Pre-training teaches general language ability. Fine-tuning adapts the model to specific behavior. Instruction tuning and RLHF help the model follow instructions and behave more helpfully.

Important generation settings include temperature, top-p, and max tokens. System prompts, user prompts, and assistant responses form the structure of chat-based LLM applications. LLMs are powerful but can hallucinate, make reasoning mistakes, and produce unsafe output if not properly controlled. In production, an LLM must be combined with application logic, retrieval, validation, security, monitoring, and human review.

6.33 Key Terms

Term

Meaning

LLM

Large Language Model; an AI model trained to process and generate language.

Token

A small unit of text processed by the model.

Parameter

A learned numerical value inside the model.

Training data

Data used to train the model.

Pre-training

Large-scale training to learn general language patterns.

Fine-tuning

Additional training to adapt a model for a task or style.

Instruction tuning

Training that helps the model follow instructions.

RLHF

Reinforcement Learning from Human Feedback.

Inference

Using a trained model to generate output.

Temperature

Setting that controls randomness or creativity.

Top-p

Setting that controls the candidate token pool.

Max tokens

Limit on generated response length.

System prompt

High-priority instruction defining assistant behavior.

Hallucination

False or unsupported answer generated by the model.

Open-weight model

Model whose trained weights are available under a license.

Cloud LLM

Model accessed through managed cloud/API service.

Local LLM

Model hosted on your own computer or server.

 

6.34 Practice Exercises

Complete the following exercises before moving to the next chapter.

  1. Explain LLM to a non-technical person in five sentences.
  2. Write one example each of system prompt, user prompt, and assistant response for an HR policy assistant.
  3. Create a table comparing temperature 0.1, 0.5, and 0.9 for the same prompt.
  4. Take a customer support ticket and ask an LLM to classify it. Then manually check whether the classification is correct.
  5. Write a prompt that instructs an LLM to say “I do not know” when the answer is not present in the provided context.
  6. List three cases where an LLM should not be allowed to make the final decision.
  7. Design a simple chatbot flow for a college admission assistant.
  8. Research one open-weight model and one cloud API model. Compare them on cost, ease of use, and privacy.

6.35 Checklist: Am I Ready for the Next Chapter?

  • I can explain what an LLM is in simple language.
  • I understand tokens, parameters, training data, pre-training, fine-tuning, instruction tuning, RLHF, and inference.
  • I know what temperature, top-p, and max tokens do.
  • I understand the difference between system prompt, user prompt, and assistant response.
  • I understand why hallucinations happen and how to reduce them.
  • I can compare open-source/open-weight and closed-source/API model options.
  • I can describe a basic chatbot response flow.
  • I know that production LLM systems need validation, monitoring, security, and human review.

6.36 Suggested Free Practice Strategy

Use free or low-cost tools to practice LLM concepts. Start with a chatbot interface to understand prompting. Then move to notebooks or simple Python scripts. Practice with small examples before building large applications.

Practice Area

Free / Low-Cost Method

Prompting

Use any free chatbot interface and test prompt variations.

Token awareness

Use tokenizer visualizers or estimate token usage manually.

API basics

Use trial credits or local mock functions before paid calls.

Local models

Experiment with small open-weight models if your machine supports it.

Evaluation

Create a spreadsheet of questions, expected answers, and actual model answers.

RAG preparation

Collect 3-5 PDFs and plan how they would be chunked and searched.

 

End of Chapter 6

In the next chapter, we will explain Transformer Architecture in simple terms. Transformers are the model architecture behind many modern LLMs. You will learn attention, self-attention, encoder-only models, decoder-only models, and why transformers became so important for language AI.