Books → Learning Ai From Scratch → Introduction → Chapter 10


Chapter 10

Chapter 10

Semantic Search

Meaning-based search using embeddings, vector similarity, metadata filters, hybrid retrieval, re-ranking, and quality evaluation

Learning Objectives

  • Understand what semantic search is and why it is important in modern AI applications.
  • Compare semantic search with keyword search using practical business examples.
  • Understand query embeddings, document embeddings, similarity search, Top-K retrieval, metadata filtering, hybrid search, and re-ranking.
  • Design a simple semantic search flow for documents, tickets, and learning content.
  • Evaluate semantic search quality using relevance, precision, recall, coverage, latency, and user feedback.
  • Identify common challenges and avoid production mistakes.

10.1 Introduction: Why Search Needs Meaning

Search is one of the most common features in software. We search emails, documents, support tickets, product catalogs, company policies, reports, code repositories, and learning material. Traditional search is usually based on keywords. It looks for words that appear in the user query and tries to find those same words inside the stored content.

Keyword search is useful, but it fails when the user and the document use different words for the same idea. An employee may search for “vacation rules”, while the HR document may use “annual leave entitlement”. A customer may search for “money deducted but order not placed”, while the support ticket says “payment captured but transaction failed”. A student may search for “how machines understand language”, while the course lesson is titled “natural language processing”.

Semantic search solves this problem by searching based on meaning. It does not depend only on exact word matching. It tries to understand the intent behind the query and find information that is semantically related. This makes it extremely useful for Generative AI, RAG systems, knowledge assistants, document Q&A, recommendation engines, duplicate ticket detection, fraud investigation, and enterprise search.

10.2 What is Semantic Search?

Semantic search is a search technique that finds results based on meaning rather than only exact words. The word “semantic” means meaning. Therefore, semantic search can be understood as meaning-based search.

In semantic search, text is converted into numerical vectors called embeddings. A vector is a list of numbers. These numbers represent the meaning of the text in a mathematical form. If two sentences have similar meaning, their vectors are close to each other. If two sentences have different meaning, their vectors are far apart.

Simple idea

User Query:       "How many days of vacation can I take?"
Document A:       "Employees are eligible for 21 days of annual leave."
Document B:       "Office laptops must be returned during exit clearance."

Semantic Search Result:
Document A is selected because "vacation" and "annual leave" have related meaning.

A semantic search system usually has two stages. First, documents are processed and stored as embeddings. Second, when the user enters a query, that query is also converted into an embedding. The system then compares the query embedding with document embeddings and returns the closest matches.

Term

Simple Meaning

Example

Query

The text entered by the user

“leave policy for new employee”

Document

The searchable content stored in the system

HR policy PDF, ticket, article, course lesson

Embedding

Numerical representation of meaning

[0.18, -0.42, 0.77, ...]

Similarity search

Finding vectors that are closest to the query vector

Return top 5 most related policy sections

Top-K

Number of best results to return

Top 3 or Top 10 documents

 

10.3 Keyword Search vs Semantic Search

Keyword search checks whether query words appear in the document. Semantic search checks whether the meaning of the query is close to the meaning of the document. Both are useful, but they solve different problems.

Aspect

Keyword Search

Semantic Search

Main idea

Searches exact or similar words

Searches meaning and intent

Example query

“vacation policy”

“How many paid leaves do I get?”

Can find synonyms?

Limited unless configured manually

Yes, if the embedding model understands the relation

Works well for

Names, IDs, codes, exact terms, legal references

Natural language questions, similar meaning, user intent

Weakness

Misses results when words differ

May return related but not exact results

Typical technology

SQL LIKE, full-text index, BM25, Elasticsearch keyword search

Embeddings, vector database, similarity search

Best enterprise use

Finding invoice number, customer ID, exact product code

Finding policy answers, similar tickets, related documents

 

For example, if a user searches for “PF withdrawal rules”, keyword search will perform well if the document also contains “PF withdrawal”. But if the document says “Provident Fund settlement process”, keyword search may not find it. Semantic search can understand that PF, Provident Fund, withdrawal, and settlement are related concepts.

10.4 How Semantic Search Understands Meaning

Semantic search does not understand meaning the way humans do. It uses an embedding model trained on large amounts of text. During training, the model learns patterns between words, phrases, and contexts. It learns that words like “car”, “vehicle”, and “automobile” are related. It also learns that “refund”, “payment reversal”, and “money returned” can be related in customer support situations.

When the embedding model receives text, it converts it into a vector. This vector works like a mathematical fingerprint of meaning. A search system can compare these fingerprints to find similar ideas.

Meaning comparison in vector space

Text 1: "annual leave policy"       -> Vector A
Text 2: "vacation rules"            -> Vector B
Text 3: "laptop replacement process"-> Vector C

Vector A is close to Vector B.
Vector A is far from Vector C.

Therefore, Text 1 and Text 2 are likely about the same topic.

This is why semantic search is powerful for user-friendly search. A user does not need to know the exact wording used by the company, school, bank, hospital, or product team. They can ask naturally, and the search engine can still find useful results.

10.5 Query Embedding

A query embedding is the vector representation of the user query. When a user types a search question, the system sends that text to an embedding model. The model returns a vector that represents the meaning of the query.

Input: User query text, such as “How can I apply for maternity leave?”

Processing: The embedding model converts the query into a vector.

Output: A query vector, such as [0.12, -0.08, 0.44, ...].

The query embedding is temporary. It is created at search time and used to find similar document embeddings already stored in the vector database.

# Pseudo-code: query embedding
query = "How can I apply for maternity leave?"
query_vector = embedding_model.embed(query)
results = vector_database.search(query_vector, top_k=5)

10.6 Document Embedding

A document embedding is the vector representation of a document or part of a document. In most real systems, large documents are not embedded as one giant block. They are split into smaller pieces called chunks. Each chunk is converted into an embedding and stored in a vector database with useful metadata.

Input: A document chunk, such as one section of an HR policy.

Processing: The embedding model converts the chunk into a vector.

Output: A document vector stored with document ID, title, page number, department, date, and source link.

Document embeddings are usually created before the user searches. This process is called indexing. The indexing pipeline may run once, daily, hourly, or whenever new documents are uploaded.

Document embedding pipeline

Company PDF / DOCX / HTML / CSV
        |
        v
Document Loader
        |
        v
Text Extraction and Cleaning
        |
        v
Chunking
        |
        v
Embedding Model
        |
        v
Vector Database + Metadata

10.7 Similarity Search

Similarity search means finding stored vectors that are closest to the query vector. The search engine compares the query vector with many document vectors and calculates how similar they are. The most similar vectors are returned as search results.

A common similarity measurement is cosine similarity. You do not need deep mathematics at this stage. The simple idea is that cosine similarity checks whether two vectors point in a similar direction. If the direction is similar, the meaning is considered similar. A score closer to 1 usually means more similar, while a lower score means less similar.

Query

Document Chunk

Similarity Score

Interpretation

How many annual leaves do I get?

Employees receive 21 days of annual leave per year.

0.91

Very relevant

How many annual leaves do I get?

Employees must submit laptop serial number during onboarding.

0.22

Not relevant

How many annual leaves do I get?

Leave encashment is calculated during final settlement.

0.72

Related, but may not answer directly

 

Similarity search is the heart of semantic search. However, similarity alone is not always enough. In production systems, we also use metadata filters, hybrid search, and re-ranking to improve accuracy.

10.8 Top-K Results

Top-K means the number of best matching results returned by the search system. If K is 5, the system returns the top 5 most similar chunks. If K is 10, it returns the top 10.

Choosing Top-K is important. If K is too small, the system may miss useful information. If K is too large, the system may include irrelevant information, increase cost, and confuse the LLM in a RAG system.

Top-K Value

When Useful

Risk

Top 1

Very precise FAQ search where one answer is expected

May miss the correct answer

Top 3

Short policy questions or simple support articles

May not include enough context

Top 5

Common default for document Q&A

Usually balanced

Top 10+

Complex research questions or broad topics

May add noisy results and increase token cost

 

In RAG applications, Top-K controls how many retrieved chunks are passed to the LLM as context. More chunks are not always better. Good retrieval means passing the right context, not the maximum context.

10.9 Filtering with Metadata

Metadata is additional information stored with each document chunk. It helps the search system filter results before or after similarity search. Metadata can include department, document type, language, region, effective date, security level, product name, customer type, and source file.

Metadata Field

Example Value

How It Helps

department

HR

Search only HR policies

country

India

Return policies applicable to India

document_type

Policy

Exclude old emails or informal notes

effective_date

2026-04-01

Prefer latest valid policy

security_level

Internal

Prevent unauthorized access

source_url

hr_policy_2026.pdf page 4

Show citation or source

 

Metadata filtering is very important in enterprise AI. Without metadata filters, a user may receive information from the wrong country, old policy version, wrong department, or unauthorized document.

# Pseudo-code: semantic search with metadata filter
query_vector = embedding_model.embed("What is the leave policy for India employees?")
results = vector_database.search(
    vector=query_vector,
    top_k=5,
    filter={"department": "HR", "country": "India", "status": "active"}
)

10.10 Hybrid Search

Hybrid search combines keyword search and semantic search. This is useful because keyword search and semantic search have different strengths. Keyword search is excellent for exact terms like invoice numbers, product codes, names, error codes, and legal clause numbers. Semantic search is excellent for meaning-based natural language questions.

A hybrid system may run both searches and combine the scores. For example, if a user searches “Oracle error ORA-12514 listener issue”, keyword search can strongly match ORA-12514, while semantic search can find articles about Oracle listener configuration even if the exact phrase is missing.

Hybrid search flow

User Query
   |
   +--> Keyword Search Engine ----+
   |                              |
   +--> Semantic Vector Search ---+--> Score Fusion / Merge
                                  |
                                  v
                            Ranked Results

 

 

Search Type

Best For

Example

Keyword search

Exact IDs, codes, names, product numbers

INV-2026-00981, ORA-12514, GSTIN

Semantic search

Natural language meaning

“payment failed but money deducted”

Hybrid search

Mixed business queries

“SBI PO eligibility age relaxation rules”

 

Many production AI systems use hybrid search because users rarely search in only one style. Sometimes they use exact terms. Sometimes they use vague natural language. Hybrid search gives better coverage.

10.11 Re-ranking

Re-ranking is the process of taking initial search results and sorting them again using a stronger relevance model or additional business logic. The first retrieval stage is usually optimized for speed. It may return 20 or 50 candidate results. The re-ranker then evaluates those candidates more carefully and selects the best results.

A simple vector search may retrieve chunks that are broadly related but not the best answer. A re-ranker can compare the query and each candidate chunk in more detail. This often improves answer quality in RAG systems.

Re-ranking flow

User Query
   |
   v
Vector Search returns Top 30 candidates quickly
   |
   v
Re-ranker checks query-document relevance more deeply
   |
   v
Final Top 5 results sent to user or LLM

Stage

Goal

Typical Output

Initial retrieval

Fast broad search

Top 20 or Top 50 candidate chunks

Re-ranking

Improve relevance order

Best 3 to 5 chunks

Final response

Use best evidence

Answer with citations or ranked search results

 

Re-ranking increases quality but may also increase cost and latency. It should be used when search accuracy is important, such as legal documents, medical knowledge bases, banking policies, and enterprise assistants.

 

 

10.12 Semantic Search Architecture Diagram

Complete semantic search architecture

                         INDEXING PIPELINE

 Documents / Tickets / Web Pages / PDFs / CSVs
                  |
                  v
          Document Loader
                  |
                  v
      Text Extraction and Cleaning
                  |
                  v
              Chunking
                  |
                  v
          Embedding Model
                  |
                  v
        Vector Database / Index
        + Metadata + Source Info

                         SEARCH PIPELINE

 User Search Query
        |
        v
 Query Preprocessing
        |
        v
 Query Embedding Model
        |
        v
 Vector Similarity Search + Metadata Filters
        |
        v
 Optional Hybrid Search Merge
        |
        v
 Optional Re-ranking
        |
        v
 Top-K Results
        |
        v
 User Search Results OR RAG Context for LLM

The architecture has two main parts: the indexing pipeline and the search pipeline. The indexing pipeline prepares the knowledge base before the user searches. The search pipeline runs whenever the user asks a question or enters a search query.

10.13 Step-by-Step Semantic Search Flow

  1. Collect documents from sources such as PDFs, DOCX files, HTML pages, ticket systems, databases, and learning platforms.
  2. Extract readable text from each source. For scanned files, use OCR to convert images into text.
  3. Clean the text by removing unnecessary headers, footers, repeated menus, broken characters, and irrelevant noise.
  4. Split long content into smaller chunks. Each chunk should be large enough to contain meaning but small enough to retrieve precisely.
  5. Generate embeddings for each chunk using an embedding model.
  6. Store chunk text, vector embedding, metadata, and source reference in a vector database.
  7. When a user searches, generate an embedding for the user query.
  8. Search the vector database for chunks with vectors closest to the query vector.
  9. Apply metadata filters such as department, region, date, user permission, or document type.
  10. Optionally combine results with keyword search using hybrid search.
  11. Optionally re-rank results for better relevance.
  12. Return the Top-K results to the user, or pass them as context to an LLM in a RAG application.

10.14 Example 1: Searching Company Leave Policy

Imagine a company has many HR documents: leave policy, maternity policy, travel policy, reimbursement policy, work-from-home policy, and exit policy. Employees do not remember exact document names. They ask natural questions.

User Query

Document Text

Why Semantic Search Helps

Can I take leave during probation?

Employees under probation are eligible for casual leave subject to manager approval.

Query uses “take leave”; document uses “eligible for casual leave”.

How many vacation days do I get?

Employees receive 21 days of annual leave per calendar year.

Query uses “vacation”; document uses “annual leave”.

Can I carry forward unused leaves?

A maximum of 10 unused annual leave days can be carried forward.

Query and document use related but not identical phrasing.

 

A good semantic search system should also use metadata filters. For example, if the user is in India, the system should prefer India HR policies. If the user is a contractor, it should not return full-time employee benefits unless clearly applicable.

Leave policy search flow

Employee Query: "Can I carry forward my vacation balance?"
        |
        v
Query Embedding
        |
        v
Vector Search in HR Policy Chunks
        |
        v
Filter: country=India, status=active, audience=employee
        |
        v
Top Results:
1. Annual Leave Carry Forward Section
2. Leave Encashment Section
3. HR FAQ on Leave Balance

10.15 Example 2: Searching Customer Complaints

Customer support teams receive thousands of complaints. Many complaints are similar but written in different words. Semantic search can find related complaints, similar past resolutions, duplicate tickets, and knowledge base articles.

Customer Complaint

Related Stored Ticket

Semantic Relation

Money debited but order failed

Payment captured but transaction unsuccessful

Same issue with different words

App keeps closing after login

Application crashes after authentication

Same technical meaning

Delivery boy did not come but status says delivered

False delivery confirmation reported by customer

Related operational complaint

 

Support agents can use semantic search to quickly find previous cases and suggested resolutions. Managers can use it to group complaints by theme, identify recurring issues, and improve product quality.

Customer complaint semantic search

New Ticket: "Amount deducted but booking not confirmed"
        |
        v
Semantic Search over Past Tickets
        |
        v
Top Similar Tickets:
- Payment captured but booking failed
- Transaction success at bank but failure in booking service
- Refund pending for failed booking
        |
        v
Suggested Routing: Payments Support Team
Suggested Action: Check payment gateway status and refund workflow

10.16 Example 3: Searching Learning Content

In an education platform, students may not know exact technical terminology. A beginner may search “how AI remembers words”, while the course content may contain “embeddings and vector representation”. Semantic search helps students discover relevant lessons even when their vocabulary is incomplete.

Student Search

Relevant Lesson

Why It Matches

How AI remembers meaning of words

Embeddings and Semantic Meaning

The intent is about word meaning representation.

How chatbot finds answers from PDF

RAG Architecture and Document Retrieval

The intent is document-based question answering.

How to compare two sentences

Cosine Similarity and Vector Search

The intent is sentence similarity.

 

Semantic search also powers recommendation. If a student reads a lesson about embeddings, the platform can recommend semantic search, vector databases, and RAG architecture because these topics are conceptually related.

10.17 Search Quality Evaluation

A semantic search system should not be judged only by whether it returns something. It should be judged by whether it returns the right results, in the right order, quickly, safely, and consistently. Search quality evaluation helps teams improve retrieval before building bigger AI applications on top of it.

Metric

Meaning

Simple Example

Relevance

Are the returned results useful for the query?

Leave policy query returns leave policy section.

Precision

How many returned results are actually relevant?

Out of 5 results, 4 are useful.

Recall

Did the system find all important relevant results?

It found both annual leave and carry-forward rules.

Ranking quality

Are the best results at the top?

The exact answer appears as result 1, not result 8.

Coverage

Can the system search all important content sources?

Policies, FAQs, tickets, and manuals are included.

Latency

How fast are results returned?

Search completes in less than 1 second.

Freshness

Are latest documents preferred?

2026 policy appears before 2023 policy.

Permission safety

Does the system hide unauthorized content?

Employee cannot see confidential HR notes.

 

A practical evaluation approach is to create a golden dataset. A golden dataset contains sample queries and expected correct results. For example, “How many days of annual leave do I get?” should return the annual leave entitlement section. Whenever the search system changes, test it again against the golden dataset.

# Example golden dataset row
query: "Can unused leave be carried forward?"
expected_document: "HR Leave Policy 2026"
expected_section: "Leave Carry Forward"
minimum_similarity_score: 0.75
must_not_return: "Travel Reimbursement Policy"

 

 

10.18 Common Challenges in Semantic Search

Challenge

What Happens

How to Reduce the Problem

Poor chunking

Search returns incomplete or confusing text.

Use meaningful section-based chunks and overlap where needed.

Old documents

System returns outdated policies.

Use metadata such as version, status, and effective date.

Wrong Top-K

Too few or too many results are retrieved.

Tune Top-K using evaluation queries.

No permission filtering

Users may see restricted content.

Apply access control filters before returning results.

Weak embedding model

Similar meanings are not matched well.

Choose an embedding model suitable for your language and domain.

Domain vocabulary

Model may not understand internal terms.

Add glossary, metadata, hybrid search, or domain-specific fine-tuning when needed.

Duplicate content

Same result appears multiple times.

Deduplicate documents and merge near-identical chunks.

No re-ranking

Useful results appear lower in the list.

Use re-ranking for high-value use cases.

No evaluation

Team cannot measure improvement.

Create test queries and expected results.

 

Semantic search is powerful, but it is not magic. Most failures happen because of poor data preparation, weak metadata, bad chunking, missing security filters, or lack of evaluation. A good AI engineer treats semantic search as an engineering system, not only as a model call.

10.19 Practical Design Example: Semantic Search API

A simple semantic search application usually exposes an API. The frontend sends a query to the backend. The backend generates query embedding, searches the vector database, applies filters, and returns ranked results.

# Pseudo-code: semantic search API

def semantic_search(user_id, query, filters):
    user_permissions = get_user_permissions(user_id)

    query_vector = embedding_model.embed(query)

    safe_filters = merge_filters(
        filters,
        {"allowed_roles": user_permissions.roles, "status": "active"}
    )

    candidates = vector_db.search(
        vector=query_vector,
        top_k=20,
        filter=safe_filters
    )

    ranked_results = reranker.rank(query, candidates)

    return ranked_results[:5]

 

This design separates responsibilities clearly. Authentication checks who the user is. Authorization decides what the user is allowed to search. The embedding model converts meaning into vectors. The vector database retrieves candidates. The re-ranker improves ordering. The API returns results with title, snippet, source, and score.

10.20 Business Use Cases of Semantic Search

Industry / Area

Use Case

Business Benefit

Banking

Search policy, KYC, fraud alerts, customer complaints

Faster resolution, better compliance support

Customer Support

Find similar tickets and solutions

Reduced handling time and better first-call resolution

Education

Search lessons using student-friendly language

Improved discovery and personalized learning

Healthcare

Search clinical guidelines and patient education material

Faster information access with proper safety controls

E-commerce

Meaning-based product search and recommendation

Better product discovery and conversion

Enterprise IT

Search runbooks, incidents, logs, and SOPs

Faster incident resolution

Legal

Search clauses, judgments, contracts, and obligations

Faster research, but requires strict validation

 

10.21 Best Practices

  • Start with a clear use case. Do not build a generic search engine without knowing what users need to find.
  • Clean the source data before embedding. Bad input produces bad search results.
  • Use meaningful chunks. Chunk by headings, sections, paragraphs, or logical units whenever possible.
  • Store strong metadata such as source, department, date, version, region, document type, and permission level.
  • Tune Top-K using real user questions, not guesswork.
  • Use hybrid search when exact codes, names, IDs, or technical terms matter.
  • Use re-ranking for important use cases where accuracy matters more than small latency increase.
  • Create a golden dataset of expected query-result pairs.
  • Monitor failed searches, low-score searches, and user feedback.
  • Never ignore security. Permission filtering must be applied before showing results or sending context to an LLM.

10.22 Mini-Project: Build a Leave Policy Semantic Search Design

This mini-project is a design exercise. You do not need to write full production code. The goal is to understand the components and decisions required for a semantic search system.

Problem statement: Build a semantic search feature that helps employees search company leave policy documents using natural language.

Input documents: Leave policy PDF, maternity policy DOCX, holiday calendar CSV, HR FAQ HTML page.

Users: Employees, HR admins, managers.

Expected output: Top 5 relevant sections with title, short snippet, source file, page number, and confidence score.

Suggested design

Mini-project design

HR Documents
   |
   v
Extract Text -> Clean Text -> Chunk by Section
   |
   v
Generate Embeddings
   |
   v
Store in Vector DB with Metadata
   |
   v
Employee Query -> Query Embedding
   |
   v
Search with Filters: country, employee_type, status
   |
   v
Return Top 5 Results with Source

Practice tasks

  1. List five sample employee questions for leave policy search.
  2. Define metadata fields you will store for each document chunk.
  3. Decide whether Top-K should be 3, 5, or 10 and explain why.
  4. Write two examples where keyword search may fail but semantic search can work.
  5. Create five golden dataset rows with query, expected document, and expected section.
  6. Think of one security rule that must be applied before returning results.

10.23 Chapter Summary

Semantic search is meaning-based search. It allows users to find relevant information even when the words in the query are different from the words in the document. It works by converting queries and documents into embeddings and then using similarity search to find related vectors.

Semantic search is different from keyword search. Keyword search is excellent for exact terms, IDs, codes, and names. Semantic search is excellent for natural language questions and meaning-based discovery. In production, many systems use hybrid search to combine both strengths.

Important concepts in semantic search include query embeddings, document embeddings, similarity search, Top-K results, metadata filtering, hybrid search, and re-ranking. A good semantic search system also needs evaluation, monitoring, security filters, and regular data refresh.

Semantic search is one of the foundation stones of RAG architecture. When an LLM answers questions from private company documents, semantic search is usually responsible for finding the right context before the LLM generates the final answer.

10.24 Key Terms

Term

Meaning

Semantic search

Search based on meaning and intent rather than only exact words.

Keyword search

Search based on exact words or terms.

Query embedding

Vector representation of the user query.

Document embedding

Vector representation of a document or chunk.

Vector

A list of numbers representing meaning.

Similarity search

Finding vectors closest to the query vector.

Top-K

The number of top matching results returned.

Metadata

Additional information about a document chunk, such as department, date, source, and permissions.

Hybrid search

Combination of keyword search and semantic search.

Re-ranking

Reordering initial search results using a stronger relevance model or business logic.

Golden dataset

A set of test queries with expected correct results used for evaluation.

 

10.25 Exercises

Exercise 1: Keyword vs Semantic Search

Write five pairs of queries and document sentences where the words are different but the meaning is similar. Example: Query “vacation balance” and document “annual leave entitlement”.

Exercise 2: Metadata Design

Assume you are building semantic search for a bank knowledge base. Define at least ten metadata fields that should be stored with each document chunk.

Exercise 3: Top-K Decision

For each use case below, decide whether Top-K should be 3, 5, or 10 and explain your choice: HR policy Q&A, technical troubleshooting, legal document research, product recommendation, and student learning search.

Exercise 4: Golden Dataset

Create a golden dataset with ten search queries for a company policy chatbot. For each query, write the expected document and expected section.

Exercise 5: Architecture Explanation

Draw your own semantic search architecture using text boxes. Include document ingestion, chunking, embedding generation, vector database, query embedding, metadata filtering, re-ranking, and result display.