Chapter 10
Semantic Search
Meaning-based search using embeddings, vector similarity, metadata filters, hybrid retrieval, re-ranking, and quality evaluation
Search is one of the most common features in software. We search emails, documents, support tickets, product catalogs, company policies, reports, code repositories, and learning material. Traditional search is usually based on keywords. It looks for words that appear in the user query and tries to find those same words inside the stored content.
Keyword search is useful, but it fails when the user and the document use different words for the same idea. An employee may search for “vacation rules”, while the HR document may use “annual leave entitlement”. A customer may search for “money deducted but order not placed”, while the support ticket says “payment captured but transaction failed”. A student may search for “how machines understand language”, while the course lesson is titled “natural language processing”.
Semantic search solves this problem by searching based on meaning. It does not depend only on exact word matching. It tries to understand the intent behind the query and find information that is semantically related. This makes it extremely useful for Generative AI, RAG systems, knowledge assistants, document Q&A, recommendation engines, duplicate ticket detection, fraud investigation, and enterprise search.
Semantic search is a search technique that finds results based on meaning rather than only exact words. The word “semantic” means meaning. Therefore, semantic search can be understood as meaning-based search.
In semantic search, text is converted into numerical vectors called embeddings. A vector is a list of numbers. These numbers represent the meaning of the text in a mathematical form. If two sentences have similar meaning, their vectors are close to each other. If two sentences have different meaning, their vectors are far apart.
Simple idea
User Query: "How many days of vacation can I take?"
Document A: "Employees are eligible for 21 days of annual leave."
Document B: "Office laptops must be returned during exit clearance."
Semantic Search Result:
Document A is selected because "vacation" and "annual leave" have related meaning.
A semantic search system usually has two stages. First, documents are processed and stored as embeddings. Second, when the user enters a query, that query is also converted into an embedding. The system then compares the query embedding with document embeddings and returns the closest matches.
|
Term |
Simple Meaning |
Example |
|
Query |
The text entered by the user |
“leave policy for new employee” |
|
Document |
The searchable content stored in the system |
HR policy PDF, ticket, article, course lesson |
|
Embedding |
Numerical representation of meaning |
[0.18, -0.42, 0.77, ...] |
|
Similarity search |
Finding vectors that are closest to the query vector |
Return top 5 most related policy sections |
|
Top-K |
Number of best results to return |
Top 3 or Top 10 documents |
Keyword search checks whether query words appear in the document. Semantic search checks whether the meaning of the query is close to the meaning of the document. Both are useful, but they solve different problems.
|
Aspect |
Keyword Search |
Semantic Search |
|
Main idea |
Searches exact or similar words |
Searches meaning and intent |
|
Example query |
“vacation policy” |
“How many paid leaves do I get?” |
|
Can find synonyms? |
Limited unless configured manually |
Yes, if the embedding model understands the relation |
|
Works well for |
Names, IDs, codes, exact terms, legal references |
Natural language questions, similar meaning, user intent |
|
Weakness |
Misses results when words differ |
May return related but not exact results |
|
Typical technology |
SQL LIKE, full-text index, BM25, Elasticsearch keyword search |
Embeddings, vector database, similarity search |
|
Best enterprise use |
Finding invoice number, customer ID, exact product code |
Finding policy answers, similar tickets, related documents |
For example, if a user searches for “PF withdrawal rules”, keyword search will perform well if the document also contains “PF withdrawal”. But if the document says “Provident Fund settlement process”, keyword search may not find it. Semantic search can understand that PF, Provident Fund, withdrawal, and settlement are related concepts.
Semantic search does not understand meaning the way humans do. It uses an embedding model trained on large amounts of text. During training, the model learns patterns between words, phrases, and contexts. It learns that words like “car”, “vehicle”, and “automobile” are related. It also learns that “refund”, “payment reversal”, and “money returned” can be related in customer support situations.
When the embedding model receives text, it converts it into a vector. This vector works like a mathematical fingerprint of meaning. A search system can compare these fingerprints to find similar ideas.
Meaning comparison in vector space
Text 1: "annual leave policy" -> Vector A
Text 2: "vacation rules" -> Vector B
Text 3: "laptop replacement process"-> Vector C
Vector A is close to Vector B.
Vector A is far from Vector C.
Therefore, Text 1 and Text 2 are likely about the same topic.
This is why semantic search is powerful for user-friendly search. A user does not need to know the exact wording used by the company, school, bank, hospital, or product team. They can ask naturally, and the search engine can still find useful results.
A query embedding is the vector representation of the user query. When a user types a search question, the system sends that text to an embedding model. The model returns a vector that represents the meaning of the query.
Input: User query text, such as “How can I apply for maternity leave?”
Processing: The embedding model converts the query into a vector.
Output: A query vector, such as [0.12, -0.08, 0.44, ...].
The query embedding is temporary. It is created at search time and used to find similar document embeddings already stored in the vector database.
# Pseudo-code: query embedding
query = "How can I apply for maternity leave?"
query_vector = embedding_model.embed(query)
results = vector_database.search(query_vector, top_k=5)
A document embedding is the vector representation of a document or part of a document. In most real systems, large documents are not embedded as one giant block. They are split into smaller pieces called chunks. Each chunk is converted into an embedding and stored in a vector database with useful metadata.
Input: A document chunk, such as one section of an HR policy.
Processing: The embedding model converts the chunk into a vector.
Output: A document vector stored with document ID, title, page number, department, date, and source link.
Document embeddings are usually created before the user searches. This process is called indexing. The indexing pipeline may run once, daily, hourly, or whenever new documents are uploaded.
Document embedding pipeline
Company PDF / DOCX / HTML / CSV
|
v
Document Loader
|
v
Text Extraction and Cleaning
|
v
Chunking
|
v
Embedding Model
|
v
Vector Database + Metadata
Similarity search means finding stored vectors that are closest to the query vector. The search engine compares the query vector with many document vectors and calculates how similar they are. The most similar vectors are returned as search results.
A common similarity measurement is cosine similarity. You do not need deep mathematics at this stage. The simple idea is that cosine similarity checks whether two vectors point in a similar direction. If the direction is similar, the meaning is considered similar. A score closer to 1 usually means more similar, while a lower score means less similar.
|
Query |
Document Chunk |
Similarity Score |
Interpretation |
|
How many annual leaves do I get? |
Employees receive 21 days of annual leave per year. |
0.91 |
Very relevant |
|
How many annual leaves do I get? |
Employees must submit laptop serial number during onboarding. |
0.22 |
Not relevant |
|
How many annual leaves do I get? |
Leave encashment is calculated during final settlement. |
0.72 |
Related, but may not answer directly |
Similarity search is the heart of semantic search. However, similarity alone is not always enough. In production systems, we also use metadata filters, hybrid search, and re-ranking to improve accuracy.
Top-K means the number of best matching results returned by the search system. If K is 5, the system returns the top 5 most similar chunks. If K is 10, it returns the top 10.
Choosing Top-K is important. If K is too small, the system may miss useful information. If K is too large, the system may include irrelevant information, increase cost, and confuse the LLM in a RAG system.
|
Top-K Value |
When Useful |
Risk |
|
Top 1 |
Very precise FAQ search where one answer is expected |
May miss the correct answer |
|
Top 3 |
Short policy questions or simple support articles |
May not include enough context |
|
Top 5 |
Common default for document Q&A |
Usually balanced |
|
Top 10+ |
Complex research questions or broad topics |
May add noisy results and increase token cost |
In RAG applications, Top-K controls how many retrieved chunks are passed to the LLM as context. More chunks are not always better. Good retrieval means passing the right context, not the maximum context.
Metadata is additional information stored with each document chunk. It helps the search system filter results before or after similarity search. Metadata can include department, document type, language, region, effective date, security level, product name, customer type, and source file.
|
Metadata Field |
Example Value |
How It Helps |
|
department |
HR |
Search only HR policies |
|
country |
India |
Return policies applicable to India |
|
document_type |
Policy |
Exclude old emails or informal notes |
|
effective_date |
2026-04-01 |
Prefer latest valid policy |
|
security_level |
Internal |
Prevent unauthorized access |
|
source_url |
hr_policy_2026.pdf page 4 |
Show citation or source |
Metadata filtering is very important in enterprise AI. Without metadata filters, a user may receive information from the wrong country, old policy version, wrong department, or unauthorized document.
# Pseudo-code: semantic search with metadata filter
query_vector = embedding_model.embed("What is the leave policy for India employees?")
results = vector_database.search(
vector=query_vector,
top_k=5,
filter={"department": "HR", "country": "India", "status": "active"}
)
Hybrid search combines keyword search and semantic search. This is useful because keyword search and semantic search have different strengths. Keyword search is excellent for exact terms like invoice numbers, product codes, names, error codes, and legal clause numbers. Semantic search is excellent for meaning-based natural language questions.
A hybrid system may run both searches and combine the scores. For example, if a user searches “Oracle error ORA-12514 listener issue”, keyword search can strongly match ORA-12514, while semantic search can find articles about Oracle listener configuration even if the exact phrase is missing.
Hybrid search flow
User Query
|
+--> Keyword Search Engine ----+
| |
+--> Semantic Vector Search ---+--> Score Fusion / Merge
|
v
Ranked Results
|
Search Type |
Best For |
Example |
|
Keyword search |
Exact IDs, codes, names, product numbers |
INV-2026-00981, ORA-12514, GSTIN |
|
Semantic search |
Natural language meaning |
“payment failed but money deducted” |
|
Hybrid search |
Mixed business queries |
“SBI PO eligibility age relaxation rules” |
Many production AI systems use hybrid search because users rarely search in only one style. Sometimes they use exact terms. Sometimes they use vague natural language. Hybrid search gives better coverage.
Re-ranking is the process of taking initial search results and sorting them again using a stronger relevance model or additional business logic. The first retrieval stage is usually optimized for speed. It may return 20 or 50 candidate results. The re-ranker then evaluates those candidates more carefully and selects the best results.
A simple vector search may retrieve chunks that are broadly related but not the best answer. A re-ranker can compare the query and each candidate chunk in more detail. This often improves answer quality in RAG systems.
Re-ranking flow
User Query
|
v
Vector Search returns Top 30 candidates quickly
|
v
Re-ranker checks query-document relevance more deeply
|
v
Final Top 5 results sent to user or LLM
|
Stage |
Goal |
Typical Output |
|
Initial retrieval |
Fast broad search |
Top 20 or Top 50 candidate chunks |
|
Re-ranking |
Improve relevance order |
Best 3 to 5 chunks |
|
Final response |
Use best evidence |
Answer with citations or ranked search results |
Re-ranking increases quality but may also increase cost and latency. It should be used when search accuracy is important, such as legal documents, medical knowledge bases, banking policies, and enterprise assistants.
Complete semantic search architecture
INDEXING PIPELINE
Documents / Tickets / Web Pages / PDFs / CSVs
|
v
Document Loader
|
v
Text Extraction and Cleaning
|
v
Chunking
|
v
Embedding Model
|
v
Vector Database / Index
+ Metadata + Source Info
SEARCH PIPELINE
User Search Query
|
v
Query Preprocessing
|
v
Query Embedding Model
|
v
Vector Similarity Search + Metadata Filters
|
v
Optional Hybrid Search Merge
|
v
Optional Re-ranking
|
v
Top-K Results
|
v
User Search Results OR RAG Context for LLM
The architecture has two main parts: the indexing pipeline and the search pipeline. The indexing pipeline prepares the knowledge base before the user searches. The search pipeline runs whenever the user asks a question or enters a search query.
Imagine a company has many HR documents: leave policy, maternity policy, travel policy, reimbursement policy, work-from-home policy, and exit policy. Employees do not remember exact document names. They ask natural questions.
|
User Query |
Document Text |
Why Semantic Search Helps |
|
Can I take leave during probation? |
Employees under probation are eligible for casual leave subject to manager approval. |
Query uses “take leave”; document uses “eligible for casual leave”. |
|
How many vacation days do I get? |
Employees receive 21 days of annual leave per calendar year. |
Query uses “vacation”; document uses “annual leave”. |
|
Can I carry forward unused leaves? |
A maximum of 10 unused annual leave days can be carried forward. |
Query and document use related but not identical phrasing. |
A good semantic search system should also use metadata filters. For example, if the user is in India, the system should prefer India HR policies. If the user is a contractor, it should not return full-time employee benefits unless clearly applicable.
Leave policy search flow
Employee Query: "Can I carry forward my vacation balance?"
|
v
Query Embedding
|
v
Vector Search in HR Policy Chunks
|
v
Filter: country=India, status=active, audience=employee
|
v
Top Results:
1. Annual Leave Carry Forward Section
2. Leave Encashment Section
3. HR FAQ on Leave Balance
Customer support teams receive thousands of complaints. Many complaints are similar but written in different words. Semantic search can find related complaints, similar past resolutions, duplicate tickets, and knowledge base articles.
|
Customer Complaint |
Related Stored Ticket |
Semantic Relation |
|
Money debited but order failed |
Payment captured but transaction unsuccessful |
Same issue with different words |
|
App keeps closing after login |
Application crashes after authentication |
Same technical meaning |
|
Delivery boy did not come but status says delivered |
False delivery confirmation reported by customer |
Related operational complaint |
Support agents can use semantic search to quickly find previous cases and suggested resolutions. Managers can use it to group complaints by theme, identify recurring issues, and improve product quality.
Customer complaint semantic search
New Ticket: "Amount deducted but booking not confirmed"
|
v
Semantic Search over Past Tickets
|
v
Top Similar Tickets:
- Payment captured but booking failed
- Transaction success at bank but failure in booking service
- Refund pending for failed booking
|
v
Suggested Routing: Payments Support Team
Suggested Action: Check payment gateway status and refund workflow
In an education platform, students may not know exact technical terminology. A beginner may search “how AI remembers words”, while the course content may contain “embeddings and vector representation”. Semantic search helps students discover relevant lessons even when their vocabulary is incomplete.
|
Student Search |
Relevant Lesson |
Why It Matches |
|
How AI remembers meaning of words |
Embeddings and Semantic Meaning |
The intent is about word meaning representation. |
|
How chatbot finds answers from PDF |
RAG Architecture and Document Retrieval |
The intent is document-based question answering. |
|
How to compare two sentences |
Cosine Similarity and Vector Search |
The intent is sentence similarity. |
Semantic search also powers recommendation. If a student reads a lesson about embeddings, the platform can recommend semantic search, vector databases, and RAG architecture because these topics are conceptually related.
A semantic search system should not be judged only by whether it returns something. It should be judged by whether it returns the right results, in the right order, quickly, safely, and consistently. Search quality evaluation helps teams improve retrieval before building bigger AI applications on top of it.
|
Metric |
Meaning |
Simple Example |
|
Relevance |
Are the returned results useful for the query? |
Leave policy query returns leave policy section. |
|
Precision |
How many returned results are actually relevant? |
Out of 5 results, 4 are useful. |
|
Recall |
Did the system find all important relevant results? |
It found both annual leave and carry-forward rules. |
|
Ranking quality |
Are the best results at the top? |
The exact answer appears as result 1, not result 8. |
|
Coverage |
Can the system search all important content sources? |
Policies, FAQs, tickets, and manuals are included. |
|
Latency |
How fast are results returned? |
Search completes in less than 1 second. |
|
Freshness |
Are latest documents preferred? |
2026 policy appears before 2023 policy. |
|
Permission safety |
Does the system hide unauthorized content? |
Employee cannot see confidential HR notes. |
A practical evaluation approach is to create a golden dataset. A golden dataset contains sample queries and expected correct results. For example, “How many days of annual leave do I get?” should return the annual leave entitlement section. Whenever the search system changes, test it again against the golden dataset.
# Example golden dataset row
query: "Can unused leave be carried forward?"
expected_document: "HR Leave Policy 2026"
expected_section: "Leave Carry Forward"
minimum_similarity_score: 0.75
must_not_return: "Travel Reimbursement Policy"
|
Challenge |
What Happens |
How to Reduce the Problem |
|
Poor chunking |
Search returns incomplete or confusing text. |
Use meaningful section-based chunks and overlap where needed. |
|
Old documents |
System returns outdated policies. |
Use metadata such as version, status, and effective date. |
|
Wrong Top-K |
Too few or too many results are retrieved. |
Tune Top-K using evaluation queries. |
|
No permission filtering |
Users may see restricted content. |
Apply access control filters before returning results. |
|
Weak embedding model |
Similar meanings are not matched well. |
Choose an embedding model suitable for your language and domain. |
|
Domain vocabulary |
Model may not understand internal terms. |
Add glossary, metadata, hybrid search, or domain-specific fine-tuning when needed. |
|
Duplicate content |
Same result appears multiple times. |
Deduplicate documents and merge near-identical chunks. |
|
No re-ranking |
Useful results appear lower in the list. |
Use re-ranking for high-value use cases. |
|
No evaluation |
Team cannot measure improvement. |
Create test queries and expected results. |
Semantic search is powerful, but it is not magic. Most failures happen because of poor data preparation, weak metadata, bad chunking, missing security filters, or lack of evaluation. A good AI engineer treats semantic search as an engineering system, not only as a model call.
A simple semantic search application usually exposes an API. The frontend sends a query to the backend. The backend generates query embedding, searches the vector database, applies filters, and returns ranked results.
# Pseudo-code: semantic search API
def semantic_search(user_id, query, filters):
user_permissions = get_user_permissions(user_id)
query_vector = embedding_model.embed(query)
safe_filters = merge_filters(
filters,
{"allowed_roles": user_permissions.roles, "status": "active"}
)
candidates = vector_db.search(
vector=query_vector,
top_k=20,
filter=safe_filters
)
ranked_results = reranker.rank(query, candidates)
return ranked_results[:5]
This design separates responsibilities clearly. Authentication checks who the user is. Authorization decides what the user is allowed to search. The embedding model converts meaning into vectors. The vector database retrieves candidates. The re-ranker improves ordering. The API returns results with title, snippet, source, and score.
|
Industry / Area |
Use Case |
Business Benefit |
|
Banking |
Search policy, KYC, fraud alerts, customer complaints |
Faster resolution, better compliance support |
|
Customer Support |
Find similar tickets and solutions |
Reduced handling time and better first-call resolution |
|
Education |
Search lessons using student-friendly language |
Improved discovery and personalized learning |
|
Healthcare |
Search clinical guidelines and patient education material |
Faster information access with proper safety controls |
|
E-commerce |
Meaning-based product search and recommendation |
Better product discovery and conversion |
|
Enterprise IT |
Search runbooks, incidents, logs, and SOPs |
Faster incident resolution |
|
Legal |
Search clauses, judgments, contracts, and obligations |
Faster research, but requires strict validation |
This mini-project is a design exercise. You do not need to write full production code. The goal is to understand the components and decisions required for a semantic search system.
Problem statement: Build a semantic search feature that helps employees search company leave policy documents using natural language.
Input documents: Leave policy PDF, maternity policy DOCX, holiday calendar CSV, HR FAQ HTML page.
Users: Employees, HR admins, managers.
Expected output: Top 5 relevant sections with title, short snippet, source file, page number, and confidence score.
Mini-project design
HR Documents
|
v
Extract Text -> Clean Text -> Chunk by Section
|
v
Generate Embeddings
|
v
Store in Vector DB with Metadata
|
v
Employee Query -> Query Embedding
|
v
Search with Filters: country, employee_type, status
|
v
Return Top 5 Results with Source
Semantic search is meaning-based search. It allows users to find relevant information even when the words in the query are different from the words in the document. It works by converting queries and documents into embeddings and then using similarity search to find related vectors.
Semantic search is different from keyword search. Keyword search is excellent for exact terms, IDs, codes, and names. Semantic search is excellent for natural language questions and meaning-based discovery. In production, many systems use hybrid search to combine both strengths.
Important concepts in semantic search include query embeddings, document embeddings, similarity search, Top-K results, metadata filtering, hybrid search, and re-ranking. A good semantic search system also needs evaluation, monitoring, security filters, and regular data refresh.
Semantic search is one of the foundation stones of RAG architecture. When an LLM answers questions from private company documents, semantic search is usually responsible for finding the right context before the LLM generates the final answer.
|
Term |
Meaning |
|
Semantic search |
Search based on meaning and intent rather than only exact words. |
|
Keyword search |
Search based on exact words or terms. |
|
Query embedding |
Vector representation of the user query. |
|
Document embedding |
Vector representation of a document or chunk. |
|
Vector |
A list of numbers representing meaning. |
|
Similarity search |
Finding vectors closest to the query vector. |
|
Top-K |
The number of top matching results returned. |
|
Metadata |
Additional information about a document chunk, such as department, date, source, and permissions. |
|
Hybrid search |
Combination of keyword search and semantic search. |
|
Re-ranking |
Reordering initial search results using a stronger relevance model or business logic. |
|
Golden dataset |
A set of test queries with expected correct results used for evaluation. |
Write five pairs of queries and document sentences where the words are different but the meaning is similar. Example: Query “vacation balance” and document “annual leave entitlement”.
Assume you are building semantic search for a bank knowledge base. Define at least ten metadata fields that should be stored with each document chunk.
For each use case below, decide whether Top-K should be 3, 5, or 10 and explain your choice: HR policy Q&A, technical troubleshooting, legal document research, product recommendation, and student learning search.
Create a golden dataset with ten search queries for a company policy chatbot. For each query, write the expected document and expected section.
Draw your own semantic search architecture using text boxes. Include document ingestion, chunking, embedding generation, vector database, query embedding, metadata filtering, re-ranking, and result display.