Chapter 11
Vector Databases
Storing, indexing, filtering, and searching embeddings for semantic search and RAG applications
In the previous chapters, you learned about embeddings and semantic search. An embedding converts text, images, audio, or other information into a vector, which is a list of numbers representing meaning. Semantic search compares these vectors to find content that is similar in meaning to the user query.
A vector database is the storage and search engine that makes this possible in real applications. It stores millions or billions of vectors and quickly finds the closest vectors when a user asks a question. Without a vector database, a RAG chatbot, document assistant, semantic search engine, recommendation engine, or duplicate ticket finder would become slow and difficult to manage.
Think of a vector database as a library designed for meaning-based search. A normal library may organize books by title, author, or category. A vector database organizes information by meaning. If a user searches for “employee vacation rules”, it can find “annual leave policy” even if the exact words are different.
Simple mental model
Normal Database:
Find rows where name = "Annual Leave Policy"
Find rows where employee_id = 1001
Find rows where status = "Open"
Vector Database:
Find documents whose meaning is closest to:
"How many paid leaves can an employee take?"
A vector database is a specialized database used to store, index, and search high-dimensional vectors. These vectors are usually created by embedding models. Each vector represents the meaning of a piece of content such as a paragraph, sentence, PDF chunk, image, product description, support ticket, or user query.
A vector database does not only store the vector. It usually stores three important things together: the vector, the original or reference content, and metadata. Metadata may include document name, page number, department, created date, customer ID, language, access level, category, product type, or any other information useful for filtering and security.
|
Stored item |
Purpose |
Example |
|
Vector |
Numerical representation of meaning |
[0.12, -0.77, 0.35, ...] |
|
Text or reference |
Original chunk or link to original source |
HR Policy, page 4, paragraph 2 |
|
Metadata |
Extra information for filtering and governance |
department=HR, country=India, access=employee |
|
ID |
Unique identifier for update, delete, and traceability |
doc_123_chunk_005 |
In a RAG system, the vector database is used during retrieval. The user question is converted into an embedding. The vector database compares the query vector with stored document vectors and returns the most relevant chunks. These chunks are then passed to the LLM as context.
Relational databases such as MySQL, Oracle, SQL Server, and PostgreSQL are excellent for structured business data. They are designed to search by exact values, ranges, joins, and indexes. For example, they can quickly answer questions like “Find all orders for customer 1001” or “Find employees where department is Finance”.
Vector search has a different requirement. It must compare one query vector with many stored vectors and find the nearest vectors. A single vector may have hundreds or thousands of dimensions. A table may contain millions of such vectors. Comparing the query vector with every stored vector one by one is too slow for production applications.
A normal database can store arrays or JSON values, but storage alone is not enough. The difficult part is fast similarity search. Vector databases solve this using specialized indexes and approximate nearest neighbor algorithms.
|
Requirement |
Normal database strength |
Vector database strength |
|
Exact lookup |
Excellent |
Possible, but not the main purpose |
|
Transaction processing |
Excellent |
Usually limited |
|
Joins and reports |
Excellent |
Not designed for complex joins |
|
Meaning-based similarity search |
Weak without special extensions |
Excellent |
|
Search millions of embeddings quickly |
Difficult |
Designed for this |
|
Metadata filtering with vector search |
Possible with extensions |
Usually built in |
Why brute force becomes slow
User query vector: Q
Stored vectors: V1, V2, V3, ... V10,000,000
Brute force approach:
Compare Q with every vector one by one.
This can be too slow when vectors and users increase.
Vector index approach:
Use an index to quickly navigate to likely nearest vectors.
A vector index is a data structure that helps the database find similar vectors quickly. Just as a book index helps you find a topic without reading the whole book, a vector index helps the database find nearest vectors without comparing the query vector with every vector in storage.
Vector indexes are optimized for similarity search. They organize vectors in a way that makes nearest-neighbor lookup faster. Different vector databases support different index types. Common concepts include graph-based indexes, inverted file indexes, quantization, clustering, and hierarchical navigation.
|
Index idea |
Simple explanation |
Practical meaning |
|
Graph-based index |
Vectors are connected like a network of nearby points |
Fast search by moving through nearby nodes |
|
Clustering |
Similar vectors are grouped together |
Search starts in the most likely groups |
|
Quantization |
Vectors are compressed approximately |
Saves memory and improves speed, with some accuracy trade-off |
|
Flat index |
Compares more directly |
More accurate but slower for large datasets |
Choosing an index is a trade-off. A highly accurate index may be slower or use more memory. A very fast index may miss some relevant results. Production systems balance speed, cost, memory, and accuracy.
Nearest neighbor search means finding vectors closest to the query vector. Approximate nearest neighbor search, often called ANN, means finding nearly closest vectors very quickly instead of always finding the mathematically exact closest vectors.
This approximation is acceptable in many AI applications because users usually care about useful results, not perfect mathematical ranking. For example, if a user asks “How do I claim travel reimbursement?”, the system should return relevant travel reimbursement policy chunks. It does not matter if the first and second chunks are swapped, as long as both are useful and correct.
Exact search: Very accurate but may be slow for large datasets because it compares many vectors.
Approximate search: Much faster and usually accurate enough for semantic search and RAG.
Production goal: Return useful results quickly with acceptable relevance.
ANN intuition
Imagine finding the nearest shop in a city.
Exact method: Measure distance to every shop in the city.
Approximate method: First go to the correct market area, then check nearby shops.
Vector databases use a similar idea to reduce unnecessary comparisons.
Metadata is extra information stored with each vector. It does not represent meaning directly, but it helps filter, secure, organize, and explain search results. In enterprise AI, metadata is as important as embeddings because it controls which content can be retrieved and how results are interpreted.
|
Metadata field |
Example value |
Why it is useful |
|
source_file |
employee_handbook.pdf |
Shows where the answer came from |
|
page_number |
12 |
Useful for citations and verification |
|
department |
HR |
Filter HR documents only |
|
country |
India |
Return country-specific policy |
|
access_level |
employee |
Prevent unauthorized retrieval |
|
created_at |
2026-06-01 |
Refresh and version control |
|
document_version |
v3 |
Avoid outdated documents |
|
chunk_id |
doc45_chunk08 |
Trace answer to exact chunk |
For example, a company may have leave policies for India, USA, and Germany. If an employee in India asks about leave rules, the system should not retrieve the USA policy. Metadata filtering helps apply country, department, role, and access restrictions before or during vector search.
Filtering means limiting search results using metadata conditions. Vector search finds content by meaning, while filtering controls the allowed scope. Good filtering improves relevance, security, and cost efficiency.
Example query: “What is the maternity leave policy?”
Possible filter: department = HR, country = India, access_level <= employee, status = active
Result: The system retrieves only active Indian HR policy chunks that the employee is allowed to see.
Filters can be applied before vector search, after vector search, or during vector search depending on the database. Pre-filtering reduces the search space before similarity search. Post-filtering searches first and then removes results that do not match filters. Many production systems prefer databases that support efficient filtering together with vector search.
|
Filtering type |
How it works |
Advantage |
Risk |
|
Pre-filtering |
Apply metadata filter first, then vector search |
Secure and efficient for strict filters |
May reduce search space too much |
|
Post-filtering |
Vector search first, filter results later |
May find semantically strong candidates |
Can return too few results after filtering |
|
Integrated filtering |
Database combines vector search and metadata conditions |
Best balance for many enterprise use cases |
Depends on database capability |
Collections and namespaces are ways to organize vectors. The exact terminology differs across vector databases, but the idea is similar: group related vectors so they can be managed, searched, updated, and secured properly.
Collection: A group of vectors that usually share the same embedding model, schema, and purpose. Example: hr_policy_chunks.
Namespace: A logical subdivision inside a collection or index. Example: tenant_abc, india, production, or dev.
For a small project, you may use one collection for all documents. For a larger enterprise project, you may separate collections by application, data type, tenant, language, or access boundary.
|
Design option |
Example |
When useful |
|
One collection for all documents |
company_knowledge_base |
Small internal chatbot |
|
Collection per domain |
hr_docs, finance_docs, it_docs |
Different business ownership and filters |
|
Namespace per tenant |
customer_a, customer_b |
SaaS applications with multiple clients |
|
Collection per embedding model |
docs_embedding_v1, docs_embedding_v2 |
Migration from one model to another |
|
Separate environments |
dev, test, prod |
Safe testing and deployment |
Indexing strategy means deciding how vectors will be created, stored, organized, refreshed, and searched. A weak indexing strategy can make a good AI application unreliable. A strong indexing strategy improves relevance, speed, cost, and maintainability.
Before indexing, you should decide the chunking strategy, embedding model, metadata fields, collection structure, update process, deletion process, and evaluation method. These decisions are connected. For example, if chunks are too large, search results may be broad and less precise. If chunks are too small, the LLM may not receive enough context.
|
Decision |
Question to ask |
Example choice |
|
Chunk size |
How much text should each vector represent? |
500-800 tokens for policy documents |
|
Chunk overlap |
Should nearby chunks share text? |
100 tokens overlap to preserve context |
|
Embedding model |
Which model creates vectors? |
One text embedding model for all English documents |
|
Metadata fields |
What filters and citations are needed? |
file, page, department, country, access_level |
|
Index type |
What speed and accuracy are needed? |
ANN index for production search |
|
Refresh method |
How will updates be handled? |
Incremental update on document change |
|
Deletion policy |
How will outdated vectors be removed? |
Delete by document_id and version |
Indexing pipeline
Source Documents
|
v
Load files: PDF, DOCX, HTML, CSV
|
v
Clean and split into chunks
|
v
Generate embeddings for each chunk
|
v
Attach metadata: file, page, owner, access, version
|
v
Store vectors + metadata in vector database
|
v
Build or update vector index
There are many vector database and vector search options. Some are lightweight and easy for learning. Some are designed for large-scale production. Some are managed cloud services. Some are libraries that run inside your application. Some are extensions of existing databases.
The best choice depends on your project size, cost, deployment preference, security requirements, operational skills, and integration needs. For learning, simple local tools are often enough. For production, you must think about scaling, backup, monitoring, access control, latency, filtering, and maintenance.
Chroma is a beginner-friendly vector database often used for local experiments, prototypes, and small RAG applications. It is popular in tutorials because it is easy to install and integrate with Python frameworks.
Good for: Learning, local RAG prototypes, small document search demos, quick proof of concept.
Be careful about: Production operations, large-scale requirements, backup, monitoring, and enterprise governance depending on deployment setup.
FAISS is a vector search library rather than a full enterprise database. It is widely used for efficient similarity search and supports different index types. It is strong when you want local control and high-performance search, but you may need to build surrounding database features yourself.
Good for: Research, local search, custom vector indexing, high-performance experiments.
Be careful about: Metadata management, distributed operations, access control, and application-level integration work.
Pinecone is a managed vector database service. It is designed to reduce operational work because infrastructure, scaling, and availability are handled by the service provider. It is useful when a team wants to build production vector search without managing servers directly.
Good for: Managed production vector search, teams that prefer cloud service operations, scalable RAG applications.
Be careful about: Cloud cost, data residency, vendor dependency, and compliance requirements.
Weaviate is an open-source vector database with features for vector search, metadata filtering, hybrid search, and schema-based organization. It can be self-hosted or used as a managed service depending on project needs.
Good for: Semantic search platforms, hybrid search, applications needing schema and filtering features.
Be careful about: Operational setup, cluster management, and correct schema design for production.
Milvus is an open-source vector database designed for large-scale similarity search. It is often considered for high-volume vector workloads where performance and scalability are important.
Good for: Large-scale vector search, high-volume applications, self-hosted production environments.
Be careful about: Infrastructure complexity, operational skills, resource planning, and monitoring.
Qdrant is a vector database focused on vector search with payload metadata filtering. It is used for semantic search, recommendation, and RAG workloads. It can be self-hosted or used as a managed service.
Good for: Applications needing strong metadata filtering, semantic search, recommendation, and production-friendly APIs.
Be careful about: Sizing, indexing configuration, backup, and cost planning depending on deployment model.
pgvector is a PostgreSQL extension that allows vectors to be stored and searched inside PostgreSQL. This is useful when your application already uses PostgreSQL and you want to keep structured data and vector search close together.
Good for: Small to medium applications, teams already using PostgreSQL, combining relational filters with vector search.
Be careful about: Very large vector workloads, scaling limits, index tuning, and performance compared with specialized vector databases.
|
Option |
Type |
Best for |
Main advantage |
Main caution |
|
Chroma |
Vector database / local-first tool |
Learning and prototypes |
Easy to start |
Validate production needs carefully |
|
FAISS |
Vector search library |
Research and custom local search |
High-performance indexing |
Not a complete database by itself |
|
Pinecone |
Managed vector database |
Production cloud RAG |
Less infrastructure work |
Cost and vendor dependency |
|
Weaviate |
Vector database |
Semantic and hybrid search |
Schema, filtering, hybrid capabilities |
Requires good deployment design |
|
Milvus |
Vector database |
Large-scale vector search |
Designed for scale |
Operational complexity |
|
Qdrant |
Vector database |
Search with metadata payload filters |
Filtering and production APIs |
Sizing and operations planning |
|
PostgreSQL pgvector |
Database extension |
Apps already using PostgreSQL |
Keep relational and vector data together |
May not fit very large vector workloads |
For beginners, Chroma and FAISS are good for learning. PostgreSQL pgvector is useful if you already know SQL and PostgreSQL. For production cloud applications, managed services or production-grade vector databases become more important. The decision should not be based only on popularity; it should be based on your architecture, budget, security, and scale.
|
Scenario |
Suggested option |
Reason |
|
Student learning RAG locally |
Chroma or FAISS |
Simple setup and low cost |
|
Small internal chatbot with PostgreSQL backend |
PostgreSQL pgvector |
Keeps app data and vectors in one database |
|
Enterprise document assistant with managed operations |
Pinecone or managed vector service |
Reduces infrastructure management |
|
Open-source production search with filtering |
Qdrant or Weaviate |
Good APIs and metadata filtering options |
|
Very large vector workloads |
Milvus or managed scalable vector DB |
Designed for scale and high volume |
|
Research or custom search algorithm |
FAISS |
Flexible vector index library |
A practical beginner rule is simple: start with the easiest tool, learn the concepts, then upgrade when the project requirements demand it. Do not begin with complex infrastructure if your goal is to understand RAG and semantic search. First build a working prototype. Then improve storage, security, scaling, and monitoring.
A vector database schema defines what you store with each vector. Some vector databases are schema-light, while others use a clear schema. Even if the database does not force a strict schema, your application should still follow a consistent structure.
Collection: company_documents
Fields:
- id: string
- vector: float[]
- text: string
- source_file: string
- source_type: string # pdf, docx, html, csv
- page_number: integer
- chunk_number: integer
- department: string
- country: string
- access_level: string
- document_version: string
- created_at: datetime
- updated_at: datetime
- status: string # active, archived, deleted
For example, an HR policy PDF may be split into 20 chunks. Each chunk gets one vector and one metadata record. If the policy is updated, the system can delete or archive all old chunks for that document version and insert new chunks.
|
Field |
Example |
Purpose |
|
id |
hr_leave_v3_chunk_004 |
Unique chunk identity |
|
text |
Employees are eligible for... |
Context passed to LLM |
|
source_file |
leave_policy.pdf |
Citation and traceability |
|
page_number |
7 |
Source reference |
|
department |
HR |
Domain filtering |
|
country |
India |
Local policy filtering |
|
access_level |
employee |
Security control |
|
status |
active |
Avoid retrieving deleted content |
A vector database normally participates in two flows: the indexing flow and the search flow. The indexing flow prepares and stores content. The search flow retrieves relevant content when a user asks a question.
Vector DB architecture diagram
INDEXING FLOW
Documents / Records / Images / Tickets
|
v
Loader and Parser
|
v
Cleaning and Chunking
|
v
Embedding Model
|
v
Vector + Text + Metadata Records
|
v
Vector Database Index
--------------------------------------------------
SEARCH FLOW
User Query
|
v
Embedding Model
|
v
Query Vector
|
v
Vector Database Search + Metadata Filters
|
v
Top-K Similar Chunks
|
v
RAG Context Builder
|
v
LLM Generates Answer with Sources
In production, this architecture also includes authentication, authorization, logging, monitoring, feedback collection, caching, and evaluation. The vector database is not a standalone magic component. It must be connected properly with the application backend, data pipeline, LLM orchestration, and security layer.
Let us design a simple RAG use case using a vector database. The company wants an assistant that answers employee questions from HR, IT, and finance policy documents. The assistant should answer using company documents, show sources, and avoid answering from unauthorized documents.
Employees ask questions such as “How many casual leaves do I get?”, “What is the laptop replacement process?”, or “How do I claim travel expenses?” The company has answers in PDF and DOCX policy documents, but employees find it difficult to search manually.
1. Collect HR, IT, and finance policy documents.
2. Parse PDF and DOCX files into clean text.
3. Split documents into meaningful chunks.
4. Generate embeddings for each chunk.
5. Attach metadata such as department, country, page number, document version, and access level.
6. Store vectors, text, and metadata in the vector database.
7. Create or update the vector index.
1. Employee enters a question in the chatbot.
2. Application checks user identity and role.
3. Question is converted into a query embedding.
4. Vector database searches for top relevant chunks with metadata filters such as country and access level.
5. The top chunks are sent to the LLM as context.
6. The LLM generates an answer based only on retrieved context.
7. The answer includes source file and page number.
8. Logs and feedback are stored for monitoring and improvement.
RAG with vector database
Employee Question:
"How many paid leaves do I get?"
|
v
Check employee role and country
|
v
Create query embedding
|
v
Search Vector DB:
filter country = India
filter department = HR
filter status = active
|
v
Return Top-K policy chunks
|
v
LLM receives context and generates answer
|
v
Answer:
"According to the Leave Policy, employees are eligible for ..."
Source: leave_policy.pdf, page 7
Vector databases introduce cost and scaling considerations. Cost may come from storage, compute, memory, managed service fees, network usage, backups, and embedding generation. Scaling depends on number of vectors, vector dimensions, query volume, filtering complexity, index type, and latency requirements.
|
Cost factor |
Why it matters |
Optimization idea |
|
Number of vectors |
More chunks mean more storage and search work |
Use sensible chunking and remove duplicate content |
|
Vector dimension |
Higher dimension can use more memory |
Choose embedding model carefully |
|
Query volume |
More users increase compute cost |
Cache common queries and tune Top-K |
|
Top-K value |
Returning too many results increases processing |
Start with 3-5 for RAG and evaluate |
|
Metadata filters |
Complex filters can affect performance |
Design metadata fields intentionally |
|
Index type |
Different indexes have memory and accuracy trade-offs |
Benchmark with real data |
|
Managed service |
Convenient but recurring cost |
Monitor usage and set budgets |
A common beginner mistake is storing every sentence as a separate vector without thinking about cost. Another mistake is using very large chunks and then wondering why answers are vague. Good vector database design starts with good data preparation.
AI applications must handle changing data. Policies change, product catalogs update, support articles become outdated, and customer records may need deletion. If the vector database is not refreshed correctly, the AI system may answer using old or unauthorized information.
Full refresh: Delete and rebuild all vectors. Simple but may be slow and expensive for large datasets.
Incremental refresh: Only update vectors for changed documents. More efficient but requires tracking document IDs, versions, and timestamps.
Soft delete: Mark vectors as inactive or deleted using metadata. Useful for audit and recovery.
Hard delete: Physically remove vectors. Required in some privacy and compliance scenarios.
|
Situation |
Recommended action |
Reason |
|
Document updated |
Delete old version chunks and insert new version |
Avoid mixed answers from old and new policy |
|
Document archived |
Set status=archived or remove from active collection |
Prevent outdated retrieval |
|
User requests data deletion |
Hard delete vectors linked to that user if required |
Privacy compliance |
|
Embedding model changed |
Re-embed documents into new collection |
Vectors from different models should not be mixed casually |
|
Access rule changed |
Update metadata or re-index affected chunks |
Prevent data leakage |
When a document changes, you should not simply add new vectors and leave old vectors active. This creates contradictory answers. Always design a versioning and deletion strategy before production.
|
Mistake |
Why it is a problem |
Better approach |
|
Using vector DB before understanding chunking |
Poor chunks create poor search results |
Design and test chunking first |
|
No metadata |
Cannot filter, secure, or cite results |
Store source, page, role, version, status |
|
Mixing vectors from different embedding models |
Similarity may become unreliable |
Use separate collections or re-embed consistently |
|
Returning too many chunks to the LLM |
Higher cost and noisy context |
Tune Top-K and re-ranking |
|
No deletion strategy |
Old policies keep appearing |
Use document IDs and versioning |
|
Ignoring access control |
Sensitive data may leak |
Apply authorization-aware retrieval |
|
Assuming vector search is always better |
Exact IDs and codes need keyword search |
Use hybrid search where needed |
Suppose an IT support team wants to find similar old tickets when a new ticket arrives. This helps agents reuse solutions and route tickets faster.
New ticket: “VPN disconnects every 10 minutes after password reset.”
Vector search goal: Find previous tickets with similar meaning, even if wording is different.
Metadata filters: product=VPN, region=India, status=resolved, created_at within last 12 months.
Output: Top 5 similar resolved tickets with resolution notes.
# Pseudo-code: ticket similarity search
new_ticket = "VPN disconnects every 10 minutes after password reset"
query_vector = embedding_model.embed(new_ticket)
results = vector_db.search(
vector=query_vector,
top_k=5,
filter={
"product": "VPN",
"status": "resolved",
"region": "India"
}
)
for ticket in results:
print(ticket.id, ticket.score, ticket.metadata["resolution_summary"])
This is not exactly the same as a chatbot. The LLM may not even be required in the first version. Vector search alone can provide value by finding similar historical cases.
|
Term |
Meaning |
|
Vector database |
A database designed to store and search embeddings using vector similarity. |
|
Vector |
A list of numbers representing the meaning of content. |
|
Embedding |
A vector produced by an embedding model from text, image, audio, or other content. |
|
Vector index |
A data structure that speeds up vector similarity search. |
|
ANN |
Approximate nearest neighbor search, used to find nearly closest vectors quickly. |
|
Metadata |
Additional fields stored with vectors for filtering, security, and traceability. |
|
Collection |
A group of related vectors, often sharing the same purpose and embedding model. |
|
Namespace |
A logical partition inside a vector index or collection. |
|
Top-K |
The number of best matching results returned by search. |
|
Filtering |
Limiting search results using metadata conditions. |
|
RAG |
Retrieval-Augmented Generation, where retrieved context is passed to an LLM. |
|
Re-indexing |
Rebuilding or updating vector indexes after data or model changes. |
1. Design a vector database schema for an HR policy chatbot. Include at least 10 metadata fields.
2. Take one PDF document and decide how you would split it into chunks. Write your chunking rule.
3. Create a table comparing Chroma, FAISS, and PostgreSQL pgvector for a beginner project.
4. Write a pseudo-code search function that retrieves Top 5 chunks using department and country filters.
5. List five risks of using a vector database in an enterprise application and write one mitigation for each.
6. Design a refresh strategy for a company document assistant where policies change monthly.
7. Explain when keyword search may be better than vector search using three examples.
In this mini-project, you do not need to write full production code. Your task is to design the plan for a local vector search prototype.
End of Chapter 11