The short answer
Quick answer: A vector database stores embeddings, the lists of numbers that represent the meaning of text, images and other content, and finds the stored vectors most similar to a query vector. Ordinary database indexes are built for exact matches and ranges, not for "closest in meaning" across hundreds of dimensions. Comparing a query with every stored vector is too slow at scale, so vector databases use approximate nearest neighbour (ANN) indexes such as HNSW, which find almost all of the closest matches in milliseconds. AI applications use them for semantic search, recommendations, and retrieving context for language models.
The problem
AI models turn content into embeddings: vectors with hundreds or thousands of numbers, arranged so that similar content has similar vectors. See how embeddings work.
The core operation is nearest neighbour search: given a query vector, find the k stored vectors closest to it.
The straightforward way is to compare the query against every vector. For ten thousand vectors that is instant. For a hundred million vectors of 1,500 numbers each, every query would mean hundreds of billions of arithmetic operations.
Traditional indexes do not help. A B-tree sorts values along one dimension. In hundreds of dimensions there is no useful sort order, and tree-based spatial indexes break down too, a problem known as the curse of dimensionality.
The key idea: approximate search
Exact nearest neighbour search in high dimensions is slow. But applications rarely need the mathematically exact top 10. Getting 95 to 99% of them, a hundred times faster, is an excellent trade.
That is approximate nearest neighbour (ANN) search. The measure of quality is recall: the fraction of the true nearest neighbours that the search returned.
Every ANN index lets you tune a three-way trade-off between recall, speed and memory.
How the main indexes work
HNSW: a navigable graph
Hierarchical Navigable Small World graphs, introduced in this 2016 paper, are the most widely used index.
- Every vector is a node, connected to a number of its near neighbours.
- The graph has layers. The top layer has few nodes with long-range links; each layer below has more nodes and shorter links. The bottom layer contains everything.
- To search, start at the top, hop greedily towards the query, drop down a layer, and repeat.
It works like finding an address: take the motorway to the right city, main roads to the right district, then side streets to the door. Search time grows roughly with the logarithm of the collection size.
Trade-offs: very fast with high recall, but memory-hungry, and building the index takes time.
IVF: search only the nearby clusters
An inverted file index groups vectors into clusters, using an algorithm such as k-means. Each cluster has a centre point.
- To search, find the few cluster centres nearest the query.
- Compare the query only with vectors in those clusters.
Trade-offs: uses less memory than HNSW and is quick to build. Recall depends on how many clusters you probe, and it needs a training step.
Quantisation: smaller vectors
Product quantisation and related techniques compress each vector into a short code, shrinking memory use many times over at some cost in accuracy. Systems often search the compressed vectors first, then re-score the best candidates with the full vectors.
| Index | Speed | Recall | Memory | Notes |
|---|---|---|---|---|
| Flat (brute force) | Slow at scale | Exact | High | Fine for small collections |
| HNSW | Very fast | High | High | The common default |
| IVF | Fast | Good, tunable | Medium | Needs training on the data |
| IVF + quantisation | Fast | Lower | Low | For very large collections |
| Disk-based graphs | Fast | High | Low RAM | Keep most data on SSD |
What a vector database adds on top
An index alone is a library. A database wraps it with the features applications need:
- Metadata filtering. "Find similar documents, but only in English, from this year, that this user may access." Combining filters with ANN search efficiently is one of the hard engineering problems, and a key point on which products differ.
- Hybrid search. Combining vector similarity with keyword search, which catches exact terms such as names and codes.
- Create, update and delete. Graph indexes are awkward to modify in place, so this takes real engineering.
- Storage, durability, replication and sharding, as in any database. See sharding explained.
- Multi-tenancy. Keeping each customer's data separate.
The options
| Category | Examples | When it fits |
|---|---|---|
| Purpose-built vector databases | Pinecone, Weaviate, Milvus, Qdrant, Chroma | Large collections, heavy query load, advanced filtering |
| Extensions to databases you already use | pgvector for PostgreSQL; vector search in Elasticsearch, OpenSearch, Redis, MongoDB | Your data already lives there; moderate scale |
| Libraries | FAISS, hnswlib, ScaNN | Embedding search directly in an application or pipeline |
For many projects, the simplest good answer is pgvector. It adds a vector column type and HNSW and IVF indexes to PostgreSQL, so your embeddings sit next to your ordinary tables. You get joins, transactions and filters with plain SQL, and one less system to run:
CREATE TABLE docs (id bigserial PRIMARY KEY, body text, embedding vector(1536));
CREATE INDEX ON docs USING hnsw (embedding vector_cosine_ops);
SELECT id, body
FROM docs
ORDER BY embedding <=> $1 -- cosine distance to the query vector
LIMIT 5;
A dedicated vector database earns its place when you have many millions of vectors, high query volume, or filtering and tenancy requirements the simpler option cannot meet.
What they are used for
- Retrieval-augmented generation. Find the passages relevant to a question and give them to a language model. See how RAG works.
- Semantic search over documents, products or support tickets.
- Recommendations. "Items similar to this one"; see how recommendation systems work.
- Image and audio search. Find visually similar products or matching sounds.
- Deduplication and anomaly detection.
- Memory for AI assistants. Recalling relevant past conversations or notes.
Things to get right
- Use one embedding model consistently. Query and stored vectors must come from the same model. Changing model means re-embedding everything.
- Match the distance metric to the one the embedding model was trained with (cosine, dot product or Euclidean).
- Measure recall on your own data. Compare ANN results against exact search on a sample.
- Budget memory. A million 1,536-dimension vectors of 32-bit floats take about 6 GB before index overhead.
- Plan filtering early. A very selective filter combined with ANN search can return too few results or slow queries.
- Remember that similarity is not relevance. The nearest vectors are the most similar, which is not always the most useful. Reranking and hybrid search help.
Do you need one at all?
Not always.
- For a few thousand vectors, a brute-force scan in memory is fast and exact.
- If you already run PostgreSQL, start with pgvector.
- If your problem is exact-term lookup, keyword search may be the right tool.
Adopt a specialised system when measurements show you need it.
Frequently asked questions
What is a vector database in simple terms?
A database that stores embeddings and quickly finds the ones most similar to a given query, enabling search by meaning instead of exact words.
What is approximate nearest neighbour search?
A family of algorithms that find most of the closest vectors much faster than checking every one, accepting a small loss of accuracy.
What is HNSW?
A graph-based index in which vectors are linked to their neighbours across several layers, allowing a search to hop quickly towards the query's nearest matches.
Is a vector database the same as an embedding model?
No. The embedding model creates the vectors. The vector database stores and searches them.
Conclusion
Vector databases exist because "find the most similar" is a different problem from "find the exact match", and hundreds of dimensions defeat ordinary indexes. Approximate indexes such as HNSW make similarity search fast, and the surrounding database features make it usable. Start with the simplest option that works, often an extension to the database you already have.
Related articles
- How Embeddings Turn Words Into Numbers
- How RAG (Retrieval-Augmented Generation) Works
- How a Database Index Makes Queries 1000x Faster
- How Recommendation Systems Know What You Want to Watch
