Vector Databases and Hybrid Search in RAG
A vector database stores your embedded chunks and finds the ones nearest to a query vector, fast — usually by cosine similarity, across millions of chunks in milliseconds.
Hybrid search combines that semantic search with keyword search, so you get meaning and exact terms. For most real corpora it's the first and biggest upgrade past naive retrieval.
What a vector database does
A vector database stores your embedded chunks and answers one question extremely quickly: given this query vector, which stored vectors are nearest? "Nearest" is usually cosine similarity, which compares the angle between vectors and so measures similarity of meaning regardless of length. The query returns the top matches, and those become your candidate chunks.
Why not just compare everything?
With a few thousand chunks you could compare the query against every one — brute force — and it would be fine. With millions, that's too slow per query. Vector databases use approximate nearest-neighbor algorithms that trade a tiny bit of accuracy for enormous speed, finding almost-certainly-the-closest vectors in milliseconds. For RAG, that trade is almost always worth it.
What to actually store
A vector database entry is more than a vector. Store the embedding, the original chunk text, and its metadata — source, section, date, tags you'll filter on. The text is what you feed the model; the metadata is what lets you restrict a search and cite the result. A vector with no text attached retrieves nothing useful.
index.add(
vector=embed(chunk.text),
text=chunk.text,
metadata={'source': chunk.doc, 'date': chunk.date},
)
# query with a metadata filter
hits = index.search(embed(question), top_k=20,
filter={'date': {'gte': '2026-01-01'}})
Hybrid search: the first big upgrade
Pure semantic search is strong on meaning and weak on exact terms — product codes, names, error strings, rare jargon. Pure keyword search is the opposite. Hybrid search runs both and combines the results, so a query mentioning a specific error code and a general concept gets the best of each. For most real corpora, hybrid search is a clear, reliable win over either alone, and it's the first upgrade to make past naive retrieval.
Semantic search finds meaning.|Keyword search finds exact terms.|Hybrid search finds both.
Filtering is a superpower
Metadata filters narrow the search before similarity is even computed — by permissions, recency, source, or type. This is how you enforce access control cheaply (only retrieve what a user may see), keep answers fresh (only recent documents), and stop retrieval from wandering into irrelevant corners of the corpus. Structure you already have is retrieval signal you're not yet using.
Choosing a vector store
The market offers many vector databases, from in-process libraries to managed cloud services, and for most projects the choice matters less than it seems. The fundamentals — store vectors, filter on metadata, return nearest neighbors fast — are table stakes everywhere. What differs is operational: how they scale, handle updates, and cost at your volume. Pick for your operational reality, and know the simple interface is easy to swap later.