An embedding is a list of numbers that represents a piece of text, an image, or another object. Similar meanings land close together in that number space. Search then becomes a comparison of distances, not only a match on exact words.
Why keywords are not enough
A keyword search for “late payment” misses a paragraph that says “the borrower fell behind on installments.” An embedding model maps both phrases to nearby vectors because they mean the same thing. Retrieval can then return the second paragraph even though the words differ.
How a vector database is used
- Split documents into chunks small enough to fit in a prompt, often a paragraph or a section.
- Embed each chunk and store the vector with the original text and a source link.
- Embed the user’s question the same way.
- Return the chunks whose vectors are closest to the question.
What to watch
| Choice | Why it matters |
|---|---|
| Chunk size | Too small loses context. Too large buries the relevant sentence. |
| Same embedding model | Queries and documents must use the same model, or distances are meaningless. |
| Metadata filters | Restrict by product, date, or access rights before you rank by similarity. |
| Freshness | Re-embed a document when the source changes. |
A practical example
A policy library holds hundreds of credit procedures. An analyst asks, “What documentation is required for a top-up loan?” The vector search returns the three closest procedure chunks. Those chunks, not the whole library, are what a language model should read before it answers.
What to remember
- Embeddings turn meaning into coordinates so similar items sit near each other.
- A vector database stores those coordinates and finds the nearest chunks.
- This search step is the “retrieval” half of retrieval-augmented generation.