Embeddings give us points in space. Now we need somewhere to keep millions of them and a way to ask, in milliseconds, "which stored points are closest to this new one?" A relational database is built to match exact values (WHERE dept = 'HR'), not to rank rows by distance in 1,536 dimensions. That is the gap a vector database fills.
Anatomy of a record
Every chunk becomes one record with four parts:
| Part | Example | Why it is there |
|---|---|---|
| ID | handbook-2026_working-hours_chunk-1 | Uniquely identifies the chunk, and makes updates possible (Module 6) |
| Vector | 1,536 floats | What similarity search compares |
| Text | "Office hours are 9:00 to 18:00…" | What the language model actually reads |
| Metadata | {department: "HR", year: 2026, source_url: "…", section: "Working hours"} | Filtering, citations and versioning |
Beginners often store only the vector and the text. In production, the metadata is where much of the value lies. It lets you search only HR documents, show the user a link to the source, and keep this year's policy apart from last year's. Module 4 covers it in depth. For now, the rule is simple: whatever you might later want to filter on, cite or update by, store it at ingestion time, because adding it afterwards means re-ingesting everything.
Indexes and namespaces
Most vector stores organise records in two levels.
- An index (Pinecone calls it an index, Milvus a collection, Weaviate a class or collection) holds vectors of one fixed dimension compared with one similarity metric. When you create it, you declare both: "1,536 dimensions, cosine". Every vector written to it must match. A model that outputs 3,072 numbers cannot write into a 1,536-dimension index.
- A namespace (or partition) is a subdivision of an index: HR, Finance and Medical in separate namespaces. A query aimed at one namespace never touches the others, which is both faster and a simple way to keep unrelated content from polluting results. Namespaces are also a common way to separate tenants: one namespace per customer, so one customer's documents can never appear in another's answers.
Writing: upsert
Records are written with an upsert, short for update or insert. If a record with the same ID already exists, it is overwritten, and if not, a new one is created. This small detail decides how you will keep the index fresh. If IDs are random, every re-ingest creates duplicates. If IDs are derived deterministically from the document (name + section + chunk number), re-ingesting a changed document overwrites exactly the chunks that changed. Module 6 builds on this.
Reading: query
A query sends a vector and asks for the top-k nearest records, optionally:
- restricted to a namespace,
- filtered by metadata (
department = "Medical" AND year >= 2025), - returning only selected fields (just the text and source URL, not the vector itself).
Each result comes back with a similarity score, such as 0.89 or 0.71. Treat scores as a ranking signal, not as calibrated probabilities. A score of 0.8 means "closer than 0.7", but it does not mean "80% relevant", and the typical score range differs between embedding models.
The landscape
Vector stores come in three broad flavours.
| Flavour | Examples | When it fits |
|---|---|---|
| Managed, serverless | Pinecone, cloud vector services (e.g. S3 Vectors, managed OpenSearch) | You want zero operations and pay per use |
| Self-hosted open source | Milvus, Qdrant, Weaviate, Chroma | Data must stay on-premises or in your own VPC, or you need fine control |
| Vector extension to an existing DB | PostgreSQL with pgvector, MongoDB Atlas Vector Search | Vectors live next to relational data you already run |
They differ in features that matter later in this course: hybrid (keyword + vector) search, advanced metadata filtering, built-in rerankers, partitioning and graph queries. Regulated organisations often run an open-source store inside their own infrastructure precisely because the text of every chunk sits in the database, and that text is the confidential data.
For a corpus of a few thousand chunks, almost any store, or even a brute-force in-memory array, is fast enough. The choice starts to matter at millions of vectors, with heavy filtering, multi-tenancy, or strict data-residency rules. Choose based on those requirements, not on benchmark headlines.