A question about disease prevention will almost never be answered by the finance report, but a pure similarity search does not know that. It compares the question with every chunk and may well rank a finance paragraph about "preventing losses" above a weaker match in the medical guide. Worse, a question about the leave policy may retrieve both the 2025 and the 2026 versions, with different numbers, and the model then has to guess which one is current.
Similarity is a measure of meaning. It knows nothing about department, year, document type, region, language or access rights. Those are facts about a chunk, and they belong in its metadata.
What to store
When you ingest a document, attach structured fields to every chunk it produces:
| Field | Example | Enables |
|---|---|---|
department | HR, Finance, Medical | Scoping a query to the right domain |
doc_type | policy, FAQ, contract, runbook | Answering "which policy…" from policies only |
year / effective_date | 2026 | Current-version answers, historical queries |
region / language | IN, EU / en, hi | Jurisdiction-specific answers |
source_url, title, section | link, "Leave Policy", "§4.2" | Citations in the answer |
doc_id, chunk_index | handbook-2026, 7 | Updates, deletions, fetching neighbouring chunks |
access_group / tenant_id | hr-managers / acme | Security filtering |
All chunks from one document inherit that document's fields, so if a PDF becomes five chunks, all five carry the same source_url.
Filtering the search
A filtered query says: "find the top-k most similar chunks, but only among those whose metadata matches this condition." Filters read much like a SQL WHERE clause: equality (department = "Medical"), sets (department IN ["HR", "Legal"]), ranges (year >= 2025), and combinations with AND and OR. Most databases also let you choose which fields come back, so you can return only the text, department and year rather than everything.
There are two ways an engine can apply a filter, and the difference matters.
- Post-filtering: find the top-k by similarity first, then drop the ones that fail the filter. It is simple, but if the filter is selective you may end up with 1 result instead of 5, or none at all, because the top 5 overall were all finance chunks.
- Pre-filtering (filtered search): restrict the candidate set first, then rank by similarity within it. You always get k matching results. Modern vector databases integrate the filter into the index traversal itself so that this stays fast.
When filters are highly selective, make sure your database performs true filtered search. Otherwise "no results" bugs will appear only for rare departments.
Where do filter values come from?
- From the application. The user is in the HR portal, or logged in as tenant
acme, so the server attaches the filter. This is the safest source, and the only acceptable one for security filters. - From the query, via an LLM. An LLM reads "what was the travel allowance in 2024?" and extracts
{year: 2024, doc_type: "policy"}as a structured filter before searching. This is often called self-querying. It is powerful, but validate the extracted values against the allowed list. - From a routing step. A classifier decides that a question is medical, and the search is scoped to the medical namespace or filter (Module 7 develops routing fully).
Three production uses
1. Scoping. Search only where the answer can be. Fewer candidates mean fewer distractors, better precision and lower latency. Namespaces are the coarse version (a separate partition per domain), and metadata filters are the fine-grained version.
2. Versioning. When the 2026 policy replaces the 2025 one, you have two choices. You can overwrite the old chunks (Module 6), or you can keep both with a year field and filter for the latest by default, while still answering "what was the policy in 2025?" when asked. Keeping history together with filters is often the better choice in regulated settings, where you must be able to show what the rule was on a given date.
3. Citations. Store each document's source_url (and title and section) in every chunk's metadata. When chunks are retrieved, pass the metadata to the model along with the text and ask it to cite them, or attach the links yourself. Users can then click through to the original document, which builds trust and lets them check the answer, and reviewers can trace any answer back to its evidence.
If access control relies on a filter such as access_group IN user.groups, that filter must be added by trusted server code from the authenticated identity. Never build it from user input or from an LLM's output. A prompt-injected or confused model must not be able to widen its own search.