Accuracy is decided before the model reads a word.
Before the model writes a word, retrieval decides what it is allowed to read. Most teams tune the model and leave retrieval at its defaults — that is backwards.
Every grounded answer is only as good as the material the system was allowed to read. Long before the model writes anything, a retrieval layer decides which documents even enter the conversation — and that quiet upstream decision determines whether the assistant answers from the current policy and the relevant passage, or improvises confidently from noise. Most teams tune the model and leave retrieval at its defaults. That is backwards: retrieval is where the leverage is.
The first decision constrains every later one
When a question arrives, the knowledge base returns a ranked list of candidates. Everything downstream — re-ranking, grounding, confidence scoring, the answer itself — operates only on what that first pass admitted. Let noise through and every subsequent control is spending effort compensating for it. Admit a clean, current, high-signal set and the whole pipeline gets sharper, faster and cheaper at once.
There is a second reason to care, which is who gets to decide. The four families of control below let somebody who understands the domain express a business judgement — how cautious, how current, how concise — as an explicit setting, rather than having it encoded incidentally in application logic by whoever wrote the retrieval call.
Relevance bands: what is eligible at all
The first dial is the band of match scores a document must land in to be considered. Set the floor too low and loosely-related material scrapes in, where the model will faithfully stretch a weak match into a confident-sounding claim. Set it sensibly and the system reasons only over sources that genuinely bear on the question.
The ceiling is less obvious and matters more than it looks. Boilerplate headers, navigation furniture and near-duplicate fragments match mechanically and can dominate a result set while adding nothing. Capping the raw score keeps them from crowding out the passage that actually answers the question.
Tightening the band raises precision and lowers coverage; loosening it does the reverse. There is no correct setting, only a trade-off — and the point of exposing it is that you make it deliberately rather than inherit it.
Freshness: favouring what is current
Knowledge bases accumulate history. Last year's pricing, a superseded policy, a deprecated specification, all still sitting there and all still matching. A retriever that treats every version as equally valid is how an assistant ends up perfectly accurate about a world that no longer exists — the most convincing failure mode there is, because nothing about the answer looks wrong.
Folding recency into the retrieval score makes newer sources rise and stale ones recede without anyone having to prune the corpus by hand. How aggressive the decay should be depends entirely on the material: gentle for reference documentation that ages slowly, steep for pricing and release notes where anything from last quarter is actively misleading.
Document limits: more context is not better context
Pushing more documents into a single call degrades three things simultaneously. Signal-to-noise collapses and the model loses the thread. Latency climbs. And cost grows with every additional token, on every request, forever.
A hard ceiling on how many documents reach the model forces the system to send only its strongest matches. Paired with a reuse window for repeated questions — the same question asked twice does not need to be retrieved twice — you cut noise and spend together, with the freshness of that reuse under your control rather than the cache's.
Snippet sizing: how much of each source
Even the right documents carry dead weight. Snippet sizing governs how much of each source is passed forward: long enough to preserve the passage that answers the question along with the context that makes it interpretable, short enough to keep the token budget disciplined.
Both failure directions are real. Too small and you cut the sentence in half, so the model receives a fragment and completes it from somewhere else. Too large and the answer-bearing line is buried in surrounding prose. This is the last precision dial in the layer, and it shapes not which sources win but how much of each one is actually read.
Tuning them as one system
These are not independent knobs, and tuning them one at a time is how teams end up chasing their own changes. Relevance decides what is eligible; freshness decides which of the eligible win; the document limit decides how many survive; snippet size decides how much of each is read. Change one and the right setting for the others moves.
Tuned together they hand the model a context set that is relevant, current, focused and concise — which is the cleanest foundation a grounded answer can have, and the thing that most reliably improves output quality without touching the model at all.
The other benefit is evidential. Because each is an explicit, versioned setting rather than logic inside an application, 'we think our retrieval is good' becomes a statement of exactly what the model was permitted to read, and why — which is a far better position to be in when somebody asks about a specific answer six months later.
Keep reading
Two more problems, in depth.
Each article is the long version of an argument a platform page makes in a paragraph.
Every article links back to the page that owns its subject.
The questions a healthcare privacy review asks about your AI layer
Four engineering questions that decide whether a clinical project clears review — what reaches the model, what record exists, what the system can see at all, and how any of it is shown.
Building an AI audit trail an examiner will accept
The hardest question is not what the system did but what it was configured to do on the day in question. Four properties that make that answerable, including the one most systems lack.