Semantic retrieval
Where the first answer to any intent is built: hybrid search, grounding checks, confidence gates and citation.
See how an answer is builtSemantic cache
Most of what your users ask has been asked before, in other words. Settle the answer once and serve it again without another model call.

How it works
Two steps between a question arriving and a settled answer leaving — neither of them a call to the model.
Each question is grouped under an intent — questions that share a meaning. NeuralSeek generates similar-meaning questions for every one it receives and matches a new phrasing exactly or fuzzily, so an old question in new words finds its intent.
An answer a subject-matter expert has edited becomes eligible for its own cache and is served directly, without going back to generation. Normal answers have a cache of their own, behind a threshold you set. Seek labels the answer Cached.
Supervised, not silent
Curate is where cached answers live: an editorial surface with indicators, filters and search, not an optimisation you discover when an answer is wrong.
Any generated answer can be edited for style and content, and the edit is marked as such. Curated questions and answers can be uploaded in bulk from a spreadsheet, with the option to improve them through Seek first.
Each intent carries indicators: whether it has new answers, whether it contains personally identifiable information, and whether the data it was answered from has gone out of date.
A hash of the source passage travels with each cached answer and is checked against the knowledge base, so a changed document surfaces as an out-of-date flag rather than a stale answer served at scale.
Intents can be searched by keyword and filtered for the ones edited, given a new answer, flagged, or found out of date — so the answers that need a person are the ones a person sees.
Measured
The Governance tab reports the cache alongside the rest of retrieval, filtered by intent, category or date.
Cache efficiency is how often answers came from the cache, from an edited answer, or were generated uncached — a share of traffic, not a promise. Coverage and confidence are charted per intent over a lookback period; drift shows up there.
Cache efficiency

Where it sits
Not a separate product. The cache is what happens between the retrieval layer and the governance layer once an expert has settled an answer.
Separate from all of this is the local cache inside agent workflows: a key-value store an agent reads, writes and searches by index and key, including phonetically. It is a tool for agents, not the answer cache described on this page.
Where the first answer to any intent is built: hybrid search, grounding checks, confidence gates and citation.
See how an answer is builtWhere the cache is measured with everything else: efficiency, coverage, confidence, tokens and cost.
See what is recordedTalk to a NeuralSeek expert about how it fits your stack, where your data has to live, and the governance your auditors already expect.