Skip to content

Semantic cache

Serve the answer you already trust.

Most of what your users ask has been asked before, in other words. Settle the answer once and serve it again without another model call.

How it works

Matched, then served.

Two steps between a question arriving and a settled answer leaving — neither of them a call to the model.

  1. Matched to an intent

    Each question is grouped under an intent — questions that share a meaning. NeuralSeek generates similar-meaning questions for every one it receives and matches a new phrasing exactly or fuzzily, so an old question in new words finds its intent.

  2. Served from the cache

    An answer a subject-matter expert has edited becomes eligible for its own cache and is served directly, without going back to generation. Normal answers have a cache of their own, behind a threshold you set. Seek labels the answer Cached.

Supervised, not silent

A cache you can read and edit.

Curate is where cached answers live: an editorial surface with indicators, filters and search, not an optimisation you discover when an answer is wrong.

  • Edited by your expert

    Any generated answer can be edited for style and content, and the edit is marked as such. Curated questions and answers can be uploaded in bulk from a spreadsheet, with the option to improve them through Seek first.

  • Flagged per intent

    Each intent carries indicators: whether it has new answers, whether it contains personally identifiable information, and whether the data it was answered from has gone out of date.

  • Out-of-date answers detected

    A hash of the source passage travels with each cached answer and is checked against the knowledge base, so a changed document surfaces as an out-of-date flag rather than a stale answer served at scale.

  • Filtered and searched

    Intents can be searched by keyword and filtered for the ones edited, given a new answer, flagged, or found out of date — so the answers that need a person are the ones a person sees.

Measured

Governance shows what the cache is doing.

The Governance tab reports the cache alongside the rest of retrieval, filtered by intent, category or date.

Cache efficiency is how often answers came from the cache, from an edited answer, or were generated uncached — a share of traffic, not a promise. Coverage and confidence are charted per intent over a lookback period; drift shows up there.

Cache efficiency

console › governance › semantic insights
The Semantic Insights view under Governance: a grid of metric tiles — semantic confidence, longest source phrase, coverage, answer length, source deviation, question resolution — and a cache-hit tile splitting answers into cached, edited and uncached, over a chosen date range.
Semantic Insights under Governance: the cache-hit split — cached, edited, uncached — in the same grid as coverage and confidence.

Where it sits

Seek answers. Curate settles. Governance measures.

Not a separate product. The cache is what happens between the retrieval layer and the governance layer once an expert has settled an answer.

Separate from all of this is the local cache inside agent workflows: a key-value store an agent reads, writes and searches by index and key, including phonetically. It is a tool for agents, not the answer cache described on this page.

  • Semantic retrieval

    Where the first answer to any intent is built: hybrid search, grounding checks, confidence gates and citation.

    See how an answer is built
  • Governance

    Where the cache is measured with everything else: efficiency, coverage, confidence, tokens and cost.

    See what is recorded

See it running inside your own boundary.

Talk to a NeuralSeek expert about how it fits your stack, where your data has to live, and the governance your auditors already expect.