Skip to content

Governance

Prove what it did. Then put it back.

An incident review asks: what did it do, what caused that, and what was it set to at the time. Each has its own record.

Plane one

The answer.

Every question and answer, with the semantic confidence and coverage behind it, kept as a distribution — minimum, average, maximum — rather than an average alone, because an average hides the tail that caused the complaint.

Semantic confidence

console › governance › semantic insights
The Semantic Insights page under Governance: a sidebar grouping the Seek Governance, mAIstro Governance and Custom Governance views, a date-range picker top right, and three rows of tiles — semantic confidence, longest source phrase, top-source coverage, total coverage, answer length, source standard deviation, source jumps, question resolution and cache hit — most with a minimum, average and maximum on a slider, the last two as a split bar.
Semantic Insights: most readings as a minimum, average and maximum over the chosen window; resolution and cache hit as a split.

Read the tiles

Each tile answers one question.

The same page, one reading at a time: what each tile measures over the window you chose, and what a high or low value is telling you.

Semantic confidenceThe semantic score of each answer — how well it aligns with the sources it was built from — kept as the lowest, average and highest across the instance. It is the score the warning and minimum-confidence thresholds can be based on, so a falling average shows up next as questions dropping below the floor.

Plane two

The agent run.

Every execution, per agent: invocations over the window, mean, median, tail and peak latency, guardrail activations, and the tree of child agents it called — so a slow agent is diagnosable, and a run is openable step by step.

Invocations over the window

console › governance › agent details
The Agent Details view under Governance: an agent picker and date range, tiles for invocations, average run time, median, tail and peak latency and guardrail activations, and an invocation tree linking a parent agent to its child agents.
Agent Details: latency per agent, and the invocation tree of the child agents it called.

Plane three

The configuration.

Every change to how the system is set up is logged with the user who made it and shown on a timeline. Any setting can be reverted from the change log: version control over what the AI may do.

Every change, with its author

console › governance › configuration insights
The Configuration Insights view under Governance: a large title above a timeline whose dated axis carries grouped change events.
Configuration Insights: change events grouped along a dated timeline.

What is recorded

The measurements that make drift visible.

A deployment rarely fails suddenly. It degrades: a source goes stale, a question pattern shifts, a model changes underneath you. These readings show it first.

  • Answer quality, as a distribution

    Confidence, coverage, answer length and how far an answer ranged across its sources, each with its minimum, average and maximum, so a bimodal result reads as bimodal rather than as a comfortable mean.

  • Drift, per intent, over a window you choose

    Coverage and confidence broken down by intent across a lookback period you slide. A regression confined to one question type shows as one shifting curve, not a barely moved global average.

  • Which sources are actually answering

    The knowledge base's own confidence and coverage, the documents and URLs most referenced, and user ratings — useful for retrieval tuning, and in an argument about which content is worth maintaining.

  • The terms that keep appearing ungrounded

    The specific terms appearing in answers without support in the sources, ranked — a leading indicator, not an incident report — with allow-listing from the same view when one is legitimate.

  • Tokens, latency and cost

    Input and generated tokens, cost per call, generation speed and tokens over time, with a model cost comparison beside them. The unit economics a finance team asks for, next to the quality data.

  • Who is using it, and how much

    Users and agent growth have their own governance views, so a caller nobody remembers provisioning shows up in the same console as the answers it produced.

Closing the loop

A dashboard that only describes a problem moves the work elsewhere. Each of these ends in a change made where you saw the reading.

  • Fix an ungrounded term where you found it

    A term flagged as hallucinated that is in fact legitimate is allow-listed from the view that surfaced it, and the allowance persists for the instance.

  • Settle a model question with evidence

    Select models, enter a question and compare the results side by side, with token cost and generation speed recorded. A procurement decision has a record behind it, not a memory of a demo.

  • Roll a configuration back

    Open the change log, see which user changed which setting, and revert it. A change that made things worse is reversible from the same screen that showed it.

  • A forensic audit trail

    Every question, answer, score and configuration change logged, timestamped, attributable, and exportable to S3, Splunk or Datadog.

Three questions from the people who have to sign it off.

A model-risk reviewer, a CISO and a finance lead ask different things of the same record.

  1. Does this record live in your console, or in ours?

    Yours, if you want it. Corporate logging streams Seek and Curate activity to an endpoint you run, so the record sits beside your other security telemetry under your own retention policy. No auditor needs a login to a vendor tool.

  2. Something went wrong in March. What can we reconstruct?

    The question and answer with their scores and sources, replayed as they appeared then; which agent ran and which child agents it called; and which user changed which setting, when. Configuration is usually the layer that changes without a trace.

  3. How do you decide which risks to govern in the first place?

    Not from our own list. NeuralSeek starts from an externally published AI risk taxonomy, because a vendor's own register tends to contain the risks it can mitigate. That argument lives on the Responsible AI page; this page is the machinery.

See it running inside your own boundary.

Talk to a NeuralSeek expert about how it fits your stack, where your data has to live, and the governance your auditors already expect.