Guardrails
Enforced at runtime, not written into a prompt.
A prompt that asks a model to behave is a request. A guardrail runs on every call, cooperative model or not, and leaves a record.

The path a request takes
A fixed order, and each stage reports.
Every request walks all of these, in this order. A check that runs after generation, or that a client can skip, is not a control.
Screening on the way in
Input is scored against a prompt-injection model. Above one threshold the offending phrases are stripped; above another the request is blocked outright. A blocked-word list of your own, and a profanity filter, sit beside it.
Sensitive data, before the model
A pattern pass and a model-based pass run over the input before anything reaches a language model. What is found is flagged, masked, hidden or deleted — your choice, and per intent.
Intent classification and routing
The request is categorised and routed through the configuration tree. Each category can carry its own configuration and its own Guardrails node, so a billing question and a clinical one need not be governed identically.
Retrieval under thresholds
A document score range, a date penalty and a cap on documents per call decide what the model is allowed to see. That is a governed decision, not a side effect of how the index scored.
Generation under limits
Model choice, token minimums and maximums, temperature and sampling bounds are set per node, so a classifier step and a synthesis step are not forced to share settings. Input length is bounded too.
Checks before release
The semantic score gates the answer: a warning caveat above one line, substitution or hand-off below another, sentences with ungrounded key words removed, and an attribute-protection scale for claims about your own organisation.
Session and context scope
How many conversation turns are carried to the model, and whether context is forced onto a follow-up, are settings. Context carried between turns is a governed surface, because it is where data quietly accumulates.
Recorded and exportable
Every question, answer and score is logged with its intent, and the stream can be sent to a logging endpoint you run. Full prompt capture is opt-in, behind an agreement.
The same guardrails, mid-flow
Guardrails are nodes too.
The controls above are also steps you place inside an agent. Protect, Profanity Filter and Remove PII act on whatever text flows into them — Protect's thresholds and the PII filters come from the Configure tab — so here the ticket is screened, and a blocked one stopped, before the first model call.

Hallucination
Four layers, because one is a slogan.
Each catches what the one before it cannot, and each reports — when an answer is stopped, you can see which check stopped it.
Constrain the perimeter
Retrieval is bounded to your knowledge base, with a stump-speech document included in every call as fallback ground truth. The cheapest hallucination to prevent is one about a subject the deployment was never meant to discuss.
Enforce grounding sentence by sentence
Re-ranking by semantic match, penalties for missing search terms and for answers that jump across sources, and removal of sentences whose key words the knowledge base does not contain.
Refuse below a floor
A warning threshold, a minimum threshold and a separate threshold for showing a link. Below your floor the platform substitutes your reply text rather than hedging, and logs the decline.
Route what is left
A low-confidence answer is a workflow trigger, not a dead end: a fallback template can notify a team, open an issue, rewrite, or hand to another agent. Nothing ships because there was nothing else to do with it.
Sensitive data
Two passes, and you choose what happens next.
Detection is half of it. The action is set per intent, so one deployment can mask an identifier in one context and refuse in another.
A pattern pass and a model pass
Pattern filters catch the structured identifiers; a model-based pass catches what patterns miss, taught from an example sentence rather than a rule. Both run before the model, not after.
Allow-lists for known-safe terms
Every deployment has terms a detector flags that a human knows are fine. Mark them as not PII, or trust words found in your own source documents, so tuning does not mean switching a detector off.
Secrets kept out of the flow
Credentials an agent needs are stored as named secrets in configuration and referenced by name inside a flow, so a key never appears in the template a reviewer reads.
PII detection and redaction
13 detector categories and five enforcement actions — mask, flag, hide, delete, or pass through with a warning — configurable per intent.
A forensic audit trail
Every question, answer, score and configuration change logged, timestamped, attributable, and exportable to S3, Splunk or Datadog.
Adversarial testing
Run the attack suite yourself.
Red Team Testing ships in the product. Pick an agent, run the test, and read back its security posture, the risks identified, the evidence and the recommended mitigations — repeatable after every change instead of once before launch.
Prompt injection

Three questions a security reviewer asks.
None of them is answered by a longer list of controls, so they are answered here instead.
Can a developer bypass these by calling the model directly?
Not through NeuralSeek: every call through the platform walks the same stages, whichever client made it. A call that never goes through it is a network question — one reason the container runs inside your boundary, under your egress controls.
Won't guardrails this strict make it useless?
They would if they were global. Thresholds live per category in the configuration tree, so a general question and a regulated one can share a deployment without sharing a floor. You tune against your own score distributions.
What do we actually hand an auditor?
The record. Questions, answers and scores are logged, exportable to your own endpoint, and replayable as they appeared at the time. Configuration changes carry who made them and can be reverted, so the auditor's question is a lookup.
See it running inside your own boundary.
Talk to a NeuralSeek expert about how it fits your stack, where your data has to live, and the governance your auditors already expect.