Before anything reaches a model you have to answer one question with confidence: does this contain personal data we are not permitted to send? The instinct is to scan for it with pattern matching, which is fast, cheap, deterministic and has been the standard for decades. It is also, used alone, quietly incomplete — and the failures are exactly the cases that turn into an incident.
What pattern matching does extremely well
Regular expressions are excellent at finding personal data with a predictable shape. A national insurance or social security number has a fixed digit grouping. A payment card is a fixed length that passes a checksum. Email addresses, telephone numbers, IP addresses, account references and postal codes all follow rigid formats that a regular expression matches in microseconds.
For these, pattern matching is fast, deterministic, auditable and effectively free — and the determinism matters as much as the speed. There is no reason to spend a model call establishing that a string formatted like a social security number is a social security number, and there is every reason to prefer a control whose behaviour you can reason about exactly.
Where it fails, quietly
Most real personal data does not arrive pre-formatted. A name is just words. A place is just a place. Consider two phrases that identify a specific human being with total reliability and contain no matchable pattern at all: 'Maria in room 4B', and 'the CFO who joined us from the Reno office last spring'.
A regular expression cannot reason about meaning, so it sails past anything contextual — names it has not been given, job titles bound to an individual, indirect identifiers, paraphrased account references, and identifiers assembled across a sentence rather than sitting in one token. It also produces false positives in the other direction, flagging any nine-digit number as an identifier, and it breaks the moment data arrives formatted in a way nobody anticipated.
The failure is silent, which is the worst property a privacy control can have. Nothing is flagged, nothing is logged, and the pipeline reports success.
What a contextual pass adds
A model-based detector reads text the way a person does. It can recognise that 'the patient', 'room 4B' and a first name together constitute identifiable health information even though no single token matches anything. It catches names it has never encountered, notices when a sentence is describing one specific individual rather than a category, and adapts to phrasing no rule anticipated.
The cost is that it is slower and probabilistic. It reasons rather than matches, which means you would not want to run it over every byte of every request if a cheap deterministic rule has already settled the question — and it means its output is a judgement, which has to be treated as one.
The answer is both, in order
The two approaches fail in opposite places, which is precisely why they belong in sequence rather than in competition. A fast deterministic pass runs first and removes the structured identifiers cheaply and predictably. A contextual pass then runs over what is left and catches the names, places and paraphrased references the patterns could not see.
Pattern matching supplies volume and precision; the contextual pass supplies nuance and recall. Layered, they close a gap neither closes alone, and both run before a single sensitive token leaves your boundary — which is the only ordering that helps, because a redaction applied after the model call has already failed.
Detection is half the control; the other half is what happens next
A detector that finds something and does nothing consistent about it is not a control. The second decision is policy: on a match, is the value masked, replaced with a stable token so the sentence still parses, removed entirely, or is the whole request refused? Different answers are right for different intents, which is why it is worth configuring per workflow rather than globally.
Two adjacent controls matter more than they look. First, the ability to define what counts as sensitive in your domain, because a term that is innocuous in one industry is an identifier in another. Second, keeping credentials and connection secrets out of model traffic entirely — because if a key can reach a prompt, it can reach a log, and the log is the artefact you hand to an auditor.
And because every detection and every action is recorded, the claim changes shape. 'We believe our data is safe' becomes a record showing what was detected, what was done about it, and when.
Keep reading
Two more problems, in depth.
Each article is the long version of an argument a platform page makes in a paragraph.
Every article links back to the page that owns its subject.
Running language models in a fully air-gapped environment
Mirrored registries, in-cluster inference on open weights, a retrieval stack that never leaves the perimeter, and the hardening that makes the isolation provable rather than asserted.
Relevance bands, freshness decay, document limits and snippet sizing. The upstream decisions that set the ceiling on accuracy, and why tuning them one at a time does not work.