Skip to content

Healthcare · 4 min read

Clinical AI does not die at the model. It dies in review.

Healthcare AI projects rarely stall on the model. They stall in privacy review — on four engineering questions, each of which can become a setting you can point at.

Most healthcare AI projects that stall do not stall because the model was not good enough. They stall in privacy and legal review, where the questions are engineering questions and the project has no engineering answers to them. This is a walkthrough of the four that come up almost every time, and of how each one becomes a setting somebody can point at, test and defend — rather than an assurance somebody gives in a meeting.

Scope note: this is engineering guidance about controls, not legal advice and not a statement of what any regulation requires. HIPAA is named here because it is the framework these reviews are conducted under, and its text — together with your own counsel and privacy office — governs what applies to you. NeuralSeek is deployed inside your environment, which means the compliance position is yours; nothing described here is a certification held by us.

Why the review is an engineering conversation

The common assumption is that the regulation forbids putting clinical data anywhere near a language model, and so the project waits for a legal answer that never quite comes. In practice the review is rarely asking for permission in the abstract. It is asking four concrete things: what reaches the model, what record exists afterwards, what the system is able to see at all, and how any of that can be demonstrated.

Each of those is answerable in configuration, and a project that arrives with the answers configured has a very different review from one that arrives with a demonstration and a promise.

One: what actually reaches the model?

The strongest possible answer to this question is 'none of it', and it is achievable more often than teams expect. If protected information is removed before the prompt is assembled, the question of what the model provider does with it stops being interesting.

The robust setups redact in two passes for the reasons set out in the article on PII detection. A fast deterministic pass removes the structured identifiers — record numbers, dates, contact details — at negligible latency. A contextual pass then catches what patterns cannot: a name in a free-text clinical note, a location phrased a dozen ways, a description that identifies a person without naming them.

The third decision is what happens on a match, and it is worth setting deliberately per workflow: mask the value, substitute a stable token so the surrounding sentence still makes sense, remove it, or refuse the request. Set once, applied consistently, and visible in the configuration rather than in someone's code.

Two: what record exists afterwards?

You cannot account for access you did not record, and a review will assume nothing about what your system did. Two complementary records do most of the work: a durable, tamper-evident organisation-wide log of every interaction, and a record of the prompts themselves so that a question raised months later can be reconstructed with the context the system was actually working from.

The practical test is whether producing that record is an export or a project. If a reviewer asks what a specific assistant said to a specific person on a specific day and the answer takes a week to assemble, the control does not really exist yet.

Three: what can the system see at all?

Redaction and logging matter little if a retrieval can reach material it should never have been able to reach. This is a boundary question rather than a login question: an assistant built for one department should be structurally incapable of retrieving another department's records, not merely discouraged from doing so.

That is enforced at the retrieval layer, by constraining which corpus each tenant or business unit can query, before any ranking happens. Alongside it, credentials and connection secrets should be kept out of prompts and responses entirely — otherwise the evidence file you produce for a reviewer becomes a place a key can leak from.

Four: how do you show it?

The last question is the one that decides how long the review takes. A configuration that is correct but undocumented gets treated as an assertion. The organisations that clear review quickly are the ones that can put each control next to the framework clause it is there to satisfy — commonly ISO 42001 or the NIST AI Risk Management Framework, both of which give a reviewer a familiar structure to read your controls in.

That mapping is the difference between a security questionnaire that becomes a months-long evidence hunt and one that is answered from a document. It is also, usefully, the artefact that survives staff turnover.

Where the responsibility actually sits

One thing worth being explicit about, because vendors are routinely vague on it. NeuralSeek ships as containers that run inside your environment. The compliance perimeter is yours, the data never has to leave it, and the controls above are ours to provide and yours to configure and operate.

SOC 2 Type II is attested at the vendor level and the report is available under NDA. Every other framework named on this site — including the ones in this article — is something the product is built to meet inside your environment, not a certification we hold. The distinction matters more in healthcare than almost anywhere, and a vendor that blurs it is telling you something.

Keep reading

Two more problems, in depth.

Each article is the long version of an argument a platform page makes in a paragraph.

Every article links back to the page that owns its subject.

  • Building an AI audit trail an examiner will accept

    The hardest question is not what the system did but what it was configured to do on the day in question. Four properties that make that answerable, including the one most systems lack.

    Read on audit trails
  • Prompt injection: direct, indirect, and how to contain both

    Why it cannot be patched at the model layer, what an indirect payload hidden in a document actually looks like, and why the containment boundary has to sit outside the thing being attacked.

    Read on prompt injection

See it running inside your own boundary.

Talk to a NeuralSeek expert about how it fits your stack, where your data has to live, and the governance your auditors already expect.