Skip to content

Deployment · 4 min read

No outbound call. Not one, not ever.

For defence, intelligence, government and parts of regulated finance, the cloud is not an option. This is the walkthrough of running AI with no path to the internet.

For organisations with the strictest data requirements — defence, intelligence, parts of government, and some regulated finance — the cloud is not an option and no amount of contractual assurance changes that. Sensitive material cannot leave the network, which rules out every API-only model and most managed platforms. Air-gapped deployment is the answer, very few platforms genuinely support it, and almost nobody has written down how to do it properly. This is that walkthrough.

What air-gapped actually means

Air-gapped means the deployment has no path to the public internet. No model API calls, no telemetry, no licence check, no package downloads at runtime, no exceptions. Everything the system needs is mirrored inside the perimeter in advance: container images, model weights, embedding models, and every dependency.

This is materially stricter than 'on-premises' or 'private cloud', both of which usually retain some outbound path and often depend on one. It is a hard isolation boundary, and the reason it is worth the difficulty is that it is provable — you can demonstrate it to an auditor with a network policy rather than describe it with a diagram.

The consequence for architecture is that isolation has to be the first constraint rather than a later configuration. Almost every design decision below follows from it.

Orchestration: the foundation

Start with what schedules and scales the workloads. Containers package each component into a portable image; Kubernetes, or OpenShift as its hardened enterprise distribution, orchestrates those images across your own servers and GPUs.

The critical detail is the registry. You run a private container registry inside the perimeter and mirror every image into it, so no pod ever attempts an external pull. This sounds trivial and is where most air-gapped attempts first fail, usually via a base image or a sidecar somebody did not know was there. OpenShift earns its place in regulated environments by defaulting to restricted security contexts, integrated image signing and built-in policy, which means the hardening below starts from a better position.

Model serving: open weights, in cluster

Because there is no hosted model to call, you serve open-weight models yourself. An inference server such as vLLM or Hugging Face's Text Generation Inference runs as an in-cluster service with GPU scheduling, request batching and autoscaling.

Weights are loaded from on-premises storage staged ahead of time and never fetched at runtime. Size the GPU pods to the model and the concurrency you actually need rather than to the largest model available, and expose inference only on the internal cluster network so nothing is reachable from outside the perimeter even in principle.

Retrieval that never leaves

A model is only as useful as the material it can ground against, so the retrieval stack has to be inside the perimeter too: a vector store, an embedding model served in-cluster, and an ingestion pipeline that indexes your documents.

Every step — embedding, storage, query — happens inside the boundary, which means the source material is never transmitted anywhere at any point. This is the difference between a generic offline chatbot, which is of limited use, and a system that can answer from classified or regulated material, which is the reason the deployment exists.

Hardening: making it defensible rather than merely isolated

Isolation gets you most of the way. Hardening is what lets you prove it, line by line, to somebody whose job is to disbelieve you.

  • Constrain east-west traffic with network policies so only the components that must talk to each other can, and so a compromised pod cannot reach the model or the index.
  • Enforce image provenance with signing and admission control, so an unverified image cannot run even if it reaches the registry.
  • Hold secrets in a sealed in-cluster store rather than in environment variables or manifests, where they end up in logs and backups.
  • Turn on comprehensive audit logging so every inference and every configuration change is attributable to an identity and a time.

The layer most guides stop before

Running the model is the infrastructure problem, and it is the easier one. Governing it does not become unnecessary because the network is isolated — an air-gapped assistant can still answer confidently from a weak retrieval, still surface one business unit's material to another, still put a credential in a log, and still leave you unable to reconstruct what it did last March.

So the same controls apply inside the perimeter as outside it: a confidence floor below which the system declines rather than answers, corpus isolation between tenants or units, redaction, and a complete interaction log. NeuralSeek deploys into the same cluster through the same container and supplies that layer without a single outbound call, including the ability to compare the open-weight models you have staged against your own material rather than against a public leaderboard.

The infrastructure makes air-gapped AI possible. The governance is what makes it safe to actually put in front of someone.

Keep reading

Two more problems, in depth.

Each article is the long version of an argument a platform page makes in a paragraph.

Every article links back to the page that owns its subject.

  • Controlling what your AI is allowed to read

    Relevance bands, freshness decay, document limits and snippet sizing. The upstream decisions that set the ceiling on accuracy, and why tuning them one at a time does not work.

    Read on retrieval grounding
  • The questions a healthcare privacy review asks about your AI layer

    Four engineering questions that decide whether a clinical project clears review — what reaches the model, what record exists, what the system can see at all, and how any of it is shown.

    Read on healthcare reviews

See it running inside your own boundary.

Talk to a NeuralSeek expert about how it fits your stack, where your data has to live, and the governance your auditors already expect.