hyperlakeDiscuss a deployment โ†—
Reference architecture ยท Physical AI

Sovereign Physical AI stack on Kubernetes: a reference architecture

Ten layers, the open components that commonly fill them, how data flows, and the controls that keep the stack sovereign.

Updated
The short answer

A sovereign Physical AI stack on Kubernetes has ten layers, from GPU compute to delivery at customer sites, with identity, policy and audit applied across all of them. Each layer can be filled by open or commercial components. What makes it sovereign is that data, models and keys stay in infrastructure you or your client control.

The stack at a glance

Layer Job Examples of components
1. Compute Run containers on GPU and CPU nodes Kubernetes, with GPU node pools
2. Storage Hold large datasets and artifacts S3-compatible object storage, a shared filesystem
3. Data and catalog Catalog sensor data and make it queryable Apache Iceberg tables with Trino, PostgreSQL for metadata
4. Simulation Test robots and policies in virtual worlds NVIDIA Isaac Sim and Isaac Lab
5. Synthetic data Generate training data beyond what was captured NVIDIA Cosmos
6. Orchestration Run multi-step data, training and evaluation workflows NVIDIA OSMO (open source), or another workflow engine
7. Model serving Serve perception and language models Open model serving on Kubernetes, vector search such as Qdrant or Milvus
8. Evaluation Check policies in simulation and on hardware Software-in-the-loop and hardware-in-the-loop runs
9. Identity, policy and audit Decide who can reach what, and record it An identity provider, OAuth/OIDC, signed tokens, OPA
10. Delivery Ship the result to customer sites and devices Helm packages, GitOps, per-site clusters

Components are examples, not requirements. NVIDIA describes OSMO as an open-source orchestrator for physical AI workflows that covers data generation, training and simulation validation, and says it runs on on-premises Kubernetes, the major clouds and edge hardware.

How does data move through the stack?

  1. Capture. Machines in the field record sensor data and events.
  2. Ingest. Data lands in object storage in the environment that owns it, and is registered in the catalog.
  3. Curate. Teams filter, label and version datasets, and query them with SQL.
  4. Simulate and generate. Simulation and synthetic data fill the gaps that real data cannot.
  5. Train. Orchestrated jobs train policies and models on GPU nodes.
  6. Evaluate. Candidates are tested in simulation, then on hardware.
  7. Package. The approved model and its application are packaged for deployment.
  8. Deploy. The package is rolled out to a fleet, site by site.
  9. Observe. Telemetry and results return to the catalog, and the loop repeats.

What makes the stack sovereign?

Sovereignty is a set of controls, not a product. Five matter most:

  • Location. Compute and storage run where the data owner requires: a region, a data center or a client's site.
  • Keys. Encryption keys and secrets live in a key service or secret store in that same environment.
  • Identity. People and workloads sign in through the owner's identity provider. Avoid shared, long-lived database passwords.
  • Policy at the access point. Access decisions are made by a policy engine, such as OPA, and enforced where data is reached.
  • Audit. Every access and every change is logged, and the log stays in the environment.

Where does Hyperlake fit?

Hyperlake sits across the layers as the control plane. It is not one of the ten layers and it does not replace any of them.

  • Deploy. Cloud-native Helm applications can be deployed through an API with typed configuration fields, previewed first and uninstalled again.
  • Govern. Sign-in goes through an OAuth/OIDC proxy to your identity provider. Cluster tokens are signed with RS256, limited to one cluster and valid for 24 hours, and policy decisions use OPA.
  • Operate. Read-only diagnostics can be run against a cluster or an application, and the same operations are available through an application, a CLI and MCP-compatible tools.
  • Repeat. The same pattern can be deployed again in another environment, so a second site starts from the first.

Hyperlake integrates tools such as NVIDIA OSMO, Omniverse, Isaac Sim and Isaac Lab, subject to licensing and validated integration. Integrations are validated per deployment.

What should you decide first?

  1. Where may the data live? Name the region, data center or site for each dataset.
  2. Who operates each environment? You, the client, or both.
  3. Is it connected? Air-gapped sites change how you deliver updates, images and models.
  4. Where does GPU capacity come from? Your hardware, a cloud account or a partner.
  5. What are the licensing terms? Simulation and generation tools have their own terms; check them for each environment you deploy to.
  6. What repeats? Package the parts that repeat across sites first.

How should you size and phase it?

Start with one cluster in one environment, with separate GPU and CPU node pools. Simulation and training are bursty and benefit from GPU capacity that can be added and released. Serving is steadier and may justify its own pool. Add a second environment only when the first is running from a repeatable package, so the second starts from a pattern and not from a rebuild.

Frequently asked questions

Which components does this reference architecture require?

None in particular. The layers describe jobs, and the tools named in each row are examples. NVIDIA tools such as Isaac Sim and OSMO are common choices for simulation and orchestration, and they carry their own licensing terms.

Can the whole stack run without internet access?

It can be designed to, but each component has to be mirrored inside the site: container images, model weights, packages and updates. Plan for how you will update the stack there before you commit to an air-gapped design.

Where does Hyperlake fit in this architecture?

Hyperlake coordinates the stack. It deploys applications from a catalog, applies identity and policy across them, and gives one control surface for operating them. It does not replace the simulator, the training code or the compute.

How many clusters does a sovereign deployment need?

Often one per environment that has its own data rules, such as your lab and each client's site, deployed from the same pattern. A single environment can start with one cluster that has separate GPU and CPU node pools.

How are secrets and keys handled?

Keep them in a key management service or secret store in the environment that owns the data, and give workloads short-lived, scoped credentials. Hyperlake's cluster tokens, for example, are limited to one cluster and expire after 24 hours.

Sources

Sources for the facts on this page, last checked October 1, 2026.

  1. NVIDIA OSMO checked October 1, 2026
  2. Physical AI and Robotics on Nebius checked October 1, 2026
  3. Hyperlake FAQ checked October 1, 2026

Start with a workload. Build the environment around it.

Tell us what you need to deploy, whose environment it must run in, and what it needs to connect to.