A 20 GB Model, Two T4s, and Not a Single H100: GPUs for Confidential Containers 

We added GPU support to podVMs, tested distributed model serving with KubeRay in an NVIDIA T4 test environment, and removed manual containerd configuration from GKE installation.


Confidential-computing platform vendor (NDA)

A Kubernetes platform that runs pods inside hardware-isolated VMs. Their memory is encrypted and inaccessible even to the cloud provider. Managed Kubernetes across multiple clouds, with GKE as the primary one. Test: a 20 GB Mistral model on 2 podVMs with NVIDIA T4.

TEAM

1 Alpacked Platform Engineer embedded in the client's product team.

PERIOD OF COLLABORATION

~2 months (August–September 2026), engagement ongoing.

CLIENT’S LOCATION

Confidential AI Without GPU Support

The client's platform ran conventional workloads in confidential containers but did not support GPU-accelerated AI, even though this is its key commercial focus. There was no way to demonstrate a distributed GPU scenario, and testing on H100/B200 was too expensive.

Three factors stood in the way of running GPU workloads:

Cloud API Adaptor limitations. The adaptor supported only expensive GPU instances and could not pass multiple GPUs into a single podVM.

Incompatibility with the NVIDIA GPU Operator. PodVMs are created outside the Kubernetes node lifecycle, so the operator cannot see them and does not install drivers.

Containerd configuration on GKE. Starting with version 1.27, containerd discards unpacked image layers, and pods in podVMs fail to start. The manual fix was reset after every node upgrade.

Key Results in Numbers

icon

20 GB – LLM on two podVMs with NVIDIA T4 (test environment)

icon

H100/B200 → T4 – cheaper functional GPU testing

icon

1 upstream PR – podVMs without public IP addresses

icon

0 – manual containerd changes during GKE installation


The Core Challenge: GPUs in podVMs Without Compromising Security

The platform could not run GPU workloads in podVMs, and installing on GKE required manual work on the nodes. At the same time, every change touched the security boundary. Solutions could not weaken isolation or attestation, and they had to become part of the product rather than one-off scripts. The problem came down to three factors:

1

Challenge #1: The Stack Wasn't Ready for GPUs in podVMs.

The adaptor worked only with expensive GPU instances and could not pass multiple GPUs into a single podVM. The GPU Operator could not see podVMs, so there was no standard way to install drivers. The process for building custom podVM images was undocumented, and every engineer had to figure it out on their own. On top of that, any drivers in the image became part of an environment that has to stay minimal and auditable. 

2

Challenge #2: Settings That Vanished After Upgrades.

After a node upgrade or restart, GKE reverted containerd to its default configuration, and pods in podVMs stopped starting. Linking the failure to the upgrade was not easy.

3

Challenge #3: Public IP Addresses for VMs.

The adaptor created VMs with public IP addresses, which expanded the network attack surface of a product that sells a secure environment.

The Solution: Distributed Serving Instead of Multiple GPUs in One VM

We worked through the entire path of a GPU workload step by step: instance provisioning, image build, node configuration, and model serving. Blockers at every stage were fixed directly in the product.

The key architectural decision concerned passing GPUs into podVMs. A 20 GB model does not fit into the memory of a single T4 (16 GB). Teaching the adaptor to pass multiple GPUs into one VM would have required significant changes to a security-critical component. So we split the model across two podVMs, each with a single GPU, using KubeRay and tensor parallelism.

Splitting a model across multiple GPUs is a common task in LLM deployment, but here it had to be solved inside podVMs. This turned a platform limitation into an orchestration problem with a proven solution. Distributed serving also gives the platform a path for models that need more resources than a single GPU can provide.

Production workloads with hardware-backed confidentiality guarantees require confidential instances with Hopper- or Blackwell-class GPUs. The N1 + T4 test environment does not provide these guarantees: it lets you verify how the platform stack works, but not memory protection.

Solution Components

Modified Cloud API Adaptor: creates podVMs on N1 + NVIDIA T4 instances for low-cost development and functional testing.

Custom podVM image built with mkosi: NVIDIA drivers, libraries, and initialization services are baked into the image instead of relying on the GPU Operator.

KubeRay and tensor parallelism: the model is split across multiple podVMs, each with a single GPU.

Self-healing containerd DaemonSet: monitors the containerd configuration and restores it after GKE node upgrades.

Upstream PR to Cloud API Adaptor: the ability to create podVMs without public IP addresses.

Technical Details: podVM Image, KubeRay, and containerd

Approaches That Didn't Work

NVIDIA GPU Operator. The operator works only with regular Kubernetes nodes and cannot see podVMs. This is an architectural trait, not a bug that can be fixed. So the drivers had to be built directly into the image, and the image build became the core of the work.

Multiple GPUs in one VM. The model could have been placed on a VM with several GPUs, but the adaptor did not support passing multiple GPUs into a single podVM. So we used the KubeRay-based distributed serving described above.

Manual containerd patch. The fix held only until the first GKE node upgrade, after which the failure came back. So we dropped the one-time installation step: the containerd configuration is now continuously maintained by a DaemonSet.

Working without documentation. The podVM image build process wasn't documented anywhere, and every attempt started from scratch. That's why the build guide became a deliverable in its own right.

The Hardest Part: A podVM Image as Part of the Security Boundary

On a regular Kubernetes node, installing GPU drivers is a routine operation. In a production configuration, the image becomes the foundation of a confidential VM, so everything baked into it ends up inside the trust boundary. That meant the build had to be more than functional: it had to be reproducible, minimal, and auditable.

NVIDIA drivers, libraries, and initialization services were placed into the image via a skeleton filesystem. This is a set of directories and files that mkosi copies into the image during the build. As a result, the services start correctly without external configuration management. There were no existing examples of such a build, so it took far more iterations than similar work on a regular node.

The second challenge was making changes persist in managed Kubernetes. The provider manages the nodes, there is no persistent host access, and upgrades reset the local configuration. So we turned the containerd settings into a continuously maintained state rather than a one-off change.

Custom Solutions

Cloud API Adaptor fork: creates smaller, cheaper instance types based on a predefined configuration. This allows GPU pods in podVMs to run on N1 with NVIDIA T4. 

podVM build with mkosi: drivers and their initialization services are added to the image via a skeleton filesystem.

Containerd Remediation DaemonSet: deployed via a label selector during installation. In a watch loop, it tracks the state of each worker node and reapplies the required containerd configuration if it has changed. The installer automatically detects the environments where the DaemonSet is needed.

Upstream PR: a change to Cloud API Adaptor that allows creating VMs without public IP addresses was accepted into the open-source project.

podVM image build guide: common pitfalls and recommended approaches for engineers who will build their own images.

Tech Stack

Core: Kubernetes, GKE (Google Cloud), Confidential Containers, Cloud API Adaptor, podVM, containerd, DaemonSets, Helm, Docker/OCI registries, NVIDIA Drivers & Container Toolkit, NVIDIA T4 / N1 instances, KubeRay, Ray, Mistral, Go.

Specialized: mkosi (podVM image builds), skeleton filesystem injection (baking drivers and services into the image), tensor parallelism across single-GPU podVMs, label-selected DaemonSet reconciler (continuous reconciliation of node configuration).


Results

The product gained the changes needed to run GPU workloads in podVMs, and the distributed serving scenario was validated in a test environment. Functional GPU testing became cheaper, and the described cause of failures after GKE upgrades was eliminated.

1

Installation Automation

Installing on GKE no longer requires manually changing the containerd configuration on every node and repeating it after upgrades. The installer detects the environment and deploys the DaemonSet on its own. We apply the same principle in our Kubernetes consulting for managed clusters: node configuration is maintained by the cluster itself, not by one-off manual changes. The podVM image build documentation saved every engineer days of figuring things out on their own. 

3

Stability and Security

The DaemonSet restores the required containerd configuration after node upgrades, eliminating the described cause of podVM startup failures. Previously, VMs were created with public IP addresses, which expanded the network attack surface. The accepted upstream change made it possible to create them without public addresses.

2

Development Costs

Development and functional testing of GPU scenarios moved from H100/B200-class instances to N1 with NVIDIA T4, significantly reducing the cost per GPU hour. Choosing instance types for a specific task is part of our Google Cloud consulting. More importantly, testing that used to be skipped because of the cost now actually happens, and the path from idea to validation has become shorter. In addition, the team no longer spends engineering time on manual node configuration.

4

A Foundation for Further Platform Development

The team gained a reference scenario for distributed GPU serving, validated in a test environment, and a foundation for applying it on compatible hardware. Custom podVM images are now built following a documented process.

Let's arrange a free consultation

Just fill the form below and we will contaсt you via email to arrange a free call to discuss your project and estimates.