We worked through the entire path of a GPU workload step by step: instance provisioning, image build, node configuration, and model serving. Blockers at every stage were fixed directly in the product.
The key architectural decision concerned passing GPUs into podVMs. A 20 GB model does not fit into the memory of a single T4 (16 GB). Teaching the adaptor to pass multiple GPUs into one VM would have required significant changes to a security-critical component. So we split the model across two podVMs, each with a single GPU, using KubeRay and tensor parallelism.
Splitting a model across multiple GPUs is a common task in LLM deployment, but here it had to be solved inside podVMs. This turned a platform limitation into an orchestration problem with a proven solution. Distributed serving also gives the platform a path for models that need more resources than a single GPU can provide.
Production workloads with hardware-backed confidentiality guarantees require confidential instances with Hopper- or Blackwell-class GPUs. The N1 + T4 test environment does not provide these guarantees: it lets you verify how the platform stack works, but not memory protection.

Solution Components
Modified Cloud API Adaptor: creates podVMs on N1 + NVIDIA T4 instances for low-cost development and functional testing.
Custom podVM image built with mkosi: NVIDIA drivers, libraries, and initialization services are baked into the image instead of relying on the GPU Operator.
KubeRay and tensor parallelism: the model is split across multiple podVMs, each with a single GPU.
Self-healing containerd DaemonSet: monitors the containerd configuration and restores it after GKE node upgrades.
Upstream PR to Cloud API Adaptor: the ability to create podVMs without public IP addresses.











