Managed Kubernetes is in beta. It works, teams are running on it, and the surface is still changing. Expect the feature list below to grow, and tell us what’s missing.
kubectl as you.
What you get
- A hosted control plane, run and kept available by Hyperbolic, located in a cloud region close to your GPU nodes.
- GPU worker nodes that are your rented instances — same hardware, same networking, same billing as renting them directly.
- NVIDIA GPU Operator and Network Operator preinstalled, so GPUs, drivers, and RDMA networking are exposed to pods on day one.
- Single sign-on access. You download a per-cluster kubeconfig; the first
kubectlcommand signs you in through Hyperbolic in your browser and authenticates as you — short-lived, with every action attributed to your identity instead of a shared admin credential. Org admins getcluster-admin; other members getedit. RBAC inside the cluster is yours to extend. See Connect with kubectl. - Multi-node performance without overhead. NCCL all-reduce inside a pod matches the same test run on the bare host.
Create a cluster
- In the console, open Kubernetes and choose Create cluster.
- Pick the GPU type, node count, and region. Nodes in one cluster are in the same region and on the same interconnect.
- Wait for the cluster to show Ready. Provisioning bootstraps the control plane, joins the nodes, and installs the add-ons.
- Install the
kubectl oidc-loginplugin (one-time) and click Download kubeconfig — see Connect with kubectl for the install commands. Then pointkubectlat the file:
kubectl runs as you. You should see one node per GPU instance, each advertising nvidia.com/gpu capacity.
Node pools
Nodes are grouped into pools. From the cluster’s Pools tab you can see each pool’s GPU type, size, and status, and add or remove nodes. CPU-only worker pools for non-GPU workloads (proxies, monitoring, controllers) are on the roadmap.Running GPU workloads
Request GPUs like any other resource:Exposing services
Nodes have public IPs. ANodePort or hostNetwork service is reachable from the internet only once the port is open on the node — see Opening inbound ports. Prefer an in-cluster ingress plus one open port over exposing many NodePorts.
Storage
Node-local NVMe is available to pods ashostPath or a local volume and does not survive node replacement. Shared network storage for clusters is in development; until it ships, use your own object storage for checkpoints.
Monitoring
The cluster page in the console shows the pods and events on each node and live GPU telemetry. You are free to install Prometheus, Grafana, or any observability agent in the cluster. A fuller Monitoring tab with history is in development.What Hyperbolic runs, and what you run
Managed Slurm is on the roadmap; Kubernetes is the supported orchestration path today.

