Before handover
A cluster is not delivered until it has passed:- Acceptance certification. A benchmark run across the delivered nodes — GPU health, NCCL collectives (all-reduce, reduce-scatter, all-gather), per-rail and pairwise fabric bandwidth, and correctness — scored against pass thresholds. The certification run ID and score are recorded in your document.
- Hardening validation. The effective SSH daemon configuration is audited (
sshd -T) for public-key-only authentication; a password-only login attempt must fail withPermission denied (publickey); default accounts are confirmed locked; the firewall or security-group policy is reviewed against the intended exposure. - Account cleanup. Provisioning, installer, and partner accounts are removed. Your user and public key are installed; the sudo decision (passwordless or not) is recorded.
- Data sanitization. Drives that have been used before are block-erased at the hardware level and fresh filesystems created. Nothing from a previous tenant is copied forward.
- Runtime checks. Docker and the NVIDIA Container Toolkit are verified on every node with a disposable GPU container that must see all GPUs; the test containers and images are removed afterwards.
- Cleanup. Benchmark containers, temporary files, and benchmarking artifacts are removed. No benchmark jobs are left running at handover unless coordinated with you.
- Baseline capture. OS, kernel, driver, container runtime, and RDMA port state are recorded per node.
When nodes are added to an existing cluster, the additions get access, runtime, and sanitization checks and a fresh baseline. The original certification is not automatically extended to the combined cluster — ask for a re-certification if you want a benchmark that covers the new topology.

