Planned maintenance
Maintenance we schedule — platform software on your instance’s host, firmware, BIOS, or driver updates — is tested before rollout, and:
- You’ll get at least one week’s notice by email, and two weeks where we can.
- The notice tells you what’s changing, the window, and whether your instance will be rebooted (your instance and its data come back; anything running at the time is interrupted) or reprovisioned (a fresh instance — onboard data does not carry over).
- On instances that occupy a whole node, timing is your call. Reply with a window that works, including off-hours or weekends, and we confirm before we start. Single-GPU instances share a host with other tenants, so we set that window and give you as much notice as we can.
- Afterwards we confirm the node comes back and stays reachable. If the fault the maintenance was meant to clear comes back, we come back to you about a follow-up window rather than acting on our own.
- For clusters, maintenance is rolled in batches so the whole cluster is never down at once.
Maintenance the datacenter schedules on the host underneath your instance is sometimes mandatory and comes with its own timeline. We forward the notice as soon as we receive it and tell you the window we were given. It can be hours rather than days.
Unplanned outages
Current incidents are posted on the status page. If your instance is affected, we’ll reach out with what happened and next steps. Downtime credits are reviewed case by case with support — see Getting Support.
Protecting your data
- Onboard NVMe is local to the machine. It survives a platform reboot; it does not survive reprovisioning, host replacement, or termination.
- Persistent storage volumes survive all of the above.
- Before you approve a maintenance window, before you ask us to power-cycle a node with a hardware fault, and on a schedule during long training runs, copy checkpoints to a persistent volume or your own object storage. Storage and Ports has ready-to-use
rsync and aws s3 sync cron lines. Rebooting an instance covers what a reboot does to onboard storage.
Treat onboard NVMe as scratch. If it’s the only copy of your checkpoints, you’re one host reboot away from losing them.