Instance Types
Two instance types are available, both built on H100 SXM GPUs:
The 1xH100 is suited for development, single-GPU fine-tuning, and inference workloads. The 8xH100 is designed for large-scale training, multi-GPU inference, and distributed computing. Resources scale proportionally: the 8-GPU instance has 8x the CPU cores, memory, VRAM, and storage of the single-GPU instance.
Multi-Node and InfiniBand
When you need to distribute a workload across multiple machines, provision 8xH100 instances in the same sector. Instances within a sector are connected over InfiniBand, providing ultra-low latency and high bandwidth for frameworks like PyTorch DDP, DeepSpeed, and Horovod.InfiniBand and sector placement are only available on 8xH100 instances. 1xH100 instances run as standalone machines without inter-node connectivity.
When to Use Compute vs Serverless
The two products serve different workload profiles:
Use Compute when you need sustained GPU access for hours or days at a time. Use Serverless when you need an API that scales to zero and handles traffic spikes automatically.
Getting Started
Provisioning an instance takes about 2-3 minutes. You choose an instance type, select a sector (for multi-node setups), paste your SSH public key, and click create. Once the instance is ready, you SSH in and have full control.Quickstart
Provision your first instance and run a GPU workload
Pricing
Per-hour rates by instance type