ONLINEBILLING:Per-Second Metering · $0 EgressONLINEPOWER:100 kW/Rack Direct-to-Chip Liquid CoolingONLINEUS-East:128 x H100 SXM AvailableONLINEUS-West:64 x H200 AvailableLIMITEDEU-Central:GB200 NVL72 Reserve OnlyFABRIC:GPUDirect RDMA Enabled · 3.2 Tbps InfiniBand · Data Egress: $0.00/GB
ONLINEBILLING:Per-Second Metering · $0 EgressONLINEPOWER:100 kW/Rack Direct-to-Chip Liquid CoolingONLINEUS-East:128 x H100 SXM AvailableONLINEUS-West:64 x H200 AvailableLIMITEDEU-Central:GB200 NVL72 Reserve OnlyFABRIC:GPUDirect RDMA Enabled · 3.2 Tbps InfiniBand · Data Egress: $0.00/GB

Enterprise compute

A control plane built for multi-node training, not a Docker wrapper.

kwctl handles node discovery, dynamic MIG partitioning, and pre-flight GPU health gating so a failing card never takes down a 512-GPU run.

kwctl
$ kwctl cluster create --nodes 16 --gpu b200-nvlink --network infiniband-3.2t

Bare-Metal Kubernetes, Ray & Slurm

Pre-configured Ray clusters, Slurm workload scheduling, and bare-metal Kubernetes operators tuned for multi-node distributed training.

  • Ray autoscaling groups
  • Slurm partitions & fair-share
  • K8s GPU operator + device plugin

InfiniBand & RoCE v2 Topology

Non-blocking leaf-spine fabric at 3.2 Tbps with GPUDirect RDMA so GPUs read and write remote VRAM without traversing CPU or host RAM.

  • Quantum-2 NDR InfiniBand
  • RoCE v2 lossless Ethernet
  • SHARP in-network reductions

kwctl Control Plane & API

Proprietary lightweight orchestration for automated node discovery, dynamic MIG partitioning, and pre-flight health checks that isolate failing GPUs.

  • Automated node discovery
  • MIG dynamic partitioning
  • Pre-run GPU health gating

Checkpointing & Fault Recovery

Long-run resilience with distributed checkpoint offload and provider-level failover routing when a node refuses or fails a job.

  • Async checkpoint offload
  • Multi-provider failover routing
  • Job resume from last shard

Cluster architectures

Reference topologies

ArchitectureScaleFabricPowerWorkload
GB200 NVL72 Rack72 GPUs · 1 rackNVLink 5 domain + NDR IB~120 kW liquid-cooledFrontier-scale foundation model training
B200 SXM Pod16 nodes · 128 GPUs3.2 Tbps non-blocking InfiniBand100 kW/rack direct-to-chipMulti-node pre-training with RDMA
H200 Inference Fleet8 nodes · 64 GPUsRoCE v2 lossless Ethernet60 kW/rack hybrid loopHigh-volume LLM serving at low TTFT
H100 Training Cell32 nodes · 256 GPUsQuantum-2 InfiniBand leaf-spine45 kW/rackDistributed pre-training & fine-tuning

Platform safeguards

Enforced in code, not promised on paper

Usage metering interval

Every 15 min

Automatic instance shutdown at zero balance

Job pre-authorization

100% of estimated cost

402 before any node is contacted

Rate limiting

30 req/min per key

Deterministic 429 responses

Concurrency

3 jobs per key

429 until a slot frees

Support response

Same business day

Direct email escalation to the founding team

Payment security

PCI-DSS via Stripe

Card data never touches our servers

Contact

Questions, quotes or support — one address, every team.

hello@kilawattcloud.dev