Enterprise compute
A control plane built for multi-node training, not a Docker wrapper.
kwctl handles node discovery, dynamic MIG partitioning, and pre-flight GPU health gating so a failing card never takes down a 512-GPU run.
$ kwctl cluster create --nodes 16 --gpu b200-nvlink --network infiniband-3.2tBare-Metal Kubernetes, Ray & Slurm
Pre-configured Ray clusters, Slurm workload scheduling, and bare-metal Kubernetes operators tuned for multi-node distributed training.
- › Ray autoscaling groups
- › Slurm partitions & fair-share
- › K8s GPU operator + device plugin
InfiniBand & RoCE v2 Topology
Non-blocking leaf-spine fabric at 3.2 Tbps with GPUDirect RDMA so GPUs read and write remote VRAM without traversing CPU or host RAM.
- › Quantum-2 NDR InfiniBand
- › RoCE v2 lossless Ethernet
- › SHARP in-network reductions
kwctl Control Plane & API
Proprietary lightweight orchestration for automated node discovery, dynamic MIG partitioning, and pre-flight health checks that isolate failing GPUs.
- › Automated node discovery
- › MIG dynamic partitioning
- › Pre-run GPU health gating
Checkpointing & Fault Recovery
Long-run resilience with distributed checkpoint offload and provider-level failover routing when a node refuses or fails a job.
- › Async checkpoint offload
- › Multi-provider failover routing
- › Job resume from last shard
Cluster architectures
Reference topologies
| Architecture | Scale | Fabric | Power | Workload |
|---|---|---|---|---|
| GB200 NVL72 Rack | 72 GPUs · 1 rack | NVLink 5 domain + NDR IB | ~120 kW liquid-cooled | Frontier-scale foundation model training |
| B200 SXM Pod | 16 nodes · 128 GPUs | 3.2 Tbps non-blocking InfiniBand | 100 kW/rack direct-to-chip | Multi-node pre-training with RDMA |
| H200 Inference Fleet | 8 nodes · 64 GPUs | RoCE v2 lossless Ethernet | 60 kW/rack hybrid loop | High-volume LLM serving at low TTFT |
| H100 Training Cell | 32 nodes · 256 GPUs | Quantum-2 InfiniBand leaf-spine | 45 kW/rack | Distributed pre-training & fine-tuning |
Platform safeguards
Enforced in code, not promised on paper
Usage metering interval
Every 15 min
Automatic instance shutdown at zero balance
Job pre-authorization
100% of estimated cost
402 before any node is contacted
Rate limiting
30 req/min per key
Deterministic 429 responses
Concurrency
3 jobs per key
429 until a slot frees
Support response
Same business day
Direct email escalation to the founding team
Payment security
PCI-DSS via Stripe
Card data never touches our servers
