NVIDIA B200
FP8 / FP4 Support
API-routed access to H100, H200 and A100 capacity across our supplier network, with live inventory checks and automatic failover when one refuses or fails a job.
Jobs run on provider-hosted NVIDIA® H100, H200 and A100 GPUs, optimized for CUDA® workloads.
$ kwctl auth login --api-key $KILAWATT_API_KEY$ kwctl deploy --gpu h100 --count 8 --policy zero_quota$ # example output: provider selected by live cost... failover armedUnified hardware catalog
Enterprise-grade NVIDIA® accelerators routed across live supplier inventory. B200 is marketplace-priced and self-serve for up to four GPUs only when matching stock is available.
GPU model
NVIDIA B200
FP8 / FP4 Support
NVIDIA H200
Ultra-large LLM Inference
NVIDIA H100 SXM
Distributed Training Workhorse
NVIDIA A100 SXM
Cost-efficient Training
NVIDIA L40S
Fine-tuning & Vision
Multi-application infrastructure
Kilawatt Cloud abstracts our provider network into a single API endpoint. Request workloads dynamically across multiple infrastructure suppliers, with live capacity checks and automatic failover instead of hyperscaler quota requests.
MCP Native
Multi-node orchestration
Globally routed execution
Smart order router
Kilawatt Router
SOR online
Hover or focus an infrastructure node to inspect the active path. Kilawatt continuously evaluates price, latency, quota, and health before forwarding each workload.
zero-quota primary · health check passed · failover armed
Enterprise pillars
Access dedicated NVIDIA accelerators across multiple providers, with each launch routed against current inventory rather than a hyperscaler quota queue.
When a provider refuses or fails a job, the router moves to the next candidate automatically. Every attempt is recorded with the raw provider response.
Full-cost pre-authorization before any node is contacted, per-key rate and concurrency limits, and automatic metering with shutdown when balance runs out.
$0 egress fees, pricing quoted from the real live cost of the provider that runs your job, and optional monthly spend caps you control from the console.
Software orchestration
Pre-configured Ray clusters, Slurm workload scheduling, and bare-metal Kubernetes operators tuned for multi-node distributed training.
Jobs land on the interconnect the serving provider offers — NVLink within a node, and RDMA-capable fabric where the provider supports it. Network topology is reported per job, never assumed.
Lightweight orchestration for provider selection, live cost quoting, pre-flight profitability checks and automated teardown at the end of the paid window.
Every provider attempt — winning and failing — is written to an append-only ledger with the raw response body, enforced by the database rather than application discipline.
Hardware matrix
FP8 / FP4 Support
from $8.44
floor — quoted live per request
Memory
180GB HBM3e
Bandwidth
8.0 TB/s Bandwidth
Fabric
NVLink 5 · 1.8 TB/s
Platform safeguards
Full-Cost Pre-Authorization
Every job is paid in full up front — insufficient balance returns a 402 before any node is contacted.
Per-Key Rate Limits
30 requests/min and 3 concurrent jobs per API key by default, with deterministic 429s.
Customer Spend Caps
Optional monthly spend ceiling per account and per API key, removable only by you.
Metering & Auto-Shutdown
Instances are metered every 15 minutes and terminated at the provider when balance or the paid window ends.
Scoped API Keys
kw_live_ keys are SHA-256 hashed at rest, shown once, and revocable instantly.
Execution Ledger
Every API call, charge and top-up is written to an append-only transaction log.
TCO calculator
Datacenter class · fixed rate card
B200 and A100 show a floor price: their real supply is marketplace-priced and varies, so the final rate is quoted live at request time and can be above the floor. H200, H100 SXM and L40S are fixed published rates.
Workstation class · live marketplace price
Workstation cards are sourced from live market supply, so their rate is quoted live per request and moves with the market — it is not a fixed published price. Self-serve up to 16 GPUs in a single machine.
GPU count
64
Single-node deployments can be initiated directly from the console, subject to live provider capacity. B200 is self-serve up to 4 GPUs when matching stock exists. Other datacenter cards support up to 8 GPUs, and workstation cards up to 16 GPUs when live stock exists. Larger clusters are arranged with our team.
Parallel storage
50 TB
Deployment tier
On-Demand, metered every 15 minutes. Committed multi-year terms are not currently sold — talk to us if you need one and we will quote it directly.
Estimated monthly TCO
$395,417
64 x B200 · 50 TB · 730 hrs · $0.00/GB egress
Single-node On-Demand is self-serve subject to live provider capacity: B200 up to 4 GPUs, other datacenter cards up to 8 GPUs, and workstation cards up to 16 GPUs when a matching machine is available. Larger configurations and committed terms are quoted by our team.
Pay-as-you-go
Pay for the duration you request, with zero egress charges and full-cost pre-authorization before any node is contacted.
By quote
Datacenter-class clusters above 8 GPUs, workstation machines above 16 GPUs, and any committed term are arranged directly with our team. No provider rents an H100/H200-class box bigger than 8 GPUs on demand, so we quote those per build rather than publishing a rate we cannot honour.
Talk to Enterprise SalesDeveloper quick-start
$ kwctl keys create --name enterprise-prod$ kwctl deploy --gpu h100 --count 8$ kwctl instances list --status running2.3s
Fastest measured provision
4
Providers in routing path
$0.00/GB
Egress
Questions, quotes or support — one address, every team.