01/ Pricing
Pay for the GPU-seconds you run.
Serverless inference bills per GPU-second while your code is on the GPU. Queue, scheduling, and prepare time are not charged. Training mode bills at the underlying GCP VM rate.
Reference catalogStatic rate card shown while the runtime catalog is unavailable
| GPU | VRAM | / min | Status | Inference | Train |
|---|---|---|---|---|---|
NVIDIA T4t4 | 16 GB | $0.013 | Available | On demand | Supported |
NVIDIA L4l4 | 24 GB | $0.021 | Available | On demand | Supported |
NVIDIA V100v100 | 16 GB | $0.083 | Coming soon | — | — |
NVIDIA H100h100 | 80 GB | $0.327 | Coming soon | — | — |
NVIDIA B200b200 | 180 GB | $0.600 | Coming soon | — | — |
02/ Plans
Three plans. Zero hidden fees.
Team
$TBD
/ month
- /Shared workspace (coming soon)
- /Priority inference warm pool
- /Usage export
- /SLA response
Enterprise
Custom
- /Dedicated nodes
- /Commit discounts
- /SSO and audit log (roadmap)
- /Named support
03/ Notes
Billing pipeline (invoices, auto-recharge, spend-limit) ships after the v1 console. Current run summary shows a demo charge calculated client-side from the rates above.