01/ Pricing

Pay for the GPU-seconds you run.

Serverless inference bills per GPU-second while your code is on the GPU. Queue, scheduling, and prepare time are not charged. Training mode bills at the underlying GCP VM rate.

Reference catalogStatic rate card shown while the runtime catalog is unavailable
GPUVRAM/ minStatusInferenceTrain
NVIDIA T4t4
16 GB$0.013AvailableOn demandSupported
NVIDIA L4l4
24 GB$0.021AvailableOn demandSupported
NVIDIA V100v100
16 GB$0.083Coming soon
NVIDIA H100h100
80 GB$0.327Coming soon
NVIDIA B200b200
180 GB$0.600Coming soon
02/ Plans

Three plans. Zero hidden fees.

Starter
$0
  • /Personal workspace
  • /Pay per GPU-second
  • /CLI and agent SDK
  • /Community support
Team
$TBD
/ month
  • /Shared workspace (coming soon)
  • /Priority inference warm pool
  • /Usage export
  • /SLA response
Enterprise
Custom
  • /Dedicated nodes
  • /Commit discounts
  • /SSO and audit log (roadmap)
  • /Named support
03/ Notes

Billing pipeline (invoices, auto-recharge, spend-limit) ships after the v1 console. Current run summary shows a demo charge calculated client-side from the rates above.