DEPLOYMENT PLANNER

Design your private AI deployment

Build a topology for a pilot, a team, or a resilient production farm. See how AI Gateway, workers, Kubernetes and licensing fit together before you create any cloud resources.

Deployment inputs

LIVE TOPOLOGY

Gateway with a managed worker pool

Recommended
Infrastructure estimate
Software licenses
Indicative monthly total
USD, before tax

Planning estimate, not a quote. Rates are a 21 August 2026 public-price snapshot. The estimate excludes data transfer, public IPv4, load balancers, monitoring ingestion, backups, support plans, discounts, reserved or Spot capacity, tax, and model-specific performance requirements.

Validate the design with your security and platform teams, benchmark your chosen models, confirm GPU quota and regional SKU availability, then price the final architecture in the provider calculator.

A practical starting point

Pilot

Use one CPU worker to validate networking, keys, client compatibility and operations. It proves the deployment, not production inference speed.

Team

Start with two GPU workers behind AI Gateway. You can maintain one worker while the other serves traffic, and add capacity without changing clients.

Production

Use three or more workers across failure domains, a production control-plane SLA, autoscaling guardrails, persistent model storage, TLS ingress and your standard observability stack.

Get release updates

New free AI products, major updates, and a few releases available only via this site. No spam.