The chart installs the production shape: an AI Gateway Deployment that clients call, in front of a StatefulSet of AI Server workers, with a NetworkPolicy so only the gateway can reach the workers. Gateway mode needs AI Server Commercial; each worker uses one licence seat and the gateway uses none.

AI Gateway farmClients send requests to the AI Gateway with their client key. The gateway checks keys and governance, then forwards each request to a healthy worker with the worker key. Workers can be added, drained or upgraded behind it.Clientsone URL · one keyclient keyAI Gatewayrouting · failover · governanceWorker 1GPU serverWorker 2GPU serverWorker 3 (canary)new model or buildworker key
In gateway mode one AI Server is the endpoint; workers do the inference. Clients keep one address and one key.

Before you start

  • A Kubernetes cluster with a CNI that enforces NetworkPolicy (Calico, Cilium, Antrea…).
  • For GPU workers: the NVIDIA device plugin advertising nvidia.com/gpu.
  • An AI Server Commercial licence key.
  • Two API key files: the client keys for the gateway and the worker key the gateway presents to workers. They are separate stores; a client key is not valid on a worker.

1. Issue the keys

aisuite-server-cli keys add --label clients --data ./seed/gateway
aisuite-server-cli keys add --label gw-to-workers --data ./seed/worker   # note the raw key it prints
aisuite-server-cli keys add --label prometheus --admin --data ./seed/gateway   # optional: for monitoring

Each command prints the raw key once and writes only its hash to ./seed/<tier>/credentials/server-keys.json. If you do not have the CLI installed, the images carry it from 2.5.2: docker run --rm -v "$PWD/seed/gateway:/data" --entrypoint aisuite-server-cli softwaretailor/aiserver:<version> keys add --label clients.

2. Install

helm repo add softwaretailor https://softwaretailor.com/charts
helm repo update
kubectl create namespace aisuite
kubectl -n aisuite create secret generic aisuite-license --from-literal=key=<licence-key>

helm install aisuite softwaretailor/aisuite -n aisuite \
  --set license.existingSecret=aisuite-license \
  --set-file gateway.clientKeys=./seed/gateway/credentials/server-keys.json \
  --set-file worker.keys=./seed/worker/credentials/server-keys.json \
  --set gateway.pool.dnsDiscovery.enabled=true \
  --set gateway.pool.dnsDiscovery.workerToken=<raw gw-to-workers key> \
  --set worker.engine=real \
  --set worker.persistence.size=200Gi

The chart refuses to render a configuration that would install and then not work — no licence, no keys, no worker pool, autoscaling over a static pool, CPU autoscaling on GPU workers. The error tells you which value to set.

Note: worker.engine defaults to stub, which answers with canned test replies so you can check the deployment without a GPU. Set real to serve inference.

3. Check it

kubectl -n aisuite rollout status deploy/aisuite-aisuite-gateway
kubectl -n aisuite logs deploy/aisuite-aisuite-gateway | grep -i license
kubectl -n aisuite port-forward svc/aisuite-aisuite-gateway 8080:8080
curl -H "Authorization: Bearer <admin key>" http://localhost:8080/v1/gateway/workers

Every worker should be listed as healthy. An empty pool usually means workerToken is not a key issued on the workers.

Values you will usually change

ValueDefaultNotes
worker.enginestubreal serves inference.
worker.replicaCount2Ignored when autoscaling is on.
worker.persistence.size5GiModels need tens of GB each: 200Gi is a realistic start, 500Gi for Pro AI Engine A.
gpu.enabled, gpu.countoff, 1Requests and limits nvidia.com/gpu; adds the GPU node selector and toleration.
worker.proEngineA.enabledoffUses softwaretailor/aiserver-pro; needs gpu.enabled.
gateway.pool.routingStrategyleast-connectionsOr round-robin, least-latency, weighted, sticky.
ingress.enabled, ingress.host, ingress.tlsoffTerminate TLS at the ingress; the gateway Service is ClusterIP.
networkPolicy.ingressNamespaceingress-nginxThe only namespace the gateway accepts traffic from when an Ingress is rendered.
networkPolicy.monitoringNamespacemonitoringPrometheus's namespace, admitted when serviceMonitor.enabled.
license.registrationUrldefault serverLicence activation server. Nodes need outbound HTTPS to it.
lifecycle.drainSeconds, shutdownSeconds8, 30The pod's termination grace period is computed from these.

Kustomize users: the same topology ships as reference manifests with autoscale, gpu and pro-engine-a overlays. The chart is the supported path.

Why it looks this way

  • Workers are a StatefulSet behind a headless Service. The gateway picks workers itself (health, warm models, stickiness, canary); an L4 Service underneath would undo that. Each worker keeps its own volume, so a rescheduled pod keeps its models.
  • The gateway finds workers by DNS when dnsDiscovery.enabled is set: pods an autoscaler adds join the pool on the next health check without a restart. A static pool gives each worker its own key but cannot autoscale.
  • Only the gateway can reach the workers. That is what makes it safe for workers to trust the client address the gateway forwards. The policy only works on a CNI that enforces it — check with a direct request from another pod.
  • Keys are mounted as single files, so the rest of the daemon's credentials folder stays writable.
  • The Ingress must not buffer. Streaming answers are Server-Sent Events; turn proxy buffering off and allow read timeouts of 60 seconds or more.

Zero-downtime upgrades

On SIGTERM a pod answers 503 on /readyz for the drain period, so the gateway and the Service stop sending it work, then finishes in-flight requests and hands its licence seat back. helm upgrade with a new image tag rolls the workers one at a time with no dropped requests. Keep the termination grace period at or above drain plus shutdown time — the chart computes it.

Autoscaling

Autoscaling needs DNS discovery, an HPA and a signal. CPU is the wrong signal for GPU inference — the GPU saturates while the CPU idles. Scale on the gateway's in-flight gauge (aisuite_gateway_worker_in_flight) through prometheus-adapter:

--set autoscaling.enabled=true \
--set autoscaling.customMetric.enabled=true \
--set autoscaling.minReplicas=2 --set autoscaling.maxReplicas=6

Scale-down waits 10 minutes by default: removing a worker mid-answer costs a user their reply and a new worker has to load its model again.

Monitoring

Set serviceMonitor.enabled=true with the Prometheus Operator installed. The chart scrapes /v1/server/metrics on the gateway and workers with bearer keys from the Secrets aisuite-scraper-key (a gateway key issued with --admin) and aisuite-worker-scraper-key (a worker key issued with --admin). See monitoring.

Questions

Can I run AI Server on Kubernetes without the gateway? +

Yes: run the worker image as a Deployment or StatefulSet with your own Service. You lose model-aware routing and failover, and you need only a Personal or Commercial licence per server. The chart always installs the gateway.

Does the gateway need a GPU? +

No. Only workers need GPUs; the gateway is a light proxy (250m CPU and 256Mi memory requested by default).

Is the chart on Azure and AWS marketplaces? +

Yes, as bring-your-own-licence Kubernetes applications for AKS and EKS. They install this chart with the marketplace's own UI for the values above.