The gateway is AI Server in gateway mode: it accepts the same API, authenticates callers with its own key store, and forwards each request to a worker AI Server. Configuration is described in scaling and high availability; this page explains the behaviour.

AI Gateway farmClients send requests to the AI Gateway with their client key. The gateway checks keys and governance, then forwards each request to a healthy worker with the worker key. Workers can be added, drained or upgraded behind it.Clientsone URL · one keyclient keyAI Gatewayrouting · failover · governanceWorker 1GPU serverWorker 2GPU serverWorker 3 (canary)new model or buildworker key
In gateway mode one AI Server is the endpoint; workers do the inference. Clients keep one address and one key.

Choosing a worker

For each request the gateway:

  1. Reads the model from the request (and resolves aliases).
  2. Builds the candidate set: healthy workers whose circuit is closed, that accept the request's work class (interactive or background), and that are below their in-flight cap.
  3. Splits callers into the canary slice or the stable group, then within each prefers workers that report the model installed ("warm"), then orders them by the routing strategy. Workers whose model list is not known yet count as warm.
  4. Tries the first candidate, falling back down the ordered list on failure.

Background work never makes a cold worker download a model while any warm worker exists; it waits for a warm slot instead.

Routing strategies

StrategyHow it picks
round-robinNext candidate in turn
least-latencyLowest exponentially weighted moving average of recent latency
least-connectionsFewest requests in flight
weightedWeighted round-robin: each worker appears weight times in the rotation
stickyRendezvous hashing of the caller's identity, so a caller keeps the same worker while it is healthy and moves minimally when the pool changes

Statistics are kept per gateway instance; several gateways behind a load balancer balance independently, and sticky hashing gives the same answer on every instance.

Canary

A stable hash of the caller decides whether it belongs to the canary_percent slice, so the same callers always go to canary workers and nobody flips between versions mid-conversation. A canary caller is sent to the canary even if its model is cold — trialling it is the point — and falls back to the stable group if the canary is down.

Health and circuit breaking

Every health_check_seconds the gateway calls each worker's /readyz and model list. A worker that is draining reports 503 and leaves the pool immediately. circuit_breaker_threshold consecutive failures open the circuit for circuit_breaker_cooldown_seconds, after which one trial request decides whether it closes. A client that disconnects mid-request is not counted as a worker failure.

Retries

Transient failures — connection refused, timeouts before the first byte, 502, 503, 504 — are retried on another candidate, up to failover_retries times, as long as nothing has been sent to the client yet. Request bodies are buffered so they can be replayed. 4xx answers from a worker are passed through once, unchanged. Once streaming has started, a failure ends the stream; it is never silently restarted on another worker.

Backpressure

When no candidate is below its cap, the gateway answers 503 with Retry-After immediately rather than queuing. Background work is refused before interactive work, and receives a longer Retry-After. An idle timeout per forwarded request (default 600 seconds, reset by streamed output) bounds stuck workers.

Governance at the edge

Authentication, key scopes, rate limits, quotas, the model allowlist, lifecycle policy and content rules run on the gateway before forwarding, so a request that will be refused never occupies a worker. Workers apply their own policies again. Streamed answers are metered from the stream's usage chunk, so quotas and budgets see streaming traffic too.

Credentials

Clients authenticate with gateway keys; the gateway authenticates to each worker with that worker's key. Client keys are never forwarded. The caller's address goes to workers in X-Forwarded-For, which workers trust only because only the gateway can reach them.

Discovery

  • Static: workers listed in gateway.json, each with its own key.
  • DNS: the gateway resolves a name (a Kubernetes headless Service) every health-check cycle and adds or removes workers; an empty answer keeps the previous pool. All discovered workers share one key.
  • Local network: workers announce themselves; only those whose certificate fingerprints are listed are used, so an unknown device never receives a key. Static entries win over discovered ones with the same id.

What a client sees

The same API, keys and errors as a single server, plus X-AISuite-Backend naming the worker. GET /v1/models on the gateway returns the union of the workers' models.