Sizing

All figures here are estimates; they vary with the model, its compression, document length and how bursty the use is. Pilot on real hardware and read the server's usage report (latency per model) before buying.

Personal or pilotTeam serverDepartment server
HardwareA recent PC or laptop, 16 GB memory; a GPU is optionalOne workstation with a GPU of 8–12 GB video memory, about 32 GB memoryA GPU of 16–24 GB video memory, about 64 GB memory
Model class1.5–3B parameters7–8B parameters14B parameters
Comfortable at the same moment1Roughly 5–15 active requestsRoughly 10–30 active requests
Typical team1 personDozens of people in bursty office use50–150 people

Rules of thumb:

  • A model compressed to 4 bits needs about 0.6–0.75 GB of memory per billion parameters, plus room for the context of each request in progress.
  • Text generation speed is limited by memory bandwidth: a discrete GPU is typically several times faster than a CPU for the same model.
  • "Active" is not "users": office use comes in bursts, so a server that handles ten requests at once serves far more named users.
  • Image generation and long transcriptions occupy a GPU for seconds to minutes per job; give them their own worker when they share a farm with chat.
  • Embedding models are small and run well on a CPU.

When to scale out

Add a second machine when one of these holds: the GPU is saturated during the working day (latency in the usage report rises with load), you need a model that does not fit beside the others, or a single machine is a risk you cannot accept. Scaling out means an AI Gateway (Pro Commercial).

The worker pool

The gateway's pool is gateway.json in its data folder, edited in the Windows app (Server → Server settings → Multi-host gateway) or by hand:

{
  "enabled": true,
  "routing_strategy": "least-connections",
  "health_check_seconds": 10,
  "failover_retries": 3,
  "circuit_breaker_threshold": 3,
  "circuit_breaker_cooldown_seconds": 30,
  "canary_percent": 10,
  "workers": [
    { "id": "gpu-1", "name": "Rack GPU 1", "base_url": "https://gpu-1.internal:11436", "bearer_token": "<key issued on gpu-1>", "max_in_flight": 8 },
    { "id": "gpu-2", "name": "Rack GPU 2", "base_url": "https://gpu-2.internal:11436", "bearer_token": "<key issued on gpu-2>", "max_in_flight": 8, "weight": 2 },
    { "id": "trial", "base_url": "https://gpu-3.internal:11436", "bearer_token": "<key issued on gpu-3>", "canary": true }
  ]
}
FieldMeaning
routing_strategyround-robin (default), least-latency, least-connections, weighted, sticky
health_check_secondsHow often each worker's /readyz is checked (default 30, 5–3600)
failover_retriesOther workers to try when one fails before any output (default 3)
circuit_breaker_threshold, …_cooldown_secondsFailures before a worker is taken out, and how long it rests (defaults 3 and 30 s)
workers[].max_in_flightThe most requests a worker gets at once; beyond it the gateway uses another worker or answers 503
workers[].background_max_in_flight, acceptsLimit or dedicate a worker to interactive or background work
workers[].weightShare of traffic for weighted routing
workers[].canary, canary_percentCanary workers and the share of callers sent to them
worker_dns_name, worker_dns_port, discovered_worker_tokenFind workers through DNS (Kubernetes headless Service) with one shared worker key
auto_discover_workers, discovered_worker_fingerprintsFind workers on the local network; only listed certificate fingerprints are used

Up to 25 workers per gateway on Commercial; more by Enterprise arrangement.

Routing

  • Model-aware first: the gateway knows which models each worker has installed and prefers those workers, so requests rarely wait for a model download.
  • Then the strategy: least-connections suits mixed request sizes; least-latency suits workers of different speeds; sticky keeps each caller on the same worker to reuse its warm context; weighted matches traffic to unequal hardware.

Health, failover and backpressure

  • Workers that fail health checks or trip the circuit breaker leave the pool and rejoin when healthy.
  • A request that fails on a worker before any output is sent is retried on another worker; the client never sees the failure. A client that disconnects is not counted against the worker.
  • When every worker is at its max_in_flight, the gateway answers 503 with Retry-After instead of queuing into a timeout. Background work is shed first. AI Suite clients wait and retry automatically.

Canary rollouts

Mark a worker "canary": true and set canary_percent. That share of callers — always the same callers, so nobody flips between versions — goes to canary workers. Watch errors and latency per worker on the dashboard, then promote the canary (remove the flag) or take it out of the pool.

Highly available gateways

  • Run two or more gateway instances behind your load balancer or Kubernetes Service. They hold no shared state; routing statistics are per instance and rebuild within seconds.
  • Gateways use no licence seats.
  • Upgrades: workers one at a time, then gateways one at a time — see upgrades.

Questions

Does a farm share models between workers? +

No. Each worker keeps its own models on its own disk, which is what makes it independent. Model-aware routing makes sure a model only needs to be loaded on the workers that serve it.

Can workers be on different hardware? +

Yes. Use weighted or least-latency routing, and max_in_flight to protect smaller machines.