Sizing
All figures here are estimates; they vary with the model, its compression, document length and how bursty the use is. Pilot on real hardware and read the server's usage report (latency per model) before buying.
| Personal or pilot | Team server | Department server | |
|---|---|---|---|
| Hardware | A recent PC or laptop, 16 GB memory; a GPU is optional | One workstation with a GPU of 8–12 GB video memory, about 32 GB memory | A GPU of 16–24 GB video memory, about 64 GB memory |
| Model class | 1.5–3B parameters | 7–8B parameters | 14B parameters |
| Comfortable at the same moment | 1 | Roughly 5–15 active requests | Roughly 10–30 active requests |
| Typical team | 1 person | Dozens of people in bursty office use | 50–150 people |
Rules of thumb:
- A model compressed to 4 bits needs about 0.6–0.75 GB of memory per billion parameters, plus room for the context of each request in progress.
- Text generation speed is limited by memory bandwidth: a discrete GPU is typically several times faster than a CPU for the same model.
- "Active" is not "users": office use comes in bursts, so a server that handles ten requests at once serves far more named users.
- Image generation and long transcriptions occupy a GPU for seconds to minutes per job; give them their own worker when they share a farm with chat.
- Embedding models are small and run well on a CPU.
When to scale out
Add a second machine when one of these holds: the GPU is saturated during the working day (latency in the usage report rises with load), you need a model that does not fit beside the others, or a single machine is a risk you cannot accept. Scaling out means an AI Gateway (Pro Commercial).
The worker pool
The gateway's pool is gateway.json in its data folder, edited in the Windows app (Server → Server settings → Multi-host gateway) or by hand:
{
"enabled": true,
"routing_strategy": "least-connections",
"health_check_seconds": 10,
"failover_retries": 3,
"circuit_breaker_threshold": 3,
"circuit_breaker_cooldown_seconds": 30,
"canary_percent": 10,
"workers": [
{ "id": "gpu-1", "name": "Rack GPU 1", "base_url": "https://gpu-1.internal:11436", "bearer_token": "<key issued on gpu-1>", "max_in_flight": 8 },
{ "id": "gpu-2", "name": "Rack GPU 2", "base_url": "https://gpu-2.internal:11436", "bearer_token": "<key issued on gpu-2>", "max_in_flight": 8, "weight": 2 },
{ "id": "trial", "base_url": "https://gpu-3.internal:11436", "bearer_token": "<key issued on gpu-3>", "canary": true }
]
}
| Field | Meaning |
|---|---|
routing_strategy | round-robin (default), least-latency, least-connections, weighted, sticky |
health_check_seconds | How often each worker's /readyz is checked (default 30, 5–3600) |
failover_retries | Other workers to try when one fails before any output (default 3) |
circuit_breaker_threshold, …_cooldown_seconds | Failures before a worker is taken out, and how long it rests (defaults 3 and 30 s) |
workers[].max_in_flight | The most requests a worker gets at once; beyond it the gateway uses another worker or answers 503 |
workers[].background_max_in_flight, accepts | Limit or dedicate a worker to interactive or background work |
workers[].weight | Share of traffic for weighted routing |
workers[].canary, canary_percent | Canary workers and the share of callers sent to them |
worker_dns_name, worker_dns_port, discovered_worker_token | Find workers through DNS (Kubernetes headless Service) with one shared worker key |
auto_discover_workers, discovered_worker_fingerprints | Find workers on the local network; only listed certificate fingerprints are used |
Up to 25 workers per gateway on Commercial; more by Enterprise arrangement.
Routing
- Model-aware first: the gateway knows which models each worker has installed and prefers those workers, so requests rarely wait for a model download.
- Then the strategy:
least-connectionssuits mixed request sizes;least-latencysuits workers of different speeds;stickykeeps each caller on the same worker to reuse its warm context;weightedmatches traffic to unequal hardware.
Health, failover and backpressure
- Workers that fail health checks or trip the circuit breaker leave the pool and rejoin when healthy.
- A request that fails on a worker before any output is sent is retried on another worker; the client never sees the failure. A client that disconnects is not counted against the worker.
- When every worker is at its
max_in_flight, the gateway answers 503 withRetry-Afterinstead of queuing into a timeout. Background work is shed first. AI Suite clients wait and retry automatically.
Canary rollouts
Mark a worker "canary": true and set canary_percent. That share of callers — always the same callers, so nobody flips between versions — goes to canary workers. Watch errors and latency per worker on the dashboard, then promote the canary (remove the flag) or take it out of the pool.
Highly available gateways
- Run two or more gateway instances behind your load balancer or Kubernetes Service. They hold no shared state; routing statistics are per instance and rebuild within seconds.
- Gateways use no licence seats.
- Upgrades: workers one at a time, then gateways one at a time — see upgrades.
Questions
Does a farm share models between workers? +
No. Each worker keeps its own models on its own disk, which is what makes it independent. Model-aware routing makes sure a model only needs to be loaded on the workers that serve it.
Can workers be on different hardware? +
Yes. Use weighted or least-latency routing, and max_in_flight to protect smaller machines.