Components
| Component | What it is |
|---|---|
Server (aisuite-server) | One process that serves the API. It runs as a child of the Windows app, as a Windows service, or as the entry point of the container images. Built on .NET 10 and ASP.NET Core. |
| Windows app | The operator's console: starts the server or installs it as a service, manages models, keys, governance, monitoring and licensing. |
CLI (aisuite-server-cli) | Keys, status, audit export and verification, enrollment, benchmarks — for scripts and headless servers. |
| Engines | The model runtimes the server downloads, starts and supervises. See engines. |
| AI Gateway | The same server started in gateway mode. It holds no models and forwards work to a pool of servers. |
| Clients | AI Suite apps, AI Client, and any OpenAI-compatible tool or code. |
One binary plays three roles, chosen at start: a server with real engines, a gateway, or a stub that answers with canned text for testing deployments.
Request pipeline
Every request passes the same stages, in this order:
- Forwarded headers — only from proxies you declared trusted, so client addresses cannot be spoofed.
- Host check — on a server bound to its own computer, only
localhost,127.0.0.1and[::1]are answered (blocks DNS rebinding). - CORS — browser origins you allowed get their preflight answered.
- Metrics and request id — timing,
X-Request-Id, the content-free usage record. - Authentication — API key (or a signed request from an AI Suite app on the same computer), with throttling of repeated wrong keys, endpoint and model scopes and per-key rate limits.
- Free allowance — on Free only, for third-party tools.
- Rate limits — per key and per client address.
- Quotas and budgets — per key, per day and per month.
- Scheduling — interactive work first; background work waits or receives 503 with
Retry-After. - Policy checks — model lifecycle, content rules and moderation.
- Engine — the model runs; streamed output flows back through the same path.
Every stage that refuses answers with an OpenAI-shaped error, so clients handle refusals the same way.
Engines and model routing
A model id names its runtime (runtime/family/variant). The server's engine orchestrator sends each request to the backend for that runtime, starting the engine and loading the model when needed. Local engines run as separate processes on internal ports that are not exposed; the server is the only thing clients talk to. Cloud providers, if an operator configures them, are just another runtime.
Storage
| What | Where |
|---|---|
| Control data: keys, licence lease, settings, governance, audit, logs | The data folder (%USERPROFILE%\AISuite\v2 on Windows, /data in containers) |
| Large caches: engines, models | The shared cache folder, which can be moved to another disk |
| Requests and answers | Memory only, for the time of the request |
On Windows the data folder is shared by the app, its service and the CLI, so all three see the same keys and licence.
Discovery and trust
- Same computer: apps find the server through a lock file that records its address, port and scheme.
- Local network: servers can announce themselves; apps list them and ask for an API key.
- Certificates: with HTTPS on, apps pin the server's certificate fingerprint on first use and warn if it changes.
Gateway mode
A gateway accepts client requests with the clients' keys, applies authentication, scopes, rate limits, quotas, lifecycle and content rules at the edge, then forwards each request to a worker with that worker's own key. Workers apply their own policies too. The gateway chooses a worker by health, load and which models it has installed; retries transient failures before any output is sent; and turns saturation into 503 with Retry-After. Details: gateway internals.
Process model and shutdown
- On SIGTERM (or a service stop) the server reports not-ready on
/readyz, waits for load balancers to notice, finishes in-flight work, hands back its licence seat and exits. - A server whose licence plan changes restarts itself (exit code 5) so the new plan applies everywhere at once.
- Exit codes for refusals are listed in the configuration reference.
Questions
Is AI Server a wrapper around another product? +
AI Server is its own server: the API, authentication, governance, gateway and operations are ours. It runs established open-source inference runtimes as engines underneath, which it installs, updates and supervises.
Can I run the engines on different machines from the server? +
Use the gateway for that: each worker is a full AI Server with its engines, and the gateway spreads the work.