AI Server architectureClients on the left call AI Server over its API. AI Server runs local engines and, only if an operator configures them, cloud providers. The licence server receives activation only.AI Suite appsdiscover the serverAI Clientthin client on each deskYour code and toolsOpenAI SDKs, agentsAI Serverkeys · governance · auditLocal engineschat · embeddings · imagesSpeech and visionTTS · STT · detectionCloud providersoptional, operator-setLicence serveractivation · weekly leaselicence only
One AI Server serves AI Suite apps, AI Client and your own code. Models and engines run on the server; nothing passes through Software Tailor.

Components

ComponentWhat it is
Server (aisuite-server)One process that serves the API. It runs as a child of the Windows app, as a Windows service, or as the entry point of the container images. Built on .NET 10 and ASP.NET Core.
Windows appThe operator's console: starts the server or installs it as a service, manages models, keys, governance, monitoring and licensing.
CLI (aisuite-server-cli)Keys, status, audit export and verification, enrollment, benchmarks — for scripts and headless servers.
EnginesThe model runtimes the server downloads, starts and supervises. See engines.
AI GatewayThe same server started in gateway mode. It holds no models and forwards work to a pool of servers.
ClientsAI Suite apps, AI Client, and any OpenAI-compatible tool or code.

One binary plays three roles, chosen at start: a server with real engines, a gateway, or a stub that answers with canned text for testing deployments.

Request pipeline

Request pipelineA request passes forwarded-header handling and the host check, API key authentication and scopes, the Free allowance, rate limits, quotas and budgets, scheduling, content and lifecycle checks, and then the engine. Usage and audit are recorded content-free.Host checkAPI keyRate limitsQuota and budgetSchedulingPolicy checksEngineUsage and audit records: endpoint, model, key, status, latency, tokens — never prompt or response text
Every request passes the same checks in this order before a model sees it.

Every request passes the same stages, in this order:

  1. Forwarded headers — only from proxies you declared trusted, so client addresses cannot be spoofed.
  2. Host check — on a server bound to its own computer, only localhost, 127.0.0.1 and [::1] are answered (blocks DNS rebinding).
  3. CORS — browser origins you allowed get their preflight answered.
  4. Metrics and request id — timing, X-Request-Id, the content-free usage record.
  5. Authentication — API key (or a signed request from an AI Suite app on the same computer), with throttling of repeated wrong keys, endpoint and model scopes and per-key rate limits.
  6. Free allowance — on Free only, for third-party tools.
  7. Rate limits — per key and per client address.
  8. Quotas and budgets — per key, per day and per month.
  9. Scheduling — interactive work first; background work waits or receives 503 with Retry-After.
  10. Policy checks — model lifecycle, content rules and moderation.
  11. Engine — the model runs; streamed output flows back through the same path.

Every stage that refuses answers with an OpenAI-shaped error, so clients handle refusals the same way.

Engines and model routing

A model id names its runtime (runtime/family/variant). The server's engine orchestrator sends each request to the backend for that runtime, starting the engine and loading the model when needed. Local engines run as separate processes on internal ports that are not exposed; the server is the only thing clients talk to. Cloud providers, if an operator configures them, are just another runtime.

Storage

WhatWhere
Control data: keys, licence lease, settings, governance, audit, logsThe data folder (%USERPROFILE%\AISuite\v2 on Windows, /data in containers)
Large caches: engines, modelsThe shared cache folder, which can be moved to another disk
Requests and answersMemory only, for the time of the request

On Windows the data folder is shared by the app, its service and the CLI, so all three see the same keys and licence.

Discovery and trust

  • Same computer: apps find the server through a lock file that records its address, port and scheme.
  • Local network: servers can announce themselves; apps list them and ask for an API key.
  • Certificates: with HTTPS on, apps pin the server's certificate fingerprint on first use and warn if it changes.

Gateway mode

AI Gateway farmClients send requests to the AI Gateway with their client key. The gateway checks keys and governance, then forwards each request to a healthy worker with the worker key. Workers can be added, drained or upgraded behind it.Clientsone URL · one keyclient keyAI Gatewayrouting · failover · governanceWorker 1GPU serverWorker 2GPU serverWorker 3 (canary)new model or buildworker key
In gateway mode one AI Server is the endpoint; workers do the inference. Clients keep one address and one key.

A gateway accepts client requests with the clients' keys, applies authentication, scopes, rate limits, quotas, lifecycle and content rules at the edge, then forwards each request to a worker with that worker's own key. Workers apply their own policies too. The gateway chooses a worker by health, load and which models it has installed; retries transient failures before any output is sent; and turns saturation into 503 with Retry-After. Details: gateway internals.

Process model and shutdown

  • On SIGTERM (or a service stop) the server reports not-ready on /readyz, waits for load balancers to notice, finishes in-flight work, hands back its licence seat and exits.
  • A server whose licence plan changes restarts itself (exit code 5) so the new plan applies everywhere at once.
  • Exit codes for refusals are listed in the configuration reference.

Questions

Is AI Server a wrapper around another product? +

AI Server is its own server: the API, authentication, governance, gateway and operations are ours. It runs established open-source inference runtimes as engines underneath, which it installs, updates and supervises.

Can I run the engines on different machines from the server? +

Use the gateway for that: each worker is a full AI Server with its engines, and the gateway spreads the work.