This page explains the behaviour you can observe from outside, for engineers debugging integrations or evaluating the design. It is not an internal API reference; the supported interface is the HTTP API.

Stack

AI Server is a .NET 10 ASP.NET Core application served by Kestrel. One binary runs as a server, an AI Gateway or a test stub. State lives in files in the data folder; there is no database to operate.

Request flow

Request pipelineA request passes forwarded-header handling and the host check, API key authentication and scopes, the Free allowance, rate limits, quotas and budgets, scheduling, content and lifecycle checks, and then the engine. Usage and audit are recorded content-free.Host checkAPI keyRate limitsQuota and budgetSchedulingPolicy checksEngineUsage and audit records: endpoint, model, key, status, latency, tokens — never prompt or response text
Every request passes the same checks in this order before a model sees it.

Middleware runs in a fixed order — forwarded headers, host check, CORS, metrics and request id, authentication, the Free allowance, rate limits, quotas, scheduling — and endpoint handlers apply policy (model allowlist, lifecycle, content rules) before calling the engine orchestrator. A refusal at any stage short-circuits with an OpenAI-shaped error, and every request, refused or served, writes one content-free usage record.

Chat mapping

An OpenAI chat request is mapped to a turn for the engine:

  • The first system or developer message becomes the system prompt.
  • Everything up to the last user message becomes ordered history; the last user message is the prompt.
  • If messages follow the last user message — an assistant message with tool_calls and the tool results — the whole conversation is history and the prompt is empty, so the model continues after the tool results. Engines skip the empty user turn.
  • image_url parts become image inputs; the server checks the model supports images before running it.
  • Sampling options (temperature, top_p, seed, stop, penalties, max_completion_tokens) are passed to engines that support them; output and context limits are capped on the server.
  • response_format is passed to engines that can constrain output; tool_choice naming a function narrows the tools to that one and requires a call.

Streaming

Streams are Server-Sent Events written as the engine produces tokens. Tool calls stream as tool_calls deltas. When stream_options.include_usage is set a final chunk carries token counts. When an operator scans answers with content rules, the turn is generated in full, scanned, and then sent — the one case where streaming is not incremental. A failure after the stream starts ends it with an error event and [DONE].

Model resolution

A model id is either canonical (runtime/family/variant, used as is) or a wire id (qwen2.5:7b), which is matched against every model the engines report plus the embedding catalogue. Canonical ids never cause an engine to start just to resolve a name. Unknown ids answer 404.

Engines

The orchestrator owns one backend per runtime. Local engines are child processes on loopback ports; the orchestrator starts them on demand, health-checks them, and restarts them after a crash. Engine updates are downloaded beside the running version and promoted on the next start, under a cross-process lock so two processes never update the same engine at once. Model downloads for the same model are shared between callers.

What is persisted

DataFile
API key hashes and scopescredentials/server-keys.json
Licence leaselicensing/
Settingsserver-settings.json
Governance policiesgovernance/*.json
Usage and audit recordsaudit/ (daily JSONL)
Gateway poolgateway.json
Discovery lock (address, port, scheme, process id)server.lock
Logslogs/

Prompts, answers, uploaded files and generated media are never written to disk by the server. Batch embedding jobs are held in memory and lost on restart by design.

Process lifecycle

  • Start: resolve the data folder, activate or load the licence lease, decide the bind (falling back to this computer if there is no licence or no key), load policies, start listening, write server.lock.
  • Plan change: a background check re-reads the licence every few minutes; if the plan changed, the server exits with code 5 and its supervisor starts it on the new plan.
  • Stop: /readyz turns 503, the server waits the drain period, finishes in-flight work, releases its licence seat and removes server.lock.

Observability hooks

X-Request-Id is accepted from callers in a safe format and threaded through the audit record, so you can correlate client logs with the server. openai-processing-ms reports server time; X-AISuite-Backend names the gateway worker. See monitoring.