This page explains the behaviour you can observe from outside, for engineers debugging integrations or evaluating the design. It is not an internal API reference; the supported interface is the HTTP API.
Stack
AI Server is a .NET 10 ASP.NET Core application served by Kestrel. One binary runs as a server, an AI Gateway or a test stub. State lives in files in the data folder; there is no database to operate.
Request flow
Middleware runs in a fixed order — forwarded headers, host check, CORS, metrics and request id, authentication, the Free allowance, rate limits, quotas, scheduling — and endpoint handlers apply policy (model allowlist, lifecycle, content rules) before calling the engine orchestrator. A refusal at any stage short-circuits with an OpenAI-shaped error, and every request, refused or served, writes one content-free usage record.
Chat mapping
An OpenAI chat request is mapped to a turn for the engine:
- The first
systemordevelopermessage becomes the system prompt. - Everything up to the last user message becomes ordered history; the last user message is the prompt.
- If messages follow the last user message — an
assistantmessage withtool_callsand thetoolresults — the whole conversation is history and the prompt is empty, so the model continues after the tool results. Engines skip the empty user turn. image_urlparts become image inputs; the server checks the model supports images before running it.- Sampling options (
temperature,top_p,seed,stop, penalties,max_completion_tokens) are passed to engines that support them; output and context limits are capped on the server. response_formatis passed to engines that can constrain output;tool_choicenaming a function narrows the tools to that one and requires a call.
Streaming
Streams are Server-Sent Events written as the engine produces tokens. Tool calls stream as tool_calls deltas. When stream_options.include_usage is set a final chunk carries token counts. When an operator scans answers with content rules, the turn is generated in full, scanned, and then sent — the one case where streaming is not incremental. A failure after the stream starts ends it with an error event and [DONE].
Model resolution
A model id is either canonical (runtime/family/variant, used as is) or a wire id (qwen2.5:7b), which is matched against every model the engines report plus the embedding catalogue. Canonical ids never cause an engine to start just to resolve a name. Unknown ids answer 404.
Engines
The orchestrator owns one backend per runtime. Local engines are child processes on loopback ports; the orchestrator starts them on demand, health-checks them, and restarts them after a crash. Engine updates are downloaded beside the running version and promoted on the next start, under a cross-process lock so two processes never update the same engine at once. Model downloads for the same model are shared between callers.
What is persisted
| Data | File |
|---|---|
| API key hashes and scopes | credentials/server-keys.json |
| Licence lease | licensing/ |
| Settings | server-settings.json |
| Governance policies | governance/*.json |
| Usage and audit records | audit/ (daily JSONL) |
| Gateway pool | gateway.json |
| Discovery lock (address, port, scheme, process id) | server.lock |
| Logs | logs/ |
Prompts, answers, uploaded files and generated media are never written to disk by the server. Batch embedding jobs are held in memory and lost on restart by design.
Process lifecycle
- Start: resolve the data folder, activate or load the licence lease, decide the bind (falling back to this computer if there is no licence or no key), load policies, start listening, write
server.lock. - Plan change: a background check re-reads the licence every few minutes; if the plan changed, the server exits with code 5 and its supervisor starts it on the new plan.
- Stop:
/readyzturns 503, the server waits the drain period, finishes in-flight work, releases its licence seat and removesserver.lock.
Observability hooks
X-Request-Id is accepted from callers in a safe format and threaded through the audit record, so you can correlate client logs with the server. openai-processing-ms reports server time; X-AISuite-Backend names the gateway worker. See monitoring.