Governance is how you run AI Server for people you do not personally supervise: other teams, customers, or an organisation with policies. The controls are in the Windows app's Governance page and Server page, and in JSON files in the governance folder of the server's data folder for containers. Governance needs Pro Commercial; on lower plans the files are ignored.
| Control | What it does | Where |
|---|---|---|
| Rate limits | Requests per minute per API key and per client address | Server page, AISUITE_PER_KEY_RPM / AISUITE_PER_IP_RPM |
| Daily quotas | Requests per day per key | usage-quota.json |
| Monthly budgets | Estimated spend per key per month | usage-quota.json, model-pricing.json |
| Scheduling | Interactive work first; background work waits or runs in quiet hours | qos-policy.json |
| Content rules | Block requests or answers that match your patterns | content-filter.json |
| Moderation | Check text with an on-device model or an external endpoint | content-moderation.json |
| Model lifecycle | Approve, deprecate, block or pin models | model-lifecycle.json |
| Region | Residency of cloud providers, audit retention, audit shipping | region-policy.json |
| Governance mode | Whether client apps may download models through the server | Server page |
Changes in the files apply when the server restarts. Rate limits set on the Server page need the service to be reinstalled, which the app does for you.
Rate limits
AISUITE_PER_KEY_RPM and AISUITE_PER_IP_RPM (or the Server page) cap requests per minute. A key can carry its own limit (--rpm) that overrides the server-wide one. Over the limit, clients get 429 with Retry-After. Behind a proxy, set trusted proxies so the limit sees the real client address.
Quotas and budgets
{ "enabled": true, "daily_requests_per_key": 2000, "monthly_budget_per_key": 50 }
Only real work counts — chat, embeddings, images, audio, vision and model downloads — not listing models or monitoring. Streamed answers through a gateway count too. The day resets at midnight UTC and survives restarts. The budget uses the per-model prices you set in model-pricing.json, so you can charge teams back for the hardware they use; it is an estimate, not a bill.
Scheduling (quality of service)
Scheduling gives people at a keyboard priority over unattended work such as document indexing.
{
"enabled": true,
"max_concurrent_inference": 8,
"background_share_percent": 25,
"quiet_hours": { "enabled": true, "start": "20:00", "end": "07:00" }
}
AI Suite apps mark their background work. When the server is busy it answers background requests with 503 and Retry-After, and the apps pause their queues instead of retrying hard. Batch embedding jobs (/v1/batch/embeddings) always run as background work.
Content rules and moderation
Content rules are regular expressions with a category, checked on the request, the answer or both:
{
"enabled": true, "scan_input": true, "scan_output": true,
"rules": [ { "pattern": "\\b\\d{3}-\\d{2}-\\d{4}\\b", "category": "us-ssn" } ]
}
A blocked request gets 403 (forbidden); a blocked answer is replaced by the rule's replacement text, or refused with 403 when the rule has none. Audit records keep the category, never the matched text. Moderation adds a semantic check after the rules: the on-device classifier keeps text on the server; an external endpoint necessarily receives the text it checks, so choose it deliberately.
Model lifecycle
{
"enabled": true,
"models": [
{ "model_key": "enginea/llama3.2/3b", "status": "approved" },
{ "model_key": "enginea/llama3.1/8b", "status": "deprecated", "message": "Moving to Qwen 2.5 on 1 Nov", "replacement_key": "enginea/qwen2.5/7b" },
{ "model_key": "enginea/mistral/7b", "status": "blocked" }
]
}
A deprecated model still answers, with an X-AI-Model-Deprecated header so clients can warn; a blocked model is refused with 403. Pins hold a model family to a variant. With a gateway, lifecycle and content rules are applied at the gateway as well as on the workers.
Governance modes
| Mode | Client apps may… |
|---|---|
| Federated (default) | download models through the server |
| Curated | use only models an operator installed; downloads are refused |
| Open | use the server, and download models on their own devices |
Region policy
region-policy.json records the region (EU, UK, CA, BR, KR, JP, CN, US or Other), the regions cloud providers may run in, how long the audit log is kept, and an optional HTTPS endpoint that receives each closed day's audit file for your SIEM. In region CN the server refuses to start without content moderation configured.
Per-key limits
Model and endpoint allowlists, expiry and per-key rate limits are set on the key itself — see API keys.
Questions
Are governance settings available on Pro Personal? +
No. Rate limits, quotas, budgets, scheduling, moderation, lifecycle and audit signing need Pro Commercial. Personal includes API keys and usage analytics.
Does governance read prompts? +
Content rules and moderation inspect text in memory to decide; nothing is stored. Usage and audit records never contain prompt or answer text.