| Term | Meaning |
|---|---|
| AI Gateway | A mode of AI Server that receives all requests at one address and forwards each to a worker. Pro Commercial. |
| AI Client | A lightweight desktop app with AI workspaces that uses an AI Server instead of the PC's own hardware. |
| AI Suite apps | Software Tailor's desktop AI apps; they find and use AI Server automatically. |
| API key | A secret that identifies a person, device or app to the server. Shown once; stored only as a hash. |
| Administration right | Permission on a key to read usage for all keys, metrics, request activity and audit exports. |
| Audit log | A content-free record of every request: time, endpoint, model, key, status, timing, tokens. |
| Backpressure | Answering "busy, retry after N seconds" (503 with Retry-After) instead of queuing until a timeout. |
| Batch job | A large embedding request run in the background, submitted and then polled. |
| Canary | A worker running a new model or version that receives a fixed share of callers first. |
| Circuit breaker | Taking a failing worker out of the pool for a short time so it is not hammered. |
| Content rules | Patterns that block requests or answers; records keep the rule's category, never the text. |
| Context window | How much text, in tokens, a model can consider at once — the prompt plus the answer. |
| Data folder | Where the server keeps keys, licence lease, settings, policies, audit and logs. |
| Drain | Stopping new work on a server while it finishes what is in flight, before shutdown. |
| Embedding | A list of numbers representing the meaning of a text, used for search and RAG. |
| Engine | The inference software that runs a model; AI Server installs and supervises engines. |
| Free allowance | On the Free edition, the small daily number of requests from tools other than AI Suite apps. |
| GGUF | A file format for compressed open models, used by llama.cpp-based engines. |
| Governance | Policies on who may use which models and how much: limits, quotas, budgets, lifecycle, content rules. |
| Governance mode | Whether client apps may download models through the server: federated, curated or open. |
| Lease | A signed, time-limited licence record the server receives at activation and renews weekly. |
| Lifecycle policy | The operator's list of approved, deprecated and blocked models. |
| Loopback | The computer's internal network interface (127.0.0.1); a server bound to it serves only that computer. |
| Model id | The name of a model in requests: canonical runtime/family/variant or a short wire id. |
| OpenAI-compatible | Accepting the same requests and returning the same responses as OpenAI's API. |
| Quantization | Storing model weights with fewer bits so models need less memory, at a small cost in quality. |
| Quota | A limit on requests per key per day. |
| RAG | Retrieval-augmented generation: finding relevant passages and giving them to the model as context. |
| Rate limit | A limit on requests per minute, per key or per client address. |
| Scope | The models and kinds of request a key is allowed. |
| Seat | One licensed running server. Gateways use no seat. |
| Stub engine | A test engine that answers with canned text, for checking a deployment without models. |
| TTFT | Time to first token: how long until a streamed answer starts. |
| Worker | An AI Server behind a gateway that runs the models. |
ResourcesReviewed 2026-10-05 · Beginner
AI Server glossary
Plain definitions of the terms used in the AI Server documentation — from API key and lease to worker, canary, embedding and quantization.