TermMeaning
AI GatewayA mode of AI Server that receives all requests at one address and forwards each to a worker. Pro Commercial.
AI ClientA lightweight desktop app with AI workspaces that uses an AI Server instead of the PC's own hardware.
AI Suite appsSoftware Tailor's desktop AI apps; they find and use AI Server automatically.
API keyA secret that identifies a person, device or app to the server. Shown once; stored only as a hash.
Administration rightPermission on a key to read usage for all keys, metrics, request activity and audit exports.
Audit logA content-free record of every request: time, endpoint, model, key, status, timing, tokens.
BackpressureAnswering "busy, retry after N seconds" (503 with Retry-After) instead of queuing until a timeout.
Batch jobA large embedding request run in the background, submitted and then polled.
CanaryA worker running a new model or version that receives a fixed share of callers first.
Circuit breakerTaking a failing worker out of the pool for a short time so it is not hammered.
Content rulesPatterns that block requests or answers; records keep the rule's category, never the text.
Context windowHow much text, in tokens, a model can consider at once — the prompt plus the answer.
Data folderWhere the server keeps keys, licence lease, settings, policies, audit and logs.
DrainStopping new work on a server while it finishes what is in flight, before shutdown.
EmbeddingA list of numbers representing the meaning of a text, used for search and RAG.
EngineThe inference software that runs a model; AI Server installs and supervises engines.
Free allowanceOn the Free edition, the small daily number of requests from tools other than AI Suite apps.
GGUFA file format for compressed open models, used by llama.cpp-based engines.
GovernancePolicies on who may use which models and how much: limits, quotas, budgets, lifecycle, content rules.
Governance modeWhether client apps may download models through the server: federated, curated or open.
LeaseA signed, time-limited licence record the server receives at activation and renews weekly.
Lifecycle policyThe operator's list of approved, deprecated and blocked models.
LoopbackThe computer's internal network interface (127.0.0.1); a server bound to it serves only that computer.
Model idThe name of a model in requests: canonical runtime/family/variant or a short wire id.
OpenAI-compatibleAccepting the same requests and returning the same responses as OpenAI's API.
QuantizationStoring model weights with fewer bits so models need less memory, at a small cost in quality.
QuotaA limit on requests per key per day.
RAGRetrieval-augmented generation: finding relevant passages and giving them to the model as context.
Rate limitA limit on requests per minute, per key or per client address.
ScopeThe models and kinds of request a key is allowed.
SeatOne licensed running server. Gateways use no seat.
Stub engineA test engine that answers with canned text, for checking a deployment without models.
TTFTTime to first token: how long until a streamed answer starts.
WorkerAn AI Server behind a gateway that runs the models.