AI Server does not ask you to install or update inference software. It downloads the engines a model needs, runs them as separate processes behind its own API, restarts them if they fail, and updates them in the background. Clients only ever see AI Server.
Engines
| Engine | Model id prefix | Used for | Runs on |
|---|---|---|---|
| Local AI Engine A | enginea/ | The default for chat, vision chat and embeddings; a broad catalogue of open models | Windows, Linux; CPU or GPU |
| Local AI Engine (llama.cpp) | engineb/ | Any GGUF model from Hugging Face; one model per engine process | Windows, Linux; CPU or GPU |
| Local AI Engine C | enginec/ | Models accelerated on the PC's NPU or GPU through Windows ML | Windows on the server |
| Pro AI Engine A | proenginea/ | High-throughput serving that spreads large models across GPU memory, system memory and CPU | Linux x64 with NVIDIA GPUs (aiserver-pro image); Pro subscription |
| Image generation | per model | Text-to-image, editing, upscaling, background removal | GPU strongly recommended |
| Speech | per model | Transcription (Whisper family), text-to-speech, voice cloning | CPU or GPU |
| Vision | per model | Object and face detection, background removal | CPU or GPU |
| Cloud providers | provider name | Models from OpenAI-compatible, Gemini or Anthropic accounts an operator adds | The provider's cloud |
Which engines are available depends on the platform and what the operator installs; GET /v1/models/catalog lists what this server can offer.
Model names
Every model has a canonical id runtime/family/variant — for example enginea/qwen2.5/7b — and a short wire id such as qwen2.5:7b. Requests accept either; see models.
Where models come from
Models are downloaded from their publishers' repositories (for example Hugging Face) the first time they are installed. The catalogue entry shows each model's source and licence. Operators can:
- choose which models users may download (curated mode) and approve, deprecate or block models (lifecycle policy) — see governance;
- move the model store to another disk from the Models page, without downloading again;
- on Windows, share one model store between AI Server and the AI Suite apps on the same PC.
Engine updates
Engines are updated from their upstream releases. A new release is downloaded while the current one keeps serving and applied the next time the engine starts; operators can switch automatic updates off per engine. See upgrades.
Cloud providers
An operator can add provider accounts on the Providers page. Their models then appear alongside local ones, and requests for them go to that provider under its terms. Provider credentials are stored encrypted on the server, so client apps never hold them. Nothing goes to a cloud provider unless an operator has added it and a user picks one of its models.
Choosing models
| Job | Starting point |
|---|---|
| Chat and drafting for a team | A 7–8B instruct model on a GPU |
| Agents with tools | A general (not coding-tuned) instruct model of 7–8B or larger |
| Document search | A small embedding model; same model for indexing and questions |
| Quick answers on modest hardware | A 1.5–3B model |
| Transcription | A Whisper-family model; larger is more accurate and slower |
Check each model's licence in the catalogue before approving it for business use.
Questions
Can I bring my own model? +
Through Local AI Engine (llama.cpp): any GGUF model from Hugging Face can be installed. Other formats need a matching engine.
Do engines open network ports? +
Local engines listen on internal ports on the server's own loopback interface. Only AI Server's port should be reachable by clients.