AI Server does not ask you to install or update inference software. It downloads the engines a model needs, runs them as separate processes behind its own API, restarts them if they fail, and updates them in the background. Clients only ever see AI Server.

Engines

EngineModel id prefixUsed forRuns on
Local AI Engine Aenginea/The default for chat, vision chat and embeddings; a broad catalogue of open modelsWindows, Linux; CPU or GPU
Local AI Engine (llama.cpp)engineb/Any GGUF model from Hugging Face; one model per engine processWindows, Linux; CPU or GPU
Local AI Engine Cenginec/Models accelerated on the PC's NPU or GPU through Windows MLWindows on the server
Pro AI Engine Aproenginea/High-throughput serving that spreads large models across GPU memory, system memory and CPULinux x64 with NVIDIA GPUs (aiserver-pro image); Pro subscription
Image generationper modelText-to-image, editing, upscaling, background removalGPU strongly recommended
Speechper modelTranscription (Whisper family), text-to-speech, voice cloningCPU or GPU
Visionper modelObject and face detection, background removalCPU or GPU
Cloud providersprovider nameModels from OpenAI-compatible, Gemini or Anthropic accounts an operator addsThe provider's cloud

Which engines are available depends on the platform and what the operator installs; GET /v1/models/catalog lists what this server can offer.

Model names

Every model has a canonical id runtime/family/variant — for example enginea/qwen2.5/7b — and a short wire id such as qwen2.5:7b. Requests accept either; see models.

Where models come from

Models are downloaded from their publishers' repositories (for example Hugging Face) the first time they are installed. The catalogue entry shows each model's source and licence. Operators can:

  • choose which models users may download (curated mode) and approve, deprecate or block models (lifecycle policy) — see governance;
  • move the model store to another disk from the Models page, without downloading again;
  • on Windows, share one model store between AI Server and the AI Suite apps on the same PC.

Engine updates

Engines are updated from their upstream releases. A new release is downloaded while the current one keeps serving and applied the next time the engine starts; operators can switch automatic updates off per engine. See upgrades.

Cloud providers

An operator can add provider accounts on the Providers page. Their models then appear alongside local ones, and requests for them go to that provider under its terms. Provider credentials are stored encrypted on the server, so client apps never hold them. Nothing goes to a cloud provider unless an operator has added it and a user picks one of its models.

Choosing models

JobStarting point
Chat and drafting for a teamA 7–8B instruct model on a GPU
Agents with toolsA general (not coding-tuned) instruct model of 7–8B or larger
Document searchA small embedding model; same model for indexing and questions
Quick answers on modest hardwareA 1.5–3B model
TranscriptionA Whisper-family model; larger is more accurate and slower

Check each model's licence in the catalogue before approving it for business use.

Questions

Can I bring my own model? +

Through Local AI Engine (llama.cpp): any GGUF model from Hugging Face can be installed. Other formats need a matching engine.

Do engines open network ports? +

Local engines listen on internal ports on the server's own loopback interface. Only AI Server's port should be reachable by clients.