AI Server

The private AI platform for teams. One server runs the models — every app and every device just uses them. AI Server serves chat, embeddings, image generation, speech and vision from hardware you own, through OpenAI-compatible APIs, to the whole AI Suite and to any compatible tool.

Start free on your own machine. When your team grows, serve your network; when one machine isn't enough, AI Gateway mode turns several servers into a single load-balanced endpoint — on your LAN, in Docker, or on Kubernetes. No vendor cloud in the data path at any step.

Get it on Microsoft Store See what it does ↓ Licensing & Pricing

WHAT AI SERVER DOES

One deployment powers everything

Without a server, every PC downloads its own models, needs its own GPU, and is configured and updated separately. With AI Server, models download once, run on the best hardware you have, and serve everyone.

Universal AI APIs

OpenAI-compatible endpoints for chat, embeddings, image generation, speech-to-text (including live streaming), text-to-speech, voice cloning, and vision — all served from your hardware. Point any compatible tool or your own code at it.

Powers the whole AI Suite

Every AI Suite app discovers the server and offloads its AI work to it — zero configuration. Notes, PDFs, translation, dictation, images, email and more, all sharing one set of models on one machine.

Governance built in

API keys with per-key plans, rate limits, quotas and budgets, usage analytics by app and model, and content-free audit logs with signed export. The controls IT expects, without a cloud vendor in the loop.

Scales with AI Gateway

Flip one switch and multiple servers become a single endpoint: health-checked load balancing, model-aware routing that prefers machines with the model already warm, canary rollouts, and zero-downtime upgrades.

Runs anywhere

A desktop app for the machine under the desk; containers, Kubernetes manifests, and a Helm chart for the rack or the cloud. Same product, same APIs, same license — from one laptop to a GPU farm.

Watchable by design

Every server hosts its own live dashboard and native Prometheus metrics: fleet health, latency, usage and capacity at a glance — no extra monitoring stack required to get started.

AI GATEWAY

From one server to a farm — without changing a single client

AI Gateway is AI Server running in its farm front-end mode. Clients connect to it exactly as if it were a single AI Server; behind it, a pool of worker servers does the actual inference. Add a machine and capacity grows; take one down for maintenance and traffic drains gracefully — nobody's request is dropped.

On Kubernetes, workers are discovered automatically as they scale, so an autoscaling farm needs no manual pool edits. The whole estate stays observable from the built-in dashboard and Prometheus metrics.

What the Gateway handles for you

  • Load balancing — round-robin, least-latency, least-connections, weighted, or sticky per client.
  • Model-aware routing — requests go to a machine that already has the requested model warm.
  • Failover — a failing worker is retried on a healthy one before the client ever notices.
  • Canary rollouts — send a consistent slice of traffic to new machines before trusting them fully.
  • Backpressure — when every worker is saturated, callers get an honest "retry shortly" instead of a timeout.
  • Zero-downtime upgrades — drain a worker, upgrade it, return it to the pool.
EDITIONS & TIERS

Start free. Grow when your team does.

The free tier is a real product, not a trial: a full local AI server for all AI Suite apps on your own machine, forever. Paid tiers begin exactly where value crosses machine boundaries.

Free

A complete local AI server on your own machine. Unlimited use by every AI Suite app on that machine; generic OpenAI-compatible clients are rate-limited.

Pro Personal

Serve your own devices across your network: LAN serving, run as an OS service, API keys for your laptop and phone, unlimited generic clients, and a small gateway pool.

Pro Commercial

Serve other people: commercial use, unlimited API keys, full governance (quotas, budgets, audit signing), and a production gateway farm with canary rollouts and automatic worker discovery.

Enterprise

Unlimited estates, organisation enrollment with AI Admin Console, offline and air-gapped operation, and priority support. Licensed per estate — talk to us.

PRIVATE BY ARCHITECTURE

Your prompts never leave your network — because there is nowhere for them to go

AI Server is not a proxy to a cloud API. The models run on your hardware; requests travel from your apps to your server and back, inside your perimeter. Offline and air-gapped operation are supported deployment modes, not special exceptions.

  • No vendor cloud in the data path. Prompts, documents and generated content stay between your clients and your server.
  • Content-free records. Usage analytics and audit logs record who called what, when — never the text of a prompt or a response.
  • Licensing that fails safe. If our licence service is ever unreachable, your licensed fleet keeps serving. Your uptime does not depend on ours.
  • Your keys, your rules. API keys are issued and revoked by you, on your server — not accounts in someone else's cloud.

Questions about AI Server or AI Gateway?

For deployment help, sizing advice, Kubernetes questions, or licensing, write to [email protected] — or use the contact form and pick "Product".

Get release updates

New free AI products, major updates, and a few releases available only via this site. No spam.