The private AI platform for teams. One server runs the models — every app and every device just uses them. AI Server serves chat, embeddings, image generation, speech and vision from hardware you own, through OpenAI-compatible APIs, to the whole AI Suite and to any compatible tool.
Start free on your own machine. When your team grows, serve your network; when one machine isn't enough, AI Gateway mode turns several servers into a single load-balanced endpoint — on your LAN, in Docker, or on Kubernetes. No vendor cloud in the data path at any step.
Get it on Microsoft Store See what it does ↓ Licensing & Pricing
Without a server, every PC downloads its own models, needs its own GPU, and is configured and updated separately. With AI Server, models download once, run on the best hardware you have, and serve everyone.
OpenAI-compatible endpoints for chat, embeddings, image generation, speech-to-text (including live streaming), text-to-speech, voice cloning, and vision — all served from your hardware. Point any compatible tool or your own code at it.
Every AI Suite app discovers the server and offloads its AI work to it — zero configuration. Notes, PDFs, translation, dictation, images, email and more, all sharing one set of models on one machine.
API keys with per-key plans, rate limits, quotas and budgets, usage analytics by app and model, and content-free audit logs with signed export. The controls IT expects, without a cloud vendor in the loop.
Flip one switch and multiple servers become a single endpoint: health-checked load balancing, model-aware routing that prefers machines with the model already warm, canary rollouts, and zero-downtime upgrades.
A desktop app for the machine under the desk; containers, Kubernetes manifests, and a Helm chart for the rack or the cloud. Same product, same APIs, same license — from one laptop to a GPU farm.
Every server hosts its own live dashboard and native Prometheus metrics: fleet health, latency, usage and capacity at a glance — no extra monitoring stack required to get started.
AI Gateway is AI Server running in its farm front-end mode. Clients connect to it exactly as if it were a single AI Server; behind it, a pool of worker servers does the actual inference. Add a machine and capacity grows; take one down for maintenance and traffic drains gracefully — nobody's request is dropped.
On Kubernetes, workers are discovered automatically as they scale, so an autoscaling farm needs no manual pool edits. The whole estate stays observable from the built-in dashboard and Prometheus metrics.
The free tier is a real product, not a trial: a full local AI server for all AI Suite apps on your own machine, forever. Paid tiers begin exactly where value crosses machine boundaries.
A complete local AI server on your own machine. Unlimited use by every AI Suite app on that machine; generic OpenAI-compatible clients are rate-limited.
Serve your own devices across your network: LAN serving, run as an OS service, API keys for your laptop and phone, unlimited generic clients, and a small gateway pool.
Serve other people: commercial use, unlimited API keys, full governance (quotas, budgets, audit signing), and a production gateway farm with canary rollouts and automatic worker discovery.
Unlimited estates, organisation enrollment with AI Admin Console, offline and air-gapped operation, and priority support. Licensed per estate — talk to us.
AI Server is not a proxy to a cloud API. The models run on your hardware; requests travel from your apps to your server and back, inside your perimeter. Offline and air-gapped operation are supported deployment modes, not special exceptions.
For deployment help, sizing advice, Kubernetes questions, or licensing, write to [email protected] — or use the contact form and pick "Product".