What AI Server does today, what we are working on next, what we are considering, and the changes already announced — without promised dates.
Plans change as we learn from customers, so items below carry no dates and are not commitments until they ship. Shipped changes are in the release notes. Tell us what matters to you through support.
Available today
OpenAI-compatible chat (streaming, tools, JSON output, image input), embeddings with batch jobs, image generation and editing, text-to-speech, transcription with subtitles, live captions, voice cloning and computer vision.
API keys with scopes and separate administration rights; HTTPS; governance (rate limits, quotas, budgets, scheduling, content rules, moderation, model lifecycle, region policy); signed audit export.
AI Gateway: five routing strategies, model-aware routing, failover, canary rollouts, DNS and local-network discovery, backpressure and zero-downtime upgrades.
Windows app with a Windows service; Docker images; Helm chart and Kustomize manifests; bring-your-own-licence offers on the Azure and AWS marketplaces.
Built-in dashboard, Prometheus metrics, usage reports and redacted logs.
Next
Item
Why
/v1/responses, /v1/completions and /v1/moderations
Newer OpenAI clients and older tools use them.
Image, speech and vision models listed in /v1/models with their capabilities
Clients discover every model type from one call.
Compressed audio (MP3, M4A) accepted by local transcription
Today local models take WAV only.
A published OpenAPI description of the API
Code generation and API gateways.
Downloading a model without making it active
Safer model installs on a busy shared server.
A tamper-evident chain across audit records
Detect a removed or edited record, not only an altered export.
macOS and Linux desktop apps
Today Linux runs the containers, and there is no macOS server app.
Being considered
Longer licence terms for sites with very restricted internet access (Enterprise).
Routing by data-residency zone and by organisation across a gateway farm.