AI Server speaks the OpenAI REST API. Any OpenAI SDK, and any tool with an "OpenAI-compatible base URL" setting, works by changing two things: the base URL and the API key.

Base URL

Where the server runsBase URL
Windows app, same computerhttp://127.0.0.1:11436/v1
Windows app, another computerhttp://<server-address>:11436/v1 (or https:// with TLS)
Docker or Kuberneteshttp://<host>:8080/v1, or your ingress URL
AI GatewayThe gateway's URL — clients cannot tell a gateway from a single server

Authentication

Authorization: Bearer <api-key>

api-key: <key> and x-api-key: <key> are accepted too. On a server that listens on its own computer only, requests from that computer need no key. Everything else needs one — see API keys.

Model names

GET /v1/models lists what the server can serve. Every endpoint that takes a model accepts either form it reports:

  • the short wire id, such as llama3.2:3b;
  • the canonical runtime/family/variant, such as enginea/llama3.2/3b.

A model that is not installed answers 404 model_not_found. Install it with POST /v1/models/pull.

Endpoints

EndpointWhat it doesPage
POST /v1/chat/completionsChat, streaming, tool calling, JSON output, image inputChat
POST /v1/embeddingsText embeddingsEmbeddings
POST /v1/batch/embeddingsLarge embedding jobs run in the backgroundEmbeddings
POST /v1/images/generationsText-to-image, image-to-image, inpaintingImages
POST /v1/images/describeDescribe an image as a text promptImages
POST /v1/audio/speechText-to-speechAudio
POST /v1/audio/transcriptionsSpeech-to-text from a file, with subtitlesAudio
POST /v1/audio/translationsSpeech in any language to English textAudio
GET /v1/audio/transcriptions/streamLive speech-to-text over WebSocketAudio
POST /v1/audio/cloneSpeak text in the voice of a reference clipAudio
POST /v1/vision/detections, /faces, /backgroundObject and face detection, background removalVision
GET /v1/models, GET /v1/models/{id}List and describe modelsModels
POST /v1/models/pullDownload a modelModels

Not supported yet: /v1/responses, /v1/completions (legacy text completions), /v1/moderations, /v1/files, fine-tuning and assistants. An unknown path answers 404 unknown_url in the OpenAI error format.

Headers on every response

HeaderMeaning
X-Request-IdThe request's id — yours if you sent a well-formed X-Request-Id, else a new one. Quote it in bug reports.
X-AI-Server-VersionServer version.
openai-processing-msTime the server spent, in milliseconds.
X-AISuite-BackendBehind a gateway: the worker that answered.
Retry-AfterOn 429 and 503: seconds to wait before retrying.
X-AISuite-Upgrade: 1On a 429 that a paid plan would lift.
X-AI-Model-DeprecatedThe model is deprecated by the operator; a replacement may be named.

Your first request

curl http://127.0.0.1:11436/v1/chat/completions \
  -H "Authorization: Bearer $AISERVER_KEY" -H "Content-Type: application/json" \
  -d '{"model":"llama3.2:3b","messages":[{"role":"user","content":"Say hello in French."}]}'
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:11436/v1", api_key="<api-key>")
reply = client.chat.completions.create(model="llama3.2:3b",
    messages=[{"role": "user", "content": "Say hello in French."}])
print(reply.choices[0].message.content)

More languages and frameworks are on SDKs and frameworks.

Calling from a browser

Browsers block calls to another origin unless the server allows it. Set AISUITE_CORS_ORIGINS to the origins of your web app (for example https://intranet.example.com); the server then answers preflight requests without a key. Prefer calling AI Server from your backend anyway: a key in page JavaScript is visible to anyone who opens the developer tools.

Questions

Is there an OpenAPI specification? +

Not yet. The endpoints follow OpenAI's own specification; the differences are listed on each endpoint page.

Does AI Server need an internet connection to answer? +

No. Local models answer from the server's own hardware. Only model downloads, licence renewal and any cloud providers an operator configures go out.