AI Server speaks the OpenAI REST API. Any OpenAI SDK, and any tool with an "OpenAI-compatible base URL" setting, works by changing two things: the base URL and the API key.
Base URL
| Where the server runs | Base URL |
|---|---|
| Windows app, same computer | http://127.0.0.1:11436/v1 |
| Windows app, another computer | http://<server-address>:11436/v1 (or https:// with TLS) |
| Docker or Kubernetes | http://<host>:8080/v1, or your ingress URL |
| AI Gateway | The gateway's URL — clients cannot tell a gateway from a single server |
Authentication
Authorization: Bearer <api-key>
api-key: <key> and x-api-key: <key> are accepted too. On a server that listens on its own computer only, requests from that computer need no key. Everything else needs one — see API keys.
Model names
GET /v1/models lists what the server can serve. Every endpoint that takes a model accepts either form it reports:
- the short wire id, such as
llama3.2:3b; - the canonical runtime/family/variant, such as
enginea/llama3.2/3b.
A model that is not installed answers 404 model_not_found. Install it with POST /v1/models/pull.
Endpoints
| Endpoint | What it does | Page |
|---|---|---|
POST /v1/chat/completions | Chat, streaming, tool calling, JSON output, image input | Chat |
POST /v1/embeddings | Text embeddings | Embeddings |
POST /v1/batch/embeddings | Large embedding jobs run in the background | Embeddings |
POST /v1/images/generations | Text-to-image, image-to-image, inpainting | Images |
POST /v1/images/describe | Describe an image as a text prompt | Images |
POST /v1/audio/speech | Text-to-speech | Audio |
POST /v1/audio/transcriptions | Speech-to-text from a file, with subtitles | Audio |
POST /v1/audio/translations | Speech in any language to English text | Audio |
GET /v1/audio/transcriptions/stream | Live speech-to-text over WebSocket | Audio |
POST /v1/audio/clone | Speak text in the voice of a reference clip | Audio |
POST /v1/vision/detections, /faces, /background | Object and face detection, background removal | Vision |
GET /v1/models, GET /v1/models/{id} | List and describe models | Models |
POST /v1/models/pull | Download a model | Models |
Not supported yet: /v1/responses, /v1/completions (legacy text completions), /v1/moderations, /v1/files, fine-tuning and assistants. An unknown path answers 404 unknown_url in the OpenAI error format.
Headers on every response
| Header | Meaning |
|---|---|
X-Request-Id | The request's id — yours if you sent a well-formed X-Request-Id, else a new one. Quote it in bug reports. |
X-AI-Server-Version | Server version. |
openai-processing-ms | Time the server spent, in milliseconds. |
X-AISuite-Backend | Behind a gateway: the worker that answered. |
Retry-After | On 429 and 503: seconds to wait before retrying. |
X-AISuite-Upgrade: 1 | On a 429 that a paid plan would lift. |
X-AI-Model-Deprecated | The model is deprecated by the operator; a replacement may be named. |
Your first request
curl http://127.0.0.1:11436/v1/chat/completions \
-H "Authorization: Bearer $AISERVER_KEY" -H "Content-Type: application/json" \
-d '{"model":"llama3.2:3b","messages":[{"role":"user","content":"Say hello in French."}]}'
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:11436/v1", api_key="<api-key>")
reply = client.chat.completions.create(model="llama3.2:3b",
messages=[{"role": "user", "content": "Say hello in French."}])
print(reply.choices[0].message.content)
More languages and frameworks are on SDKs and frameworks.
Calling from a browser
Browsers block calls to another origin unless the server allows it. Set AISUITE_CORS_ORIGINS to the origins of your web app (for example https://intranet.example.com); the server then answers preflight requests without a key. Prefer calling AI Server from your backend anyway: a key in page JavaScript is visible to anyone who opens the developer tools.
Questions
Is there an OpenAPI specification? +
Not yet. The endpoints follow OpenAI's own specification; the differences are listed on each endpoint page.
Does AI Server need an internet connection to answer? +
No. Local models answer from the server's own hardware. Only model downloads, licence renewal and any cloud providers an operator configures go out.