POST /v1/chat/completions takes the same request as OpenAI's Chat Completions API and returns the same response shape, streamed or not.
curl http://<server>/v1/chat/completions \
-H "Authorization: Bearer $AISERVER_KEY" -H "Content-Type: application/json" \
-d '{
"model": "llama3.2:3b",
"messages": [
{"role": "system", "content": "You answer in one sentence."},
{"role": "user", "content": "What is a GPU?"}
],
"temperature": 0.3
}'
Request parameters
| Parameter | Supported | Notes |
|---|---|---|
model | Yes | A wire id or runtime/family/variant from GET /v1/models. |
messages | Yes | Roles system, developer (treated as system), user, assistant, tool. Content is a string or an array of text and image_url parts. |
stream, stream_options.include_usage | Yes | Server-Sent Events; usage arrives in a final chunk when requested. |
max_completion_tokens, max_tokens | Yes | max_completion_tokens wins when both are sent. Capped at 65536. |
temperature, top_p, seed | Yes | |
stop | Yes | A string or an array of strings. |
frequency_penalty, presence_penalty | Yes | Passed to engines that support them. |
response_format | Yes | {"type":"json_object"} or {"type":"json_schema", …} — see JSON output. |
tools, tool_choice | Yes | auto, none, required, or {"type":"function","function":{"name":…}}. |
user | Accepted | Not stored as text. |
repetition_penalty | Extension | For local engines that use it. |
context_window | Extension | Context length to load the model with, capped at 131072 tokens. |
n > 1, logprobs, logit_bias, audio output, parallel_tool_calls | No | Ignored. |
Response
{
"id": "chatcmpl-…", "object": "chat.completion", "created": 1759650000, "model": "enginea/llama3.2/3b",
"choices": [ { "index": 0, "message": { "role": "assistant", "content": "A GPU is …" }, "finish_reason": "stop" } ],
"usage": { "prompt_tokens": 24, "completion_tokens": 18, "total_tokens": 42 }
}
finish_reason is stop, length or tool_calls.
Streaming
With "stream": true the answer arrives as chat.completion.chunk events with a delta, ending with data: [DONE]. Every OpenAI SDK handles this unchanged.
stream = client.chat.completions.create(model="llama3.2:3b", stream=True,
stream_options={"include_usage": True},
messages=[{"role": "user", "content": "Write a haiku about servers."}])
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")
If you put a proxy in front of AI Server, turn off response buffering or the stream arrives all at once. When an operator has turned on content rules for answers, the server checks the complete answer before sending it, so the stream arrives in one burst at the end.
Tool calling
Send tools; the model answers with tool_calls instead of text when it wants one. Your code runs the tool and sends the result back as a tool message; the server keeps the whole conversation, in order, and the model continues.
tools = [{"type": "function", "function": {
"name": "get_weather", "description": "Current weather for a city",
"parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]}}}]
messages = [{"role": "user", "content": "Do I need an umbrella in Hong Kong?"}]
reply = client.chat.completions.create(model="qwen2.5:7b", messages=messages, tools=tools)
call = reply.choices[0].message.tool_calls[0]
messages.append(reply.choices[0].message)
messages.append({"role": "tool", "tool_call_id": call.id, "content": '{"rain_mm": 12}'})
final = client.chat.completions.create(model="qwen2.5:7b", messages=messages, tools=tools)
print(final.choices[0].message.content)
- Use a general instruct model of 7–8B parameters or more for multi-step agents. Coding-tuned and very small models tend to write the call as plain text.
tool_choice: "required"forces a structured call; naming a function forces that one.- Streaming returns tool calls as
tool_callsdeltas.
JSON output
{"type": "json_object"} asks for valid JSON; {"type": "json_schema", "json_schema": {"name": …, "schema": {…}}} passes your schema to engines that can constrain their output to it. Always parse and validate the answer — smaller models can still produce JSON that does not match the schema.
Images in a message
Send images as image_url parts with a base64 data URL. The model must support images (vision); others answer 400.
{"role": "user", "content": [
{"type": "text", "text": "What is in this photo?"},
{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,/9j/4AAQ…"}}
]}
Remote https:// image URLs are passed on only to cloud providers; local models need the image inline.
Limits and errors
- Over a rate limit, quota or the Free allowance: 429 with
Retry-After. - A model the key may not use: 403. A model that is not installed: 404
model_not_found. - A deprecated model answers with the
X-AI-Model-Deprecatedheader. - Behind a gateway, a request that fails on one worker before any output is sent is retried on another.
See the error reference and rate limits.
Questions
Does the server remember conversations? +
No. Like OpenAI's API it is stateless: send the conversation so far with each request.
Can I use cloud models through the same endpoint? +
Yes, if an operator added a cloud provider on the server's Providers page. Those requests go to that provider; local models stay on the server.