Errors use OpenAI's envelope, so SDK error handling works unchanged:
{ "error": { "message": "Model 'llama3.2:3b' is not installed.", "type": "invalid_request_error", "code": "model_not_found", "param": "model" } }
Messages are written for people and may change; branch on code and the HTTP status.
Catalogue
| Status | code | type | Meaning | Retry? |
|---|---|---|---|---|
| 400 | invalid_request_error | invalid_request_error | Malformed JSON, a missing or invalid field, or an unsupported option. param names the field. | No — fix the request |
| 401 | invalid_api_key | authentication_error | Missing, wrong, revoked or expired key. | No |
| 402 | license_required | insufficient_quota | The request needs a plan this server does not have. | No |
| 403 | forbidden | permission_error | The key may not use this model or endpoint, the model is blocked by policy, the request or answer was blocked by a content rule, or the request needs administration rights. | No |
| 403 | governance_denied | — | A model download in a governance mode that does not allow it. | No |
| 404 | model_not_found | invalid_request_error | The model is not installed or unknown. | After installing it |
| 404 | unknown_url | invalid_request_error | No such endpoint. | No |
| 429 | rate_limit_exceeded | rate_limit_exceeded | A rate limit, a daily quota, a budget, the Free allowance, or too many wrong keys. | Yes, after Retry-After |
| 500 | server_error | server_error | An unexpected failure. The message carries a reference you can quote; details are in the server log. | Once |
| 501 | no_provider_configured | not_implemented | No engine is installed for this kind of request. | No |
| 503 | engine_starting | service_unavailable | The engine is starting or loading the model. | Yes, after Retry-After |
| 503 | — | — | All workers are busy, or this server is draining for an upgrade. | Yes, after Retry-After |
Extra headers help you decide:
Retry-Afteron 429 and 503 — seconds to wait.X-AISuite-Upgrade: 1on 429 when a paid plan would lift the limit (the Free allowance for third-party tools).
Retrying well
- Retry only 429, 503 and network failures, waiting at least
Retry-After, with exponential back-off and jitter. - Do not retry 4xx otherwise; the request will fail the same way.
- Behind a gateway, failures on one worker before the answer starts are already retried on another worker for you.
- Background jobs should back off longer than interactive ones. AI Suite apps park their queues when the server says it is busy; do the same.
The OpenAI SDKs retry 429 and 5xx automatically (twice by default) and honour Retry-After.
Errors in a stream
Once a streamed answer has started, the status code is already 200. If something fails mid-stream the server sends one more event, data: {"error": {…}} in the envelope above, then data: [DONE]. Events are always complete JSON. Check for an error field in each event if your client parses the stream itself.
Questions
Why do I get 401 from the same computer the server runs on? +
When the server listens on the network it requires a key for every request, including local ones, unless they come from AI Suite apps. Send the key, or switch the server to this computer only.
The message says "This API key is not authorized to use model…" +
The key has a model allowlist. Use an allowed model or ask the operator to widen the key.