Errors use OpenAI's envelope, so SDK error handling works unchanged:

{ "error": { "message": "Model 'llama3.2:3b' is not installed.", "type": "invalid_request_error", "code": "model_not_found", "param": "model" } }

Messages are written for people and may change; branch on code and the HTTP status.

Catalogue

StatuscodetypeMeaningRetry?
400invalid_request_errorinvalid_request_errorMalformed JSON, a missing or invalid field, or an unsupported option. param names the field.No — fix the request
401invalid_api_keyauthentication_errorMissing, wrong, revoked or expired key.No
402license_requiredinsufficient_quotaThe request needs a plan this server does not have.No
403forbiddenpermission_errorThe key may not use this model or endpoint, the model is blocked by policy, the request or answer was blocked by a content rule, or the request needs administration rights.No
403governance_denied—A model download in a governance mode that does not allow it.No
404model_not_foundinvalid_request_errorThe model is not installed or unknown.After installing it
404unknown_urlinvalid_request_errorNo such endpoint.No
429rate_limit_exceededrate_limit_exceededA rate limit, a daily quota, a budget, the Free allowance, or too many wrong keys.Yes, after Retry-After
500server_errorserver_errorAn unexpected failure. The message carries a reference you can quote; details are in the server log.Once
501no_provider_configurednot_implementedNo engine is installed for this kind of request.No
503engine_startingservice_unavailableThe engine is starting or loading the model.Yes, after Retry-After
503——All workers are busy, or this server is draining for an upgrade.Yes, after Retry-After

Extra headers help you decide:

  • Retry-After on 429 and 503 — seconds to wait.
  • X-AISuite-Upgrade: 1 on 429 when a paid plan would lift the limit (the Free allowance for third-party tools).

Retrying well

  • Retry only 429, 503 and network failures, waiting at least Retry-After, with exponential back-off and jitter.
  • Do not retry 4xx otherwise; the request will fail the same way.
  • Behind a gateway, failures on one worker before the answer starts are already retried on another worker for you.
  • Background jobs should back off longer than interactive ones. AI Suite apps park their queues when the server says it is busy; do the same.

The OpenAI SDKs retry 429 and 5xx automatically (twice by default) and honour Retry-After.

Errors in a stream

Once a streamed answer has started, the status code is already 200. If something fails mid-stream the server sends one more event, data: {"error": {…}} in the envelope above, then data: [DONE]. Events are always complete JSON. Check for an error field in each event if your client parses the stream itself.

Questions

Why do I get 401 from the same computer the server runs on? +

When the server listens on the network it requires a key for every request, including local ones, unless they come from AI Suite apps. Send the key, or switch the server to this computer only.

The message says "This API key is not authorized to use model…" +

The key has a model allowlist. Use an allowed model or ask the operator to widen the key.