Every limit answers in the same way: 429 with Retry-After for a limit on you, 503 with Retry-After when the server or the farm is busy. Your client handles both by waiting and retrying.
What can limit a request
| Limit | Applies to | Set by | Response |
|---|---|---|---|
| Free allowance | Third-party tools on a Free server (curl, SDKs, chat UIs, IDE plugins) | Fixed: 10 requests a day across those tools, 1 a minute per client | 429 with X-AISuite-Upgrade: 1 |
| Rate limit per key | Each API key | Operator (server-wide, or per key with --rpm) | 429 |
| Rate limit per client address | Each client IP | Operator | 429 |
| Daily quota | Each API key | Operator | 429 until midnight UTC |
| Monthly budget | Each API key | Operator, from estimated cost | 429 until next month |
| Wrong keys | A client that keeps sending wrong keys | Built in | 429 for a short time |
| Busy server or farm | Everyone; background work first | Operator's scheduling policy | 503 |
| Draining for an upgrade | A server or worker being stopped | Built in | 503 |
AI Suite apps are never subject to the Free allowance. Pro Personal and Pro Commercial have no allowance at all; their limits are only the ones the operator sets.
What counts
Real work counts: chat, embeddings (a batch submission counts once), images, audio, vision and model downloads. Listing models, health probes, polling a batch job and monitoring do not.
How to respond
import random, time
import openai
def call_with_retry(fn, attempts=5):
for attempt in range(attempts):
try:
return fn()
except openai.RateLimitError as e: # 429
wait = float(e.response.headers.get("Retry-After", 1))
except openai.APIStatusError as e:
if e.status_code != 503: raise
wait = float(e.response.headers.get("Retry-After", 2))
time.sleep(wait + random.uniform(0, wait / 2) + attempt)
raise RuntimeError("server still busy")
- Respect
Retry-After; never retry in a tight loop. - Mark unattended work with the header
X-AISuite-Priority: background(ormaintenancefor the lowest priority), or run it as a batch job, so it waits while people are working. The header can only lower a request's priority, never raise it. - Spread large jobs over time instead of firing them all at once.
Checking your usage
The operator sees usage per key, model and app on the server's Usage page or through GET /v1/server/usage with an admin key. Ask them for your key's limits.
Questions
Why does my script get 429 after ten requests? +
The server is on the Free plan, where third-party tools share 10 requests a day. The response carries X-AISuite-Upgrade: 1. A Pro plan removes the allowance.
Is there a limit on prompt size? +
The model's context window limits it; the server caps context_window at 131072 tokens and output at 65536 tokens.