Every limit answers in the same way: 429 with Retry-After for a limit on you, 503 with Retry-After when the server or the farm is busy. Your client handles both by waiting and retrying.

What can limit a request

LimitApplies toSet byResponse
Free allowanceThird-party tools on a Free server (curl, SDKs, chat UIs, IDE plugins)Fixed: 10 requests a day across those tools, 1 a minute per client429 with X-AISuite-Upgrade: 1
Rate limit per keyEach API keyOperator (server-wide, or per key with --rpm)429
Rate limit per client addressEach client IPOperator429
Daily quotaEach API keyOperator429 until midnight UTC
Monthly budgetEach API keyOperator, from estimated cost429 until next month
Wrong keysA client that keeps sending wrong keysBuilt in429 for a short time
Busy server or farmEveryone; background work firstOperator's scheduling policy503
Draining for an upgradeA server or worker being stoppedBuilt in503

AI Suite apps are never subject to the Free allowance. Pro Personal and Pro Commercial have no allowance at all; their limits are only the ones the operator sets.

What counts

Real work counts: chat, embeddings (a batch submission counts once), images, audio, vision and model downloads. Listing models, health probes, polling a batch job and monitoring do not.

How to respond

import random, time
import openai
def call_with_retry(fn, attempts=5):
    for attempt in range(attempts):
        try:
            return fn()
        except openai.RateLimitError as e:          # 429
            wait = float(e.response.headers.get("Retry-After", 1))
        except openai.APIStatusError as e:
            if e.status_code != 503: raise
            wait = float(e.response.headers.get("Retry-After", 2))
        time.sleep(wait + random.uniform(0, wait / 2) + attempt)
    raise RuntimeError("server still busy")
  • Respect Retry-After; never retry in a tight loop.
  • Mark unattended work with the header X-AISuite-Priority: background (or maintenance for the lowest priority), or run it as a batch job, so it waits while people are working. The header can only lower a request's priority, never raise it.
  • Spread large jobs over time instead of firing them all at once.

Checking your usage

The operator sees usage per key, model and app on the server's Usage page or through GET /v1/server/usage with an admin key. Ask them for your key's limits.

Questions

Why does my script get 429 after ten requests? +

The server is on the Free plan, where third-party tools share 10 requests a day. The response carries X-AISuite-Upgrade: 1. A Pro plan removes the allowance.

Is there a limit on prompt size? +

The model's context window limits it; the server caps context_window at 131072 tokens and output at 65536 tokens.