Performance depends on the model, its compression, the GPU, the length of prompts and answers, and how many requests run at once. Numbers measured on someone else's hardware tell you little about yours, so this page gives a method rather than a league table. Run it during your pilot.
What to measure
| Metric | Why it matters | How |
|---|---|---|
| Time to first token (TTFT) | How quickly a person sees the answer start | Time from sending a streamed request to the first content chunk |
| Generation speed | How fast the answer appears, per request | Output tokens ÷ (end time − first-token time) |
| Throughput | How much work the server does in total | Total output tokens per second with N requests in flight |
| p95 latency at load | What the slowest users experience | 95th percentile of request duration at a given concurrency |
| Error and 503 rate | Where the server starts shedding load | Count of non-200 answers at each concurrency |
Method
- Fix the variables: one model, the same prompt set, a fixed
max_completion_tokens, temperature 0, the server otherwise idle. - Warm up: send a few requests first so the model is loaded; the first load is not representative.
- Single-request baseline: run sequentially to measure TTFT and generation speed.
- Concurrency sweep: repeat at 1, 2, 4, 8, 16… requests in flight until p95 latency or the 503 rate becomes unacceptable. That knee is your capacity per worker.
- Use realistic prompts: a sample of your own chats or document sizes. Long prompts change the result more than anything else.
- Record the setup: server version, model id, GPU and driver, memory, operating system, and whether the run went through a gateway.
A script
import asyncio, time, statistics
from openai import AsyncOpenAI
client = AsyncOpenAI(base_url="http://<server>/v1", api_key="<api-key>")
MODEL, PROMPT, MAX_TOKENS = "qwen2.5:7b", "Summarise the benefits of a private AI server in 150 words.", 256
async def one():
start = time.perf_counter(); first = None; tokens = 0
stream = await client.chat.completions.create(model=MODEL, stream=True, temperature=0,
max_completion_tokens=MAX_TOKENS, stream_options={"include_usage": True},
messages=[{"role": "user", "content": PROMPT}])
async for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content and first is None:
first = time.perf_counter()
if chunk.usage: tokens = chunk.usage.completion_tokens
end = time.perf_counter()
return first - start, tokens / max(end - first, 1e-6), end - start, tokens
async def run(concurrency, rounds=3):
results = []
for _ in range(rounds):
started = time.perf_counter()
batch = await asyncio.gather(*[one() for _ in range(concurrency)])
wall = time.perf_counter() - started
results.append((batch, sum(r[3] for r in batch) / wall))
durations = sorted(r[2] for b, _ in results for r in b)
print(f"c={concurrency:>3} ttft={statistics.median(r[0] for b, _ in results for r in b):.2f}s "
f"gen={statistics.median(r[1] for b, _ in results for r in b):.1f} tok/s "
f"p95={durations[int(len(durations) * .95) - 1]:.1f}s "
f"throughput={statistics.median(t for _, t in results):.1f} tok/s")
async def main():
await one() # warm-up
for c in (1, 2, 4, 8, 16):
await run(c)
asyncio.run(main())
Catch openai.APIStatusError with status 503 in longer runs to count shed requests rather than stopping.
Reading the results
- TTFT grows with prompt length — the model reads the whole prompt before writing. Long documents need more GPU, or retrieval that sends only relevant passages.
- Generation speed per request falls as concurrency rises, while throughput rises — until the GPU saturates. Set
max_in_flightfor the worker a little below the knee, so the gateway sheds load before latency degrades. - If throughput stops growing early, the model may not fit entirely in GPU memory; try a smaller or more compressed model.
The gateway's own overhead
A gateway adds a network hop and a routing decision, and no model work. Measure it by running the same script against a worker directly and through the gateway with the stub engine (AISUITE_ENGINE=stub on the workers), which removes model time from the result.
Questions
Do you publish benchmark results? +
Not as general claims, because results depend so heavily on hardware and workload. We are glad to help you plan and read a benchmark run on your hardware — contact support.
Why are my first requests slow? +
The first request after start loads the model into memory. Warm up before measuring, and keep frequently used models loaded on the workers that serve them.