Performance depends on the model, its compression, the GPU, the length of prompts and answers, and how many requests run at once. Numbers measured on someone else's hardware tell you little about yours, so this page gives a method rather than a league table. Run it during your pilot.

What to measure

MetricWhy it mattersHow
Time to first token (TTFT)How quickly a person sees the answer startTime from sending a streamed request to the first content chunk
Generation speedHow fast the answer appears, per requestOutput tokens ÷ (end time − first-token time)
ThroughputHow much work the server does in totalTotal output tokens per second with N requests in flight
p95 latency at loadWhat the slowest users experience95th percentile of request duration at a given concurrency
Error and 503 rateWhere the server starts shedding loadCount of non-200 answers at each concurrency

Method

  1. Fix the variables: one model, the same prompt set, a fixed max_completion_tokens, temperature 0, the server otherwise idle.
  2. Warm up: send a few requests first so the model is loaded; the first load is not representative.
  3. Single-request baseline: run sequentially to measure TTFT and generation speed.
  4. Concurrency sweep: repeat at 1, 2, 4, 8, 16… requests in flight until p95 latency or the 503 rate becomes unacceptable. That knee is your capacity per worker.
  5. Use realistic prompts: a sample of your own chats or document sizes. Long prompts change the result more than anything else.
  6. Record the setup: server version, model id, GPU and driver, memory, operating system, and whether the run went through a gateway.

A script

import asyncio, time, statistics
from openai import AsyncOpenAI

client = AsyncOpenAI(base_url="http://<server>/v1", api_key="<api-key>")
MODEL, PROMPT, MAX_TOKENS = "qwen2.5:7b", "Summarise the benefits of a private AI server in 150 words.", 256

async def one():
    start = time.perf_counter(); first = None; tokens = 0
    stream = await client.chat.completions.create(model=MODEL, stream=True, temperature=0,
        max_completion_tokens=MAX_TOKENS, stream_options={"include_usage": True},
        messages=[{"role": "user", "content": PROMPT}])
    async for chunk in stream:
        if chunk.choices and chunk.choices[0].delta.content and first is None:
            first = time.perf_counter()
        if chunk.usage: tokens = chunk.usage.completion_tokens
    end = time.perf_counter()
    return first - start, tokens / max(end - first, 1e-6), end - start, tokens

async def run(concurrency, rounds=3):
    results = []
    for _ in range(rounds):
        started = time.perf_counter()
        batch = await asyncio.gather(*[one() for _ in range(concurrency)])
        wall = time.perf_counter() - started
        results.append((batch, sum(r[3] for r in batch) / wall))
    durations = sorted(r[2] for b, _ in results for r in b)
    print(f"c={concurrency:>3}  ttft={statistics.median(r[0] for b, _ in results for r in b):.2f}s  "
          f"gen={statistics.median(r[1] for b, _ in results for r in b):.1f} tok/s  "
          f"p95={durations[int(len(durations) * .95) - 1]:.1f}s  "
          f"throughput={statistics.median(t for _, t in results):.1f} tok/s")

async def main():
    await one()                      # warm-up
    for c in (1, 2, 4, 8, 16):
        await run(c)
asyncio.run(main())

Catch openai.APIStatusError with status 503 in longer runs to count shed requests rather than stopping.

Reading the results

  • TTFT grows with prompt length — the model reads the whole prompt before writing. Long documents need more GPU, or retrieval that sends only relevant passages.
  • Generation speed per request falls as concurrency rises, while throughput rises — until the GPU saturates. Set max_in_flight for the worker a little below the knee, so the gateway sheds load before latency degrades.
  • If throughput stops growing early, the model may not fit entirely in GPU memory; try a smaller or more compressed model.

The gateway's own overhead

A gateway adds a network hop and a routing decision, and no model work. Measure it by running the same script against a worker directly and through the gateway with the stub engine (AISUITE_ENGINE=stub on the workers), which removes model time from the result.

Questions

Do you publish benchmark results? +

Not as general claims, because results depend so heavily on hardware and workload. We are glad to help you plan and read a benchmark run on your hardware — contact support.

Why are my first requests slow? +

The first request after start loads the model into memory. Warm up before measuring, and keep frequently used models loaded on the workers that serve them.