Every AI Server node — a single server, a worker or a gateway — monitors itself. There is no agent or exporter to install. None of these views ever contain prompt or answer text.
| What | Where | Who can read it |
|---|---|---|
| Live dashboard | GET /dashboard in a browser | The page is public and empty; its data needs an admin key off this computer |
| Prometheus metrics | GET /v1/server/metrics | Admin key |
| Usage report | GET /v1/server/usage, the app's Usage page | Admin key |
| Live request activity | The app's Active requests page | This computer |
| Signed audit export | GET /v1/server/audit/export, aisuite-server-cli export-audit | Admin key |
| Gateway pool status | GET /v1/gateway/workers | Admin key |
| Health probes | GET /livez, GET /readyz | Anyone |
| Logs | stdout, and logs/aisuite-server-*.log outside containers | Server administrators |
An admin key is an API key issued with administration rights (aisuite-server-cli keys add --label monitoring --admin, or Allow server administration on the API keys page). Give monitoring tools their own key: rotating a client key should never break monitoring. See API keys.
Dashboard
Open http://<server>:<port>/dashboard. It shows whether the server is ready or draining, request rate, latency percentiles, the busiest endpoints, usage by app and model, and on a gateway every worker with its health, load against its cap, latency, canary flag and installed models. Paste an admin key into the page when you open it from another computer. Set AISUITE_DASHBOARD=0 to remove it.
Prometheus
scrape_configs:
- job_name: aisuite-gateway
metrics_path: /v1/server/metrics
authorization: { type: Bearer, credentials: <admin-key> }
static_configs: [ { targets: ["gateway.example.internal:8080"] } ]
| Series | Type | Meaning |
|---|---|---|
aisuite_requests_total | counter | Requests served. |
aisuite_errors_total | counter | Responses with status 500 or above. A 4xx (wrong key, quota) is a client problem, not an outage. |
aisuite_requests_by_endpoint_total | counter | Requests per endpoint. |
aisuite_requests_by_status_total | counter | Requests per HTTP status. |
aisuite_request_duration_seconds | summary | Latency per endpoint (p50, p95, p99). |
aisuite_gateway_pool_workers, …_pool_workers_healthy | gauge | Workers in the pool, and how many are healthy. |
aisuite_gateway_pool_in_flight, …_pool_in_flight_peak | gauge | Requests in progress across the pool. |
aisuite_gateway_worker_healthy, …_breaker_open, …_consecutive_failures | gauge | Per-worker health and circuit breaker state. |
aisuite_gateway_worker_in_flight, …_max_in_flight | gauge | Per-worker load and its cap — the right autoscaling signal. |
aisuite_gateway_worker_ewma_latency_seconds | gauge | Smoothed per-worker latency. |
aisuite_gateway_worker_requests_total | counter | Requests sent to each worker. |
aisuite_gateway_pool_info, aisuite_gateway_worker_info | gauge | Labels: routing strategy, worker id, canary. |
Starter alerts worth having:
- No healthy workers:
aisuite_gateway_pool_workers_healthy == 0for 2 minutes. This is what users feel. - Target down:
up == 0for 5 minutes — also fires when the scrape key was revoked. - Error rate:
aisuite_errors_totaloveraisuite_requests_totalabove 5% for 10 minutes.
Do not alert on /readyz returning 503: a draining pod does that on every rollout.
Usage reports
GET /v1/server/usage?from=2026-10-01&to=2026-10-31&by=key returns request counts, failures, latency, input and output tokens and an estimated cost, grouped by key, model, app or (on a gateway) worker. The default is the last 30 days by key. Costs use the per-model prices in governance/model-pricing.json, so they can feed a charge-back to teams.
Audit log
The server writes one content-free line per request — time, endpoint, model, key id, app, status, latency and token counts — to daily files in the audit folder, kept for the retention you set in the region policy.
For an auditor, export a signed bundle and let them check it offline:
aisuite-server-cli export-audit --from 2026-10-01 --to 2026-10-31 --out audit-october.json
aisuite-server-cli verify-audit --in audit-october.json # exit 0 = PASS, 2 = FAIL
Bundles are signed with ES256 using a key that stays on the server; the public key is in the bundle and at GET /v1/server/audit/public-key. rotate-audit-key replaces the signing key and keeps old public keys so older exports still verify. If you configure a timestamp authority in governance/audit-timestamp.json, exports also carry an RFC 3161 timestamp. To ship each closed day to a SIEM, set audit_sink_endpoint in the region policy.
Logs
- Containers log one JSON object per line to stdout (
AISUITE_LOG_JSON=1in the images), ready for Fluent Bit, Loki or CloudWatch. - Windows and other installs also write
logs/aisuite-server-yyyyMMdd.login the data folder: 14 files of up to 50 MB. - API keys, licence keys and other secrets are redacted before anything is written.
AISUITE_LOG_LEVEL=Debug(or Log level on the Server page) for troubleshooting;Warningfor quiet production logs.
Every response carries X-Request-Id (yours, if you sent a well-formed one) and X-AI-Server-Version; behind a gateway, X-AISuite-Backend names the worker that answered. Quote them in support requests.
Questions
Can the dashboard show who asked what? +
No. It shows which key and app made requests to which model and how they performed. Prompt and answer text are never recorded anywhere.
Why does Prometheus show the target as down after a key rotation? +
The scrape authenticates with an API key. Issue the scraper its own admin key and update the Secret or scrape config when you rotate it.