Every AI Server node — a single server, a worker or a gateway — monitors itself. There is no agent or exporter to install. None of these views ever contain prompt or answer text.

WhatWhereWho can read it
Live dashboardGET /dashboard in a browserThe page is public and empty; its data needs an admin key off this computer
Prometheus metricsGET /v1/server/metricsAdmin key
Usage reportGET /v1/server/usage, the app's Usage pageAdmin key
Live request activityThe app's Active requests pageThis computer
Signed audit exportGET /v1/server/audit/export, aisuite-server-cli export-auditAdmin key
Gateway pool statusGET /v1/gateway/workersAdmin key
Health probesGET /livez, GET /readyzAnyone
Logsstdout, and logs/aisuite-server-*.log outside containersServer administrators

An admin key is an API key issued with administration rights (aisuite-server-cli keys add --label monitoring --admin, or Allow server administration on the API keys page). Give monitoring tools their own key: rotating a client key should never break monitoring. See API keys.

Dashboard

Open http://<server>:<port>/dashboard. It shows whether the server is ready or draining, request rate, latency percentiles, the busiest endpoints, usage by app and model, and on a gateway every worker with its health, load against its cap, latency, canary flag and installed models. Paste an admin key into the page when you open it from another computer. Set AISUITE_DASHBOARD=0 to remove it.

Prometheus

scrape_configs:
  - job_name: aisuite-gateway
    metrics_path: /v1/server/metrics
    authorization: { type: Bearer, credentials: <admin-key> }
    static_configs: [ { targets: ["gateway.example.internal:8080"] } ]
SeriesTypeMeaning
aisuite_requests_totalcounterRequests served.
aisuite_errors_totalcounterResponses with status 500 or above. A 4xx (wrong key, quota) is a client problem, not an outage.
aisuite_requests_by_endpoint_totalcounterRequests per endpoint.
aisuite_requests_by_status_totalcounterRequests per HTTP status.
aisuite_request_duration_secondssummaryLatency per endpoint (p50, p95, p99).
aisuite_gateway_pool_workers, …_pool_workers_healthygaugeWorkers in the pool, and how many are healthy.
aisuite_gateway_pool_in_flight, …_pool_in_flight_peakgaugeRequests in progress across the pool.
aisuite_gateway_worker_healthy, …_breaker_open, …_consecutive_failuresgaugePer-worker health and circuit breaker state.
aisuite_gateway_worker_in_flight, …_max_in_flightgaugePer-worker load and its cap — the right autoscaling signal.
aisuite_gateway_worker_ewma_latency_secondsgaugeSmoothed per-worker latency.
aisuite_gateway_worker_requests_totalcounterRequests sent to each worker.
aisuite_gateway_pool_info, aisuite_gateway_worker_infogaugeLabels: routing strategy, worker id, canary.

Starter alerts worth having:

  • No healthy workers: aisuite_gateway_pool_workers_healthy == 0 for 2 minutes. This is what users feel.
  • Target down: up == 0 for 5 minutes — also fires when the scrape key was revoked.
  • Error rate: aisuite_errors_total over aisuite_requests_total above 5% for 10 minutes.

Do not alert on /readyz returning 503: a draining pod does that on every rollout.

Usage reports

GET /v1/server/usage?from=2026-10-01&to=2026-10-31&by=key returns request counts, failures, latency, input and output tokens and an estimated cost, grouped by key, model, app or (on a gateway) worker. The default is the last 30 days by key. Costs use the per-model prices in governance/model-pricing.json, so they can feed a charge-back to teams.

Audit log

The server writes one content-free line per request — time, endpoint, model, key id, app, status, latency and token counts — to daily files in the audit folder, kept for the retention you set in the region policy.

For an auditor, export a signed bundle and let them check it offline:

aisuite-server-cli export-audit --from 2026-10-01 --to 2026-10-31 --out audit-october.json
aisuite-server-cli verify-audit --in audit-october.json      # exit 0 = PASS, 2 = FAIL

Bundles are signed with ES256 using a key that stays on the server; the public key is in the bundle and at GET /v1/server/audit/public-key. rotate-audit-key replaces the signing key and keeps old public keys so older exports still verify. If you configure a timestamp authority in governance/audit-timestamp.json, exports also carry an RFC 3161 timestamp. To ship each closed day to a SIEM, set audit_sink_endpoint in the region policy.

Logs

  • Containers log one JSON object per line to stdout (AISUITE_LOG_JSON=1 in the images), ready for Fluent Bit, Loki or CloudWatch.
  • Windows and other installs also write logs/aisuite-server-yyyyMMdd.log in the data folder: 14 files of up to 50 MB.
  • API keys, licence keys and other secrets are redacted before anything is written.
  • AISUITE_LOG_LEVEL=Debug (or Log level on the Server page) for troubleshooting; Warning for quiet production logs.

Every response carries X-Request-Id (yours, if you sent a well-formed one) and X-AI-Server-Version; behind a gateway, X-AISuite-Backend names the worker that answered. Quote them in support requests.

Questions

Can the dashboard show who asked what? +

No. It shows which key and app made requests to which model and how they performed. Prompt and answer text are never recorded anywhere.

Why does Prometheus show the target as down after a key rotation? +

The scrape authenticates with an API key. Issue the scraper its own admin key and update the Secret or scrape config when you rotate it.