List installed models
GET /v1/models returns the models that are installed and allowed by the operator's model lifecycle policy, in OpenAI's list shape. Ids are canonical runtime/family/variant names.
{ "object": "list", "data": [
{ "id": "enginea/llama3.2/3b", "object": "model", "created": 1759650000, "owned_by": "enginea" },
{ "id": "enginea/nomic-embed-text/latest", "object": "model", "created": 1759650000, "owned_by": "enginea" }
] }
GET /v1/models/{id} returns one model, accepting either the canonical id or the wire id (llama3.2:3b); an unknown id answers 404 model_not_found.
Browse the catalogue
GET /v1/models/catalog lists every model the server knows how to get, installed or not, with what you need to choose one:
| Field | Meaning |
|---|---|
key | The id to use in requests and downloads. |
display_name, description | Human-readable name and summary. |
parameters, quantization | Size class (for example 7B) and compression. |
download_size_bytes, on_disk_size_bytes | Space needed. |
modalities, tags | What it does (chat, embeddings, vision, image generation, speech…). |
context_window_tokens | Longest context it accepts. |
license, homepage_url, source_repository | Where it comes from and its licence terms. |
is_installed, is_active, last_used_utc | Its state on this server. |
GET /v1/models/installed returns just the installed keys. Operators can refresh the catalogue sources with POST /v1/models/catalog/refresh.
Download a model
POST /v1/models/pull downloads a model and loads it, streaming progress as Server-Sent Events.
curl -N http://<server>/v1/models/pull -H "Authorization: Bearer $KEY" \
-H "Content-Type: application/json" -d '{"model_key":"enginea/qwen2.5/7b"}'
event: started
data: {"model_key":"enginea/qwen2.5/7b"}
event: progress
data: {"phase":"downloading","completed":1048576000,"total":4683073536,"bps":52428800,"eta_seconds":69}
event: done
data: {"model_key":"enginea/qwen2.5/7b","duration_ms":91234}
phase is a free-text label from the engine (it varies by engine and step); use completed and total for progress bars. A failure ends with event: error. model_key must be the canonical form. Downloads of the same model by several clients are shared, not repeated.
- The key needs the
pullendpoint permission if it has an endpoint allowlist. - Downloads are allowed in the default federated governance mode. In curated mode only operators install models, and the endpoint answers 403
governance_denied; in open mode clients download models on their own devices. - Downloading also makes the model the server's active model. On a busy shared server, install models in the app's Models page instead.
Removing models is an operator task in the app's Models page; the API's delete route refuses client requests.
Provenance
GET /v1/server/models/{id}/provenance returns the model card an operator registered — publisher, source, licence, intended use and known limitations — or 404 when none is registered. Review it before approving a model for your organisation.
Questions
Why is a model I just downloaded missing from /v1/models? +
A model blocked by the lifecycle policy is hidden. Otherwise, check that the download ended with event: done.
Which ids should my code use? +
Either works. Canonical ids name the engine explicitly, which is clearer in configuration files; wire ids match what other OpenAI-compatible servers use.