Two kinds of vision are available:

  • Understanding an image — describing it, answering questions about it, reading text in it — uses chat with an image and a vision-capable chat model.
  • Measuring an image — boxes around objects, faces, a cut-out — uses the endpoints below. They return structured data instead of text.

Every endpoint takes JSON with a model and an image, given as base64 or a data: URL. They answer 501 no_provider_configured until a vision engine is installed on the server.

Object detection

POST /v1/vision/detections

{ "model": "<detector-model-id>", "image": "data:image/jpeg;base64,…", "confidence_threshold": 0.4, "max_results": 20 }
{ "width": 1280, "height": 720,
  "objects": [ { "label": "person", "confidence": 0.93, "box": [412, 88, 228, 614] } ] }

Each box is [x, y, width, height] in pixels of the original image, from its top-left corner.

Faces

POST /v1/vision/faces finds faces; with "return_embeddings": true it adds a vector per face for matching the same person across images. Face embeddings are biometric data: process them only with a lawful basis and the people's consent, and keep them out of logs.

{ "width": 1280, "height": 720, "faces": [ { "confidence": 0.98, "box": [500, 120, 140, 180], "embedding": [0.012, …] } ] }

Background removal

POST /v1/vision/background returns the subject as a transparent PNG:

{ "image": "data:image/png;base64,…", "media_type": "image/png" }

Questions

Are images stored on the server? +

No. Images are processed in memory and the results returned; audit records hold only the endpoint, model, status and timing.

Can a key be limited to vision only? +

Yes: issue it with --allow-endpoint vision — see API keys.