POST /v1/embeddings turns text into vectors for semantic search, retrieval-augmented generation (RAG), clustering and classification. The request and response match OpenAI's.
vectors = client.embeddings.create(model="nomic-embed-text:latest",
input=["How do I reset my password?", "Opening hours are 9 to 5."])
print(len(vectors.data), len(vectors.data[0].embedding))
Models
Embedding models are small and run well on a CPU. Common choices:
| Model | Good for |
|---|---|
enginea/all-minilm/latest | Fast, English, small vectors |
enginea/nomic-embed-text/latest | General English retrieval; supports shortened vectors |
enginea/snowflake-arctic-embed/33m | Compact, retrieval-tuned |
enginea/mxbai-embed-large/latest | Higher quality English retrieval |
enginea/bge-m3/latest | Many languages |
Use the same model for your documents and your queries; vectors from different models cannot be compared. The wire form (for example nomic-embed-text:latest) works too.
Parameters
| Parameter | Notes |
|---|---|
model | Required. |
input | A string or an array of strings. Token-id arrays are not supported. |
encoding_format | float (default) or base64 (little-endian float32), which is smaller on the wire. |
dimensions | Keep the first N values and re-normalise. Use it only with models trained for shortening, such as nomic-embed-text. |
user | Accepted. |
Batch jobs
For thousands of chunks — indexing a document library, for example — submit a background job instead of many requests. The server runs jobs one at a time as background work, so they never slow down people chatting.
curl -X POST http://<server>/v1/batch/embeddings -H "Authorization: Bearer $KEY" \
-H "Content-Type: application/json" -d '{"model":"nomic-embed-text:latest","input":["chunk 1","chunk 2"]}'
# 202 {"id":"…","status":"queued"}
curl http://<server>/v1/batch/embeddings/<id> -H "Authorization: Bearer $KEY"
# status, then the full embeddings response once completed
curl -X DELETE http://<server>/v1/batch/embeddings/<id> -H "Authorization: Bearer $KEY"
- A job holds up to 4,096 inputs by default; results are kept for 24 hours (operators can change both).
- Only the key that submitted a job can read or cancel it.
- Jobs live in memory, so a server restart drops them; resubmit after a restart.
Building RAG on AI Server
- Split documents into passages of a few hundred words.
- Embed the passages once, with a batch job, and store the vectors with their text.
- For each question, embed it, find the closest passages (cosine similarity), and send them to chat as context with an instruction to answer only from them and cite them.
The open-source ai-server-doc-qa sample does this end to end without a vector database.
Questions
Are embeddings counted against quotas? +
Yes. Each request (or batch submission) counts as one request; tokens are recorded for the usage report.
Can I get embeddings from the legacy API? +
The legacy local-AI API's embedding endpoint works until 2026-12-31. Use /v1/embeddings for anything new.