POST /v1/embeddings turns text into vectors for semantic search, retrieval-augmented generation (RAG), clustering and classification. The request and response match OpenAI's.

vectors = client.embeddings.create(model="nomic-embed-text:latest",
    input=["How do I reset my password?", "Opening hours are 9 to 5."])
print(len(vectors.data), len(vectors.data[0].embedding))

Models

Embedding models are small and run well on a CPU. Common choices:

ModelGood for
enginea/all-minilm/latestFast, English, small vectors
enginea/nomic-embed-text/latestGeneral English retrieval; supports shortened vectors
enginea/snowflake-arctic-embed/33mCompact, retrieval-tuned
enginea/mxbai-embed-large/latestHigher quality English retrieval
enginea/bge-m3/latestMany languages

Use the same model for your documents and your queries; vectors from different models cannot be compared. The wire form (for example nomic-embed-text:latest) works too.

Parameters

ParameterNotes
modelRequired.
inputA string or an array of strings. Token-id arrays are not supported.
encoding_formatfloat (default) or base64 (little-endian float32), which is smaller on the wire.
dimensionsKeep the first N values and re-normalise. Use it only with models trained for shortening, such as nomic-embed-text.
userAccepted.

Batch jobs

For thousands of chunks — indexing a document library, for example — submit a background job instead of many requests. The server runs jobs one at a time as background work, so they never slow down people chatting.

curl -X POST http://<server>/v1/batch/embeddings -H "Authorization: Bearer $KEY" \
  -H "Content-Type: application/json" -d '{"model":"nomic-embed-text:latest","input":["chunk 1","chunk 2"]}'
# 202 {"id":"…","status":"queued"}
curl http://<server>/v1/batch/embeddings/<id> -H "Authorization: Bearer $KEY"
# status, then the full embeddings response once completed
curl -X DELETE http://<server>/v1/batch/embeddings/<id> -H "Authorization: Bearer $KEY"
  • A job holds up to 4,096 inputs by default; results are kept for 24 hours (operators can change both).
  • Only the key that submitted a job can read or cancel it.
  • Jobs live in memory, so a server restart drops them; resubmit after a restart.

Building RAG on AI Server

  1. Split documents into passages of a few hundred words.
  2. Embed the passages once, with a batch job, and store the vectors with their text.
  3. For each question, embed it, find the closest passages (cosine similarity), and send them to chat as context with an instruction to answer only from them and cite them.

The open-source ai-server-doc-qa sample does this end to end without a vector database.

Questions

Are embeddings counted against quotas? +

Yes. Each request (or batch submission) counts as one request; tokens are recorded for the usage report.

Can I get embeddings from the legacy API? +

The legacy local-AI API's embedding endpoint works until 2026-12-31. Use /v1/embeddings for anything new.