From OpenAI's cloud API

  1. Change the base URL to your server's …/v1 and the key to an AI Server API key.
  2. Change model names to models installed on the server (GET /v1/models). Pick by job: a 7–8B instruct model for chat and tools, an embedding model for search, an image model for pictures.
  3. Re-embed your documents: vectors from different models cannot be compared, so a vector database built with a cloud embedding model needs re-indexing.
  4. Check what you use against the supported parameters. Not available: /v1/responses, /v1/completions, /v1/moderations, /v1/files, Assistants, fine-tuning, n > 1 and logprobs.
  5. Test quality on your own prompts. Smaller local models need clearer instructions and benefit from examples in the prompt.

Keep the cloud as a fallback if you need one: an operator can add a cloud provider on the server, and your code still talks only to AI Server.

From another local model server

Most local model servers offer an OpenAI-compatible endpoint. If your code uses it:

  1. Change the base URL to AI Server and add an API key — AI Server requires keys on the network.
  2. Download the models you need again from AI Server's Models page or with POST /v1/models/pull; AI Server does not read another server's model folder.
  3. Model names may differ: use the ids AI Server lists.

If your code uses another server's own non-OpenAI API, move it to the /v1 endpoints. The comparison pages list differences in features.

From AI Server's legacy local-AI API

Servers upgraded from AI Server 2.0.1 can keep serving that generation's local-AI API (/api/chat, /api/generate, /api/embed, /api/tags…) on port 11434, so old scripts kept working. It ends on 2026-12-31.

Legacy callReplacement
POST /api/chatPOST /v1/chat/completions
POST /api/generatePOST /v1/chat/completions with one user message
POST /api/embed, /api/embeddingsPOST /v1/embeddings
GET /api/tags, /api/modelsGET /v1/models
POST /api/showGET /v1/models/{id} or GET /v1/models/catalog
POST /api/pullPOST /v1/models/pull

The legacy API has no keys, so on Free it serves its own computer only. On the /v1 API every network client needs a key. Watch the dashboard's endpoint table: when /api/* traffic stops, turn the legacy listener off.

Checklist

  • Base URL ends in /v1 and points at AI Server or the gateway.
  • Each application has its own API key.
  • Model names exist on the server; embeddings re-indexed if the model changed.
  • 429 and 503 handled with Retry-After (rate limits).
  • Streaming works through any proxy in between.
  • Unattended jobs marked as background work.