This page gives you a way to calculate the case with your own numbers. It deliberately contains no customer figures or cloud prices: both vary too much, and your own pilot will be more convincing.

The cost model

AI Server's cost is fixed; cloud AI's cost grows with use.

AI Server, per month:

servers × licence  +  hardware ÷ months of use  +  power  +  admin time
  • Licence: US$49.99 per server per month (US$499.99 a year) for Commercial; the gateway is included.
  • Hardware: often a workstation you already own for a pilot; a GPU workstation or server for a team. Spread it over its useful life (three years is common).
  • Power: the machine's average draw × hours × your electricity price.
  • Admin time: hours a month of IT time for updates and keys — typically small once it runs.

Cloud AI, per month:

people × requests per person per day × working days × (input tokens × input price + output tokens × output price)

Or, for per-seat AI subscriptions: people × price per seat.

Gather the inputs from a week of real use: how many people, how many requests each, and how long the prompts and answers are. AI Server's usage report gives exactly these numbers during a pilot (requests, input and output tokens per key and model).

Where the case is strongest

  • Steady, high volume: document search, transcription, translation and drafting used every day by many people. Cloud cost grows with every request; the server's does not.
  • Large documents: long prompts are the expensive part of per-token pricing.
  • Many light users: per-seat subscriptions charge for people who use AI occasionally; a server charges for capacity.
  • Images and speech: generating images and transcribing hours of audio are priced per item in the cloud.

Where it is weaker

  • Very low or very bursty use, where a few cloud requests a week cost less than any hardware.
  • Tasks that need the largest frontier models, which do not run on local hardware. AI Server can route those to a cloud provider an operator configures, so you can mix both.

Benefits that are not on the invoice

  • Data stays inside. Prompts and documents do not go to a third party, which removes a review step for every new use of AI.
  • Predictable budget. A fixed monthly cost instead of a usage bill.
  • Control. You decide which models are approved, who has access, and how much each team may use; usage is visible per key.
  • No vendor lock-in for your code. Applications use the OpenAI-compatible API, so they can move between AI Server and other providers.
  • Works offline. Inference does not depend on an internet connection.

Run a pilot that proves it

  1. Install the Free edition on an existing machine and try the main use cases.
  2. Move to a monthly Pro Commercial subscription and give 10–20 people AI Client or a key for their tools.
  3. After two to four weeks, take the usage report: requests and tokens per person and model.
  4. Put those numbers into both formulas above, with the hardware you would buy for the full team.
  5. Add the factors that are not on the invoice, and decide.

The deployment planner sizes servers and seats for a workload and prices the licences.

Questions

How many people can one server support? +

It depends on the model size, the GPU and how bursty the use is. A mid-range GPU serves a team of dozens for typical office chat. See sizing and measure in your pilot.

Do we need to hire AI specialists? +

No. IT staff install and run AI Server; the server downloads and runs the models.