This page gives you a way to calculate the case with your own numbers. It deliberately contains no customer figures or cloud prices: both vary too much, and your own pilot will be more convincing.
The cost model
AI Server's cost is fixed; cloud AI's cost grows with use.
AI Server, per month:
servers × licence + hardware ÷ months of use + power + admin time
- Licence: US$49.99 per server per month (US$499.99 a year) for Commercial; the gateway is included.
- Hardware: often a workstation you already own for a pilot; a GPU workstation or server for a team. Spread it over its useful life (three years is common).
- Power: the machine's average draw × hours × your electricity price.
- Admin time: hours a month of IT time for updates and keys — typically small once it runs.
Cloud AI, per month:
people × requests per person per day × working days × (input tokens × input price + output tokens × output price)
Or, for per-seat AI subscriptions: people × price per seat.
Gather the inputs from a week of real use: how many people, how many requests each, and how long the prompts and answers are. AI Server's usage report gives exactly these numbers during a pilot (requests, input and output tokens per key and model).
Where the case is strongest
- Steady, high volume: document search, transcription, translation and drafting used every day by many people. Cloud cost grows with every request; the server's does not.
- Large documents: long prompts are the expensive part of per-token pricing.
- Many light users: per-seat subscriptions charge for people who use AI occasionally; a server charges for capacity.
- Images and speech: generating images and transcribing hours of audio are priced per item in the cloud.
Where it is weaker
- Very low or very bursty use, where a few cloud requests a week cost less than any hardware.
- Tasks that need the largest frontier models, which do not run on local hardware. AI Server can route those to a cloud provider an operator configures, so you can mix both.
Benefits that are not on the invoice
- Data stays inside. Prompts and documents do not go to a third party, which removes a review step for every new use of AI.
- Predictable budget. A fixed monthly cost instead of a usage bill.
- Control. You decide which models are approved, who has access, and how much each team may use; usage is visible per key.
- No vendor lock-in for your code. Applications use the OpenAI-compatible API, so they can move between AI Server and other providers.
- Works offline. Inference does not depend on an internet connection.
Run a pilot that proves it
- Install the Free edition on an existing machine and try the main use cases.
- Move to a monthly Pro Commercial subscription and give 10–20 people AI Client or a key for their tools.
- After two to four weeks, take the usage report: requests and tokens per person and model.
- Put those numbers into both formulas above, with the hardware you would buy for the full team.
- Add the factors that are not on the invoice, and decide.
The deployment planner sizes servers and seats for a workload and prices the licences.
Questions
How many people can one server support? +
It depends on the model size, the GPU and how bursty the use is. A mid-range GPU serves a team of dozens for typical office chat. See sizing and measure in your pilot.
Do we need to hire AI specialists? +
No. IT staff install and run AI Server; the server downloads and runs the models.