Private AI guide

What does it cost to self-host an AI model?

The question has no single answer, but it has a reliable method. Here is how to get a number you can defend, using your own figures rather than ours.

Why nobody can quote you a figure

Self-hosting cost is driven by four things that differ wildly between businesses: how many people use it concurrently, how large a model the work needs, how much documentation gets indexed, and whether you buy hardware or rent a tenant.

A guide that prints "$X per month" has picked a scenario and hoped it resembles yours. Treat any such figure as marketing.

The four cost components

  • Compute. Either capital — a GPU server you own and depreciate — or operating, as a private cloud tenant. Capital is higher up front and near-zero per use; the tenant is the reverse.
  • Deployment. One-off. Assessment, infrastructure, knowledge ingestion, access control, pilot. Driven mostly by how organised your documents already are.
  • Maintenance. Ongoing. Updates, model upgrades, monitoring, adding knowledge. Either an arrangement with us or time from your own team, which is a real cost even when it is not an invoice.
  • Your owner's time. Always omitted from these calculations and never actually zero.

The calculation, step by step

  1. Current state. Seats × monthly price × 12. Then the same at your expected headcount in 24 months. That second number is usually the one that changes minds.
  2. Self-hosted year one. Hardware or tenant, plus deployment, plus maintenance.
  3. Self-hosted years two and three. Maintenance only, plus any capacity growth.
  4. Three-year totals, side by side.
  5. Add the unlocked work. What can you do once confidential documents are in scope that you cannot do today? For regulated businesses this term frequently dominates everything above it.

Worth saying: Step five is the one most businesses leave out, and for a law firm or a clinic it is often the whole case. The alternative to private AI there is not cheaper AI; it is no AI on the work that matters.

Where the crossover usually sits

Directionally, and without pretending to a precision we do not have: small teams with no data constraint rarely justify self-hosting on cost alone. Larger teams, or teams growing quickly, reach a point where fixed infrastructure is plainly cheaper than a per-head licence.

The crossover arrives much earlier when you count work you currently cannot do — which is why regulated businesses often find self-hosting justified at a headcount where an unregulated business would not.

We will run this with your actual numbers as part of an AI cost analysis, and the answer is sometimes that you should keep your subscriptions.

Questions people actually ask

Before you call

Is it cheaper than ChatGPT Enterprise?

At small scale, usually not. As headcount grows, usually yes. The crossover depends on your numbers, and counting the work you currently cannot do at all moves it earlier.

Can we use hardware we already own?

Often. Plenty of useful deployments run on existing workstation-class machines, and starting there lets real usage tell you what to buy.

What is the ongoing cost if we self-manage?

Infrastructure plus your own team's time. The time is the part that gets underestimated — somebody has to handle updates, monitoring and new knowledge.

Still have the question?

Ask us directly. We answer these on calls all day.

Certified across the platforms we build on

AWSGoogle CloudMicrosoft AzureAnthropicOpenAI