← All posts
31 July 2026

Running AI locally: when a self-hosted model beats an API

Most companies asking to run AI on their own hardware want data residency, which is a contract and a region setting. Here is when the hardware is the answer.

Someone in the business has asked whether the AI work can run on your own server. Usually the reason given is data, sometimes it is the invoice from a model provider, and occasionally it is a client contract that says nothing sensitive may leave your infrastructure.

It is a fair question and the answer you will find online is close to useless. Every guide answers it in tokens per day, a unit no business measures itself in, and the thresholds they quote disagree with each other by a factor of a thousand. Meanwhile the two costs that actually decide it, the electricity and the person who has to look after the thing, appear in almost none of them.

So here is the same question answered in units you can check against your own operation, and an honest account of which of the three things people call “running AI locally” you probably need.

The short answer

Most companies asking to self-host want data residency, and residency is a contract and a region setting rather than a server room. GDPR does not require your data to sit on your own hardware, or even in the EU: the EU-US Data Privacy Framework has held an adequacy decision since 10 July 2023, and every major provider now sells European or regional processing as a priced option. Self-hosting genuinely wins in two cases and no others: when the data must never leave your network at all, usually because a client contract or a sector regulator says so, or when one narrow, repetitive, high-volume job runs continuously enough to amortise fixed hardware. The cost case almost never survives contact with the arithmetic, because the workloads where an API bill hurts are exactly the workloads a small local model cannot do.

Three different things get called “running AI locally”

The conversation goes wrong in the first minute because the phrase covers three arrangements with nothing in common.

A model on one machine. Someone runs a model on a workstation or a laptop. No data leaves the building. It serves one person at a time, it is not backed up, nobody is on call for it, and it is a fine way to evaluate a model. It is not a system anyone should depend on.

A model on your own server, on your own network. This is what people usually mean by self-hosting. Your hardware, your building or your rack, your responsibility. It is the only arrangement that satisfies a genuine “nothing leaves the network” requirement, and it is a real infrastructure project with a real owner.

A model in your own cloud tenancy, or on a regional endpoint. The weights run on somebody else’s hardware in a region you choose, under a processor contract you signed. Nothing is installed, nothing is maintained, and for most companies with a residency requirement this is the answer. It is also the option most of the guides skip entirely, because it does not fit the self-hosted versus API framing they started with.

Deciding between these three is most of the work. Confusing them is why so many of these projects begin with a hardware quote and end with nothing running.

The cost case, in units you can check

Every published break-even is stated in tokens per day. Nobody knows how many tokens per day they run, and the figures on offer range from two million to a hundred and thirty million, which tells you the number is doing no work.

Take a workload a business recognises instead. Anthropic publishes a worked example on its own pricing page: 10,000 support-ticket conversations, averaging around 3,700 tokens each, on Claude Haiku 4.5 at $1 per million input tokens and $5 per million output tokens, costs about $37. Ten thousand conversations every month for a year is therefore about $444. Figures read from the live page on 31 July 2026.

Now the other side of the ledger, which almost nobody prices.

Electricity. NVIDIA’s RTX PRO 6000 Blackwell workstation card carries 96 GB of GDDR7 with ECC and a maximum power draw of 600 W. Eurostat puts non-household electricity at an EU average of €0.1837 per kWh in the second half of 2025, for medium consumers using 500 to 2,000 MWh a year, including all non-recoverable taxes. At the card’s rated maximum:

Running patternkWh per yearCost at EU averageCost in Germany (€0.2264)Cost in Finland (€0.0748)
Office hours (2,080 h)1,248€229€283€93
Continuous (8,760 h)5,256€965€1,190€393

Those are ceilings, since a card serving intermittent requests draws well under its maximum. They are also for one card, before you have bought it, bought the machine around it, or paid anyone.

Attention. Eurostat puts average hourly labour costs at €34.9 in the EU and €38.2 in the euro area for 2025, ranging from €12.0 in Bulgaria to €56.8 in Luxembourg. Four hours a month of somebody competent, which is optimistic for a service that needs patching, monitoring and re-benchmarking every time you change model, is €1,675 a year at the EU average.

Put those together and the shape is clear without needing an exchange rate. Before the hardware is paid for, running your own box costs somewhere between two and three thousand euro a year in power and attention. At $37 per 10,000 conversations you would need on the order of half a million conversations a year, roughly 40,000 a month, before the API bill alone matched just the attention line.

The trap inside the cost case

There is an obvious objection: those numbers use a cheap small model, and cheap small models are not what makes an API bill hurt. Correct, and that is the trap.

API spend becomes painful when you are running a frontier-class model over a lot of tokens. OpenAI’s published rates run from $0.25 per million input tokens on gpt-5-mini to $5.00 on gpt-5.6-sol, with output tokens between $2.00 and $30.00. That is a twenty-fold spread between the small model and the flagship, and the bills that trigger a self-hosting conversation are almost always at the flagship end.

But a model you can run on one affordable card is at the other end of that spread in capability too. Matching frontier reasoning locally means weights in the seventy billion to four hundred billion parameter range, which means multiple high-VRAM cards, which means a five-figure or six-figure capital purchase and a person who understands inference serving. The workloads that make the API expensive are the workloads a small local model cannot do, and the workloads a small local model can do are the ones the API charges almost nothing for.

That is the sentence missing from every comparison table in this category. The cost case for self-hosting is not wrong in principle. It is that the volume needed to reach it and the capability needed to serve it pull in opposite directions, and only large, narrow, continuous workloads sit where both conditions hold at once.

What “we need it on our own servers for GDPR” usually means

This is the most common reason given and it is usually a misreading.

GDPR does not require personal data to be processed on hardware you own, and it does not require it to stay inside the EU. Transfers to a country covered by an adequacy decision are permitted, and the European Commission’s own guidance states that under one, personal data “can flow from the EU (and Norway, Liechtenstein and Iceland) to that third country without any further safeguard being necessary”, with such transfers “assimilated to intra-EU transmissions of data”. The EU-US Data Privacy Framework adequacy decision was adopted on 10 July 2023 and remains in force, with a periodic review completed on 9 October 2024. Read on the live page on 31 July 2026.

If you want the processing to happen in Europe regardless, and plenty of buyers reasonably do, that is a purchasing decision rather than a building project. Microsoft’s EU Data Boundary commits to storing and processing customer data and personal data within the EU and EFTA for Azure, Microsoft 365, Dynamics 365 and Power Platform, and the boundary countries include Latvia. Model providers price regional processing directly: Anthropic’s pricing page lists a 1.1x multiplier for restricting inference to a single geography on its newer models, and a 10% premium for regional endpoints on partner clouds.

Two honest caveats, because this is where the category oversells.

First, Microsoft’s own page states the commitments are “subject to limited circumstances where Customer Data, personal data, and Professional Services Data will continue to be transferred outside the EU Data Boundary”, and documents them. A boundary is a strong contractual commitment, not a physical impossibility. Second, an adequacy decision is a political instrument, and this one has a history: the Commission’s page on EU-US transfers cites the Court of Justice judgment in Case C-362/14 (Schrems, 2015) and the Schrems II decision of July 2020, and describes the current framework’s US safeguards as introduced to address the concerns that second ruling raised. Building on an adequacy decision is reasonable. Assuming it is permanent is not.

Which leaves the requirement that genuinely does need your own hardware: data that must never leave your network at all. That usually comes from a client contract, a sector regulator, or a public-sector tender, not from GDPR in the abstract. If you have it in writing, self-hosting is not a preference, it is the specification. If nobody can produce the sentence that requires it, you are about to buy hardware to solve a procurement problem.

What you actually need to run something useful

VRAM is the constraint and everything else follows from it. Weights at four-bit quantisation need roughly half a byte per parameter, a little more in practice, so Qwen3-32B, at 32.8 billion parameters, needs around 20 GB loaded before a single conversation has been added to context.

That has three consequences people discover late. A model that fits on one 24 GB card fits with very little headroom, and context is not free: each concurrent conversation holds its own working memory on the card, so five people using it at once is a different hardware question from one person testing it. Quantisation is a quality decision as well as a memory one, and the quantised model you benchmarked is not the model whose evaluation scores you read. And a card is not a system: you still need redundancy for anything the business depends on, which means the honest hardware answer is two of whatever you just specified.

The licence trap

Self-hosting is often chosen for control, so it is worth reading what the weights actually permit.

The Llama 3.3 Community License (version release date 6 December 2024) requires anyone whose products exceeded 700 million monthly active users on the release date to “request a license from Meta, which Meta may grant to you in its sole discretion”, and separately requires licensees to “prominently display ‘Built with Llama’ on a related website, user interface, blogpost, about page, or product documentation”. The user threshold will never bind a mid-sized company. The attribution requirement binds everyone, and a licence granted at the licensor’s sole discretion is precisely the kind of dependency self-hosting was supposed to remove.

Other strong open-weight models carry ordinary permissive terms. Qwen3-32B is released under apache-2.0. The point is not that one licence is bad and another good, it is that “open source” is doing a lot of unearned work in this category, and if control is your reason for being here then the licence is part of the specification rather than a formality.

Which jobs a small local model can actually do

We sort AI work by job type rather than by department, because the department view groups things that behave nothing alike. Mapped onto that, a small self-hosted model is not uniformly weaker. It is strong at some jobs and unusable at others.

JobSmall local modelWhy
Classification and routingGood fitSorting into categories you defined, and you can measure accuracy on work already done before committing
Structured extractionWorkableFine on documents with a stable shape, with a check set built from records you already hold correct answers for
Retrieval and groundingWorkableAnswer quality is dominated by the retrieval layer, not the model, so a smaller model costs less here than people expect
MonitoringBest fitHigh volume, repetitive, continuous. The only shape where fixed hardware genuinely beats per-token pricing
Drafting and generationPoor fitThe gap against frontier models is widest here, and the review loop eats the saving
Agents that take actionsAvoidErrors compound across steps, and multi-step reasoning is where a small model is weakest

Read that table as a filter. If the job you want to bring in-house is monitoring or classification, self-hosting is a genuine option and the economics can work. If it is drafting or an agent that acts on systems, moving it to a small local model is a downgrade dressed as a cost saving.

What breaks after the demo works

The demo always works. That is the least informative moment in the project.

What costs time afterwards is rarely the model. It is that a monitor stops working and continues to look healthy. We hit exactly that in our own systems: a brand-citation check on this site reported zero results for weeks, and the cause was a string-matching bug comparing a spaced brand name against a domain, not an absence of citations. Nothing alerted, because the job ran successfully every day and returned a number. It was caught by a person reading the output by hand.

That failure mode is the argument for the ops line in the arithmetic above. A self-hosted model needs someone who notices when quality drifts after a quantisation change, when a driver update degrades throughput, when the card is thermally throttling, and when the thing has been answering confidently from a stale index for a fortnight. None of that is difficult. All of it needs an owner with the time to look, and a system nobody checks is not cheaper, it is just unmeasured.

So should you self-host?

Three cases, and most companies are in the first.

Use a hosted API. The default, and the right answer for the large majority of workloads at the large majority of companies. Fastest to a working system, no capital cost, and you can change model when a better one appears.

Use a regional endpoint with a processor contract. The answer when the requirement is that processing happens in Europe or in a named jurisdiction. It is a configuration setting and a contract clause, it costs a modest premium, and it delivers what most people were asking for when they said they wanted to self-host.

Self-host. The answer when the data genuinely may not leave your network and you have the clause in writing, or when one narrow, high-volume, continuous job runs at a scale where fixed hardware wins, and you have somebody whose job includes owning it. Both conditions, not either.

If you are not sure which case you are in, the fastest way to find out is to write down the specific sentence, from a contract or a regulation, that self-hosting exists to satisfy. If you can produce it, you have a specification. If you cannot, you have a preference, and the arithmetic above is what it will cost you.

We build both shapes, and the cost side of the decision is worked through in more detail in what an AI automation actually costs to run. If you want a second opinion on which case you are in before anyone quotes you for hardware, that is a conversation worth having first: AI and automation.

Tell us what’s broken.
We’ll tell you the truth.

Book a free call →
Reply within one business day · EN / LV