On-Prem LLM vs API: What It Actually Costs to Run Your Own Model
A practical method for comparing self-hosted LLM inference with model APIs: the real cost drivers on each side, a break-even formula based on GPU throughput and utilization, a worked example, and a checklist for deciding.
Most teams asking whether to run their own model start with the wrong number. They compare an API's price per million tokens with the hourly rate of a GPU and conclude one side is obviously cheaper. The real answer depends on three things the price sheets leave out: how many tokens you actually process each month, how many tokens per second a GPU delivers at the latency you need, and how much of the day those GPUs are busy.
This guide lays out the cost drivers on each side, gives a break-even method you can run with your own numbers, and covers the reasons to self-host (or not) that have nothing to do with cost.
How Much Has LLM API Pricing Changed in 2026?
API prices have fallen sharply. Industry estimates put the price of a given level of model capability down by one to two orders of magnitude since 2023, and small, fast models now cost a fraction of a dollar per million tokens. Frontier models still cost much more, and output tokens are typically priced several times higher than input tokens.
Total spend has gone the other way. Several 2026 analyses estimate that inference now accounts for the majority of enterprise AI compute spend, roughly two-thirds or more, because cheaper tokens led to more usage: longer contexts, agents that call models in loops, and reasoning models that generate long internal chains before answering. Falling unit prices and rising bills are both true at once, and that is the context for any self-hosting decision.
What Drives the Cost of an LLM API?
API cost is simple to measure and harder to predict. The main drivers:
- Per-token pricing. You pay per million input tokens and per million output tokens, at different rates. A retrieval-heavy workload with long prompts and short answers costs very differently from a drafting workload with short prompts and long outputs.
- Reasoning-model token inflation. Reasoning models bill for the tokens they generate while thinking, which can be several times the length of the visible answer. A workload that looked affordable on a standard model can multiply in cost when you switch.
- Context growth. Agents and multi-turn chats resend history on each call. Prompt caching helps where it is offered, but input volume tends to grow faster than teams expect.
- Data and compliance constraints. If documents cannot leave your perimeter, a public API may require a dedicated or regional deployment at higher prices, or may be ruled out entirely. That cost does not show up on the price page.
The upside is that API cost is fully variable. At low or bursty volume you pay only for what you use, with no idle hardware.
What Does It Cost to Self-Host an LLM?
Self-hosted cost is mostly fixed. You pay for capacity whether or not you use it.
- GPU capex or rental. Buying hardware means a large upfront cost amortized over roughly three to five years. Renting from a cloud or GPU provider converts this to an hourly rate, usually with discounts for long commitments.
- Power, cooling and space. For on-prem hardware, a dense multi-GPU server draws several kilowatts, and many existing data center racks were not designed for that density.
- Ops staff. Someone has to run the serving stack, monitor latency, handle failures, patch drivers and manage capacity. Even a fraction of two engineers is a real monthly cost.
- Utilization. This is the variable that decides everything. A GPU that is busy 20% of the time costs five times more per token than one busy 100% of the time.
- Model updates. Open-weight models improve quickly. Each upgrade means evaluation, re-quantization, re-tuning the serving configuration and regression testing.
How to Calculate Self-Hosted Cost per Million Tokens
The method needs four inputs:
- Monthly tokens (T): input plus output tokens per month, measured or forecast.
- Throughput (R): tokens per second one serving node delivers while meeting your latency target. Measure this on your own prompts. Peak benchmark throughput at unlimited latency is the wrong number.
- Utilization (U): average fraction of capacity in use. Traffic has peaks, so plan for 30–60% unless you have batch work to fill the gaps.
- Hourly cost (H): fully loaded cost per node-hour, including amortized hardware or rental, power and hosting.
The GPU cost per million tokens is then:
cost per 1M tokens = H / (R × 3,600 × U) × 1,000,000
Add ops cost spread over monthly volume (monthly ops cost ÷ T × 1,000,000) and compare the total with your blended API price for the same input/output mix. The number of nodes you need is T ÷ (R × 2,628,000 × U), rounded up, where 2,628,000 is the seconds in an average month.
A Worked Break-Even Example
The numbers below are illustrative round figures chosen to show the method. They are not quotes for any vendor, model or GPU.
Assume a workload of 20 billion tokens per month, a blended API price of $3 per million tokens, and one 8-GPU node costing $24 per hour fully loaded (about $17,500 per month). Ops costs $15,000 per month. Planned utilization is 50%.
| Scenario | Throughput per node | Nodes needed | Monthly cost | Cost per 1M tokens |
|---|---|---|---|---|
| API | n/a | n/a | $60,000 | $3.00 |
| Self-hosted, baseline serving | 10,000 tokens/s | 2 | $50,000 | $2.50 |
| Self-hosted, optimized serving | 20,000 tokens/s | 1 | $32,500 | $1.63 |
At baseline throughput, self-hosting barely beats the API, and a modest API price cut would erase the gap. Doubling throughput per node halves the hardware bill, and self-hosting now costs close to half the API price.
Run the same model at 2 billion tokens per month and the picture flips. One node plus ops still costs about $32,500, which works out to over $16 per million tokens against $3 for the API. Below a certain volume, fixed costs dominate and the API wins. Your break-even volume is roughly your monthly fixed cost divided by the API price per million tokens.
Why Throughput Optimization Moves the Break-Even Point Most
In the formula, throughput sits in the denominator next to utilization. Utilization is largely set by your traffic pattern. Throughput is set by engineering, and it is where most of the room is.
- Batching. Continuous batching keeps the GPU full by adding and removing requests mid-generation. The tradeoff is latency, so the right batch size depends on your P99 target.
- Quantization. Running weights at int8 or int4 cuts memory traffic and lets larger batches fit in memory. The question is quality loss, which has to be measured on your tasks. We cover that in detail in how much quality int4 quantization actually costs.
- Kernels and parallelism. Custom CUDA kernels for attention and mixture-of-experts layers, and a tensor parallelism layout matched to the model, remove overhead that generic serving stacks leave in place.
These gains compound. Our DeepSeek R1 FP4 inference world record, 303 tokens per second on DeepSeek R1 in FP4 on a single NVIDIA DGX Blackwell node, came from combining FP4 inference, speculative decoding and a custom TensorRT-LLM build. In the formula above, any throughput gain divides the GPU cost per token by the same factor, which can turn a marginal case into a clear one.
Reasons to Self-Host That Are Not About Cost
For many regulated organizations, cost is the secondary question.
- Data residency and privacy. Prompts, documents and outputs stay inside your perimeter. For banks this is often a hard requirement; our team built on-prem LLM tooling inside a tier-1 banking perimeter at JP Morgan.
- Sovereign AI. Some jurisdictions and public-sector buyers require models and data to stay in-country on infrastructure they control. See our guide to sovereign AI and private LLMs for banks in the UAE.
- Latency. A model next to your systems avoids network round trips and shared-queue variance. Our financial deployments run at P99 latency under 50 ms, which is hard to guarantee over a public API.
- Version control. An API model can change or be retired on the provider's schedule. A self-hosted model changes only when you decide, which matters for audit and reproducibility.
When You Should Not Self-Host
- Your volume is low or very bursty, so utilization will stay low.
- You need the strongest frontier model, and no open-weight model meets your quality bar.
- You have no one to own the serving stack after launch.
- Your requirements are still changing month to month, and committing to hardware would lock in early guesses.
The Hybrid Approach
Many organizations end up with a mix. Sensitive or high-volume, predictable workloads run on a self-hosted model sized for steady load. Low-volume, experimental or frontier-quality tasks go to an API, provided the data allows it. A routing layer sends each request to the right backend, and the self-hosted side can absorb batch jobs overnight to raise utilization. This keeps the fixed cost matched to the base load while the API handles the peaks and the edge cases.
Checklist: Should You Run Your Own LLM?
- Measure monthly input and output tokens separately, including reasoning tokens.
- Get your blended API price for that mix, including any dedicated or regional tier your compliance rules require.
- Benchmark throughput per node on your own prompts at your real latency target.
- Estimate realistic utilization from your traffic curve, and identify batch work that could fill idle hours.
- Price the fully loaded node-hour and the ops time you will actually need.
- Compute both sides with the formula above, then rerun it with API prices 50% lower.
- Confirm an open-weight model meets your quality bar on your own evaluation set.
- List the non-cost requirements: residency, sovereignty, latency, version control.
- Decide which workloads go hybrid.
If the numbers are close, throughput engineering usually decides the outcome. Our LLM engineering team runs this analysis and builds the serving stack. Contact us with your workload and we will help you size it.
Related projects
Related capabilities
- CapabilityLLM EngineeringOn-prem and air-gapped LLM deployment with custom kernels, low-precision inference and speculative decoding, the methods behind our DeepSeek R1 inference world record on NVIDIA Blackwell with Avian.io and NVIDIA.
- CapabilityResearch & DevelopmentApplied AI research for problems without a known solution. We co-author with UC Berkeley, Harvard and UCL, then deploy what works into your stack.
More articles
- EngineeringInside the DeepSeek R1 World Record: 303 Tokens per Second in FP4 on NVIDIA Blackwell
- EngineeringINT4 Quantization for LLMs: How to Measure the Quality Loss Before You Ship
- GuidesSovereign AI in the UAE: How Banks Can Deploy Private LLMs and Keep Data in the Country
- ResearchHyperspectral Heavy Metal Detection: Multiscale Spatial Deep Learning on EnMAP Satellite Imagery