Is self-hosting an LLM worth it compared with an API?
The break-even logic, a worked example you can repeat in our free calculator, and the reasons that have nothing to do with cost.
Self-hosting an LLM is worth it when steady volume keeps your GPUs busy enough that each token costs less than the API's, after staff, spare GPUs and other fixed costs, or when data, residency or model-version rules rule out an API. Below the break-even volume, a hosted API costs less, so work it out with your own numbers.
An API charges for every token. Your own GPUs cost about the same whether they are busy or idle. So the question is not which is cheaper in general, but whether your volume keeps enough GPUs busy enough.
How do you work out the break-even?
Compare the API's price per token with what a GPU costs per token at the share of its throughput you will actually use.
- API side. Monthly tokens times the API's input and output prices gives a straight line: twice the volume, twice the bill.
- Owned side. Each GPU costs a fixed amount a month (rented by the hour, or a purchase price spread over its write-off period plus power and hosting). Add spare GPUs and the people and tools needed to run them. GPUs come in whole units, so this cost rises in steps.
- Throughput. One GPU serves its measured tokens per second, times your target utilization, times the seconds in a month. Measure it on your own model and prompts;
vllm bench serve, for example, reports total token throughput (input plus output) separately from output throughput. - Break-even. Owned capacity can only win if one GPU, at your target utilization, serves more tokens' worth of API spend in a month than it costs. Above the break-even volume it costs less; below it, the API does.
The LLM API vs self-hosted GPU cost calculator does this arithmetic in your browser, with every formula on the page and no built-in prices.
What does a worked example look like?
Here is an internal assistant with long, document-heavy prompts. Every number below is an example chosen to show the arithmetic, not a current price or benchmark. Enter the same values in the calculator to reproduce it.
| Calculator input | Example value |
|---|---|
| Model requests per month | 1,500,000 |
| Average input / output tokens per request | 1,500 / 300 |
| Price per 1M input / output tokens | 2 / 8 |
| How the GPUs are paid for | Bought outright |
| Purchase price per GPU; years to write it off | 30,000; 3 |
| Power, hosting and networking per GPU per month | 500 |
| Hours the GPUs run per month | 730 |
| Throughput per GPU, tokens per second | 2,500 |
| Target utilization | 50% |
| Spare GPUs; other monthly cost | 1; 8,000 |
What comes out:
| Result | Working | Value |
|---|---|---|
| Tokens a month | 1.5M × 1,500 in + 1.5M × 300 out | 2.7 billion |
| API cost | 2,250M in × 2 + 450M out × 8, per million | 8,100 a month (3.00 per 1M tokens) |
| One GPU a month | 30,000 ÷ 36 + 500 | 1,333.33 |
| Tokens one GPU serves | 2,500 × 0.5 × 3,600 × 730 | 3.285 billion a month |
| GPUs | 1 needed + 1 spare | 2 |
| Owned cost | 2 × 1,333.33 + 8,000 | 10,666.67 a month |
| Verdict | The API is cheaper by 2,566.67 a month | |
| Utilization | used 20.5% of fleet peak; break-even needs 27.1% | Below break-even |
| Break-even volume | (1 spare GPU × 1,333.33 + 8,000 other + 2 working GPUs × 1,333.33) ÷ 0.000003 per token | 4.0 billion tokens, about 2.22M requests a month |
So at this volume self-hosting is not worth it on cost alone. If agents or new users tripled the volume to 4.5 million requests, the API would cost 24,300 a month and the owned route 13,333.33 (three GPUs plus one spare), and the answer would flip.
Which inputs move the answer most?
In this example, adding 12,000 a month of staff and tooling cost moves the break-even further than doubling the GPU price.
| Change one input, keep the rest | Break-even volume |
|---|---|
| None (the example above) | about 2.22M requests a month |
| GPU purchase price doubled to 60,000 | about 2.69M requests |
| Measured throughput halved to 1,250 tokens a second | about 2.47M requests |
| Other monthly cost raised from 8,000 to 20,000 | about 4.69M requests |
That is why an honest estimate of engineering and on-call time matters as much as a sharp GPU quote. Your own sensitivity may differ; the calculator shows it for your inputs.
What are the non-cost reasons to self-host?
Control, data location and latency can justify self-hosting even when the API is cheaper.
- Model versions. Hosted models are retired on the provider's schedule. Anthropic, for example, gives at least 60 days' notice before retiring a publicly released model, after which requests to it fail. Weights you hold keep working until you choose to change them.
- Where data is processed. Providers offer residency options, but within limits. OpenAI lists regional processing only for the United States, Europe and the UAE, with storage-only residency elsewhere, and by default keeps abuse-monitoring logs for up to 30 days unless a customer is approved for zero data retention.
- Running inside someone else's environment. If your software must run in a client's network, a hosted API may not be allowed there at all.
- Latency and placement. A model next to your data and users avoids an internet round trip. Measure time to first token on both routes rather than assuming.
When is the API the better choice?
Choose the API when:
- your volume is below the break-even, or too spiky to keep GPUs busy;
- you need a model that is only available as a hosted service;
- nobody on the team can run GPU servers, patch them and be on call;
- you are still finding out what the product needs, and the volume could change tenfold.
Hyperlake builds for teams on the other side of that line: those that need private control, deploy repeatedly or run sustained agent workloads. Its use cases page explains why per-use and owned costs can cross in either direction.
Frequently asked questions
Is a self-hosted open model as good as a hosted frontier model?
It depends on the task. Hosted and open models differ in quality, context length and features, and the gap varies by use case. Test candidate models on a sample of your real prompts before you compare costs, because a cheaper token is no saving if the answers are worse.
Should I rent GPUs or buy them?
Renting suits uncertain or growing volume, because you can stop paying when you stop using them. Buying spreads the purchase price over the years you expect to use the GPUs, plus power and hosting, and only pays off if they stay busy. The calculator handles both: choose rented or bought in step 3.
Can I mix self-hosting and an API?
Yes. A router can send sensitive or high-volume, routine work to the self-hosted model and the rest to an API. The post on on-premises and hybrid AI explains how hybrid routing works.
Where do I find my token volume?
In your API provider's usage export, in the token counts returned with each response, or in your own application logs. Use a typical month and include retries and background jobs, not only user-facing calls.
Sources
Sources for the facts on this page, last checked October 9, 2026.
- Benchmark CLI: vllm bench serve and its throughput metrics (vLLM documentation) checked October 9, 2026
- Model deprecations (Claude API documentation) checked October 9, 2026
- Data controls in the OpenAI platform (OpenAI API documentation) checked October 9, 2026
Related guides
- LLM API vs self-hosted GPU cost calculatorFree calculator for the break-even between a pay-per-token LLM API and your own GPUs. Enter your own prices and volume. Formulas shown; nothing is sent.
- How many GPUs do I need to serve an LLM?GPU count starts with memory: weights by precision, plus KV cache for every token in flight, plus overhead. Worked examples for an 8B and a 70B model.
- How do you deploy an LLM on premises?Deploy an LLM on premises in seven steps: model and licence, GPU sizing, serving engine, gateway, security, monitoring and offline updates.
- The economics of large language models
Start with a workload. Build the environment around it.
Tell us what you need to deploy, whose environment it must run in, and what it needs to connect to.