hyperlakeDiscuss a deployment ↗
Answer · Self-hosting LLMs

Is self-hosting an LLM worth it compared with an API?

The break-even logic, a worked example you can repeat in our free calculator, and the reasons that have nothing to do with cost.

Updated
The short answer

Self-hosting an LLM is worth it when steady volume keeps your GPUs busy enough that each token costs less than the API's, after staff, spare GPUs and other fixed costs, or when data, residency or model-version rules rule out an API. Below the break-even volume, a hosted API costs less, so work it out with your own numbers.

An API charges for every token. Your own GPUs cost about the same whether they are busy or idle. So the question is not which is cheaper in general, but whether your volume keeps enough GPUs busy enough.

How do you work out the break-even?

Compare the API's price per token with what a GPU costs per token at the share of its throughput you will actually use.

  1. API side. Monthly tokens times the API's input and output prices gives a straight line: twice the volume, twice the bill.
  2. Owned side. Each GPU costs a fixed amount a month (rented by the hour, or a purchase price spread over its write-off period plus power and hosting). Add spare GPUs and the people and tools needed to run them. GPUs come in whole units, so this cost rises in steps.
  3. Throughput. One GPU serves its measured tokens per second, times your target utilization, times the seconds in a month. Measure it on your own model and prompts; vllm bench serve, for example, reports total token throughput (input plus output) separately from output throughput.
  4. Break-even. Owned capacity can only win if one GPU, at your target utilization, serves more tokens' worth of API spend in a month than it costs. Above the break-even volume it costs less; below it, the API does.

The LLM API vs self-hosted GPU cost calculator does this arithmetic in your browser, with every formula on the page and no built-in prices.

What does a worked example look like?

Here is an internal assistant with long, document-heavy prompts. Every number below is an example chosen to show the arithmetic, not a current price or benchmark. Enter the same values in the calculator to reproduce it.

Calculator input Example value
Model requests per month 1,500,000
Average input / output tokens per request 1,500 / 300
Price per 1M input / output tokens 2 / 8
How the GPUs are paid for Bought outright
Purchase price per GPU; years to write it off 30,000; 3
Power, hosting and networking per GPU per month 500
Hours the GPUs run per month 730
Throughput per GPU, tokens per second 2,500
Target utilization 50%
Spare GPUs; other monthly cost 1; 8,000

What comes out:

Result Working Value
Tokens a month 1.5M × 1,500 in + 1.5M × 300 out 2.7 billion
API cost 2,250M in × 2 + 450M out × 8, per million 8,100 a month (3.00 per 1M tokens)
One GPU a month 30,000 ÷ 36 + 500 1,333.33
Tokens one GPU serves 2,500 × 0.5 × 3,600 × 730 3.285 billion a month
GPUs 1 needed + 1 spare 2
Owned cost 2 × 1,333.33 + 8,000 10,666.67 a month
Verdict The API is cheaper by 2,566.67 a month
Utilization used 20.5% of fleet peak; break-even needs 27.1% Below break-even
Break-even volume (1 spare GPU × 1,333.33 + 8,000 other + 2 working GPUs × 1,333.33) ÷ 0.000003 per token 4.0 billion tokens, about 2.22M requests a month

So at this volume self-hosting is not worth it on cost alone. If agents or new users tripled the volume to 4.5 million requests, the API would cost 24,300 a month and the owned route 13,333.33 (three GPUs plus one spare), and the answer would flip.

Which inputs move the answer most?

In this example, adding 12,000 a month of staff and tooling cost moves the break-even further than doubling the GPU price.

Change one input, keep the rest Break-even volume
None (the example above) about 2.22M requests a month
GPU purchase price doubled to 60,000 about 2.69M requests
Measured throughput halved to 1,250 tokens a second about 2.47M requests
Other monthly cost raised from 8,000 to 20,000 about 4.69M requests

That is why an honest estimate of engineering and on-call time matters as much as a sharp GPU quote. Your own sensitivity may differ; the calculator shows it for your inputs.

What are the non-cost reasons to self-host?

Control, data location and latency can justify self-hosting even when the API is cheaper.

  • Model versions. Hosted models are retired on the provider's schedule. Anthropic, for example, gives at least 60 days' notice before retiring a publicly released model, after which requests to it fail. Weights you hold keep working until you choose to change them.
  • Where data is processed. Providers offer residency options, but within limits. OpenAI lists regional processing only for the United States, Europe and the UAE, with storage-only residency elsewhere, and by default keeps abuse-monitoring logs for up to 30 days unless a customer is approved for zero data retention.
  • Running inside someone else's environment. If your software must run in a client's network, a hosted API may not be allowed there at all.
  • Latency and placement. A model next to your data and users avoids an internet round trip. Measure time to first token on both routes rather than assuming.

When is the API the better choice?

Choose the API when:

  • your volume is below the break-even, or too spiky to keep GPUs busy;
  • you need a model that is only available as a hosted service;
  • nobody on the team can run GPU servers, patch them and be on call;
  • you are still finding out what the product needs, and the volume could change tenfold.

Hyperlake builds for teams on the other side of that line: those that need private control, deploy repeatedly or run sustained agent workloads. Its use cases page explains why per-use and owned costs can cross in either direction.

Frequently asked questions

Is a self-hosted open model as good as a hosted frontier model?

It depends on the task. Hosted and open models differ in quality, context length and features, and the gap varies by use case. Test candidate models on a sample of your real prompts before you compare costs, because a cheaper token is no saving if the answers are worse.

Should I rent GPUs or buy them?

Renting suits uncertain or growing volume, because you can stop paying when you stop using them. Buying spreads the purchase price over the years you expect to use the GPUs, plus power and hosting, and only pays off if they stay busy. The calculator handles both: choose rented or bought in step 3.

Can I mix self-hosting and an API?

Yes. A router can send sensitive or high-volume, routine work to the self-hosted model and the rest to an API. The post on on-premises and hybrid AI explains how hybrid routing works.

Where do I find my token volume?

In your API provider's usage export, in the token counts returned with each response, or in your own application logs. Use a typical month and include retries and background jobs, not only user-facing calls.

Sources

Sources for the facts on this page, last checked October 9, 2026.

  1. Benchmark CLI: vllm bench serve and its throughput metrics (vLLM documentation) checked October 9, 2026
  2. Model deprecations (Claude API documentation) checked October 9, 2026
  3. Data controls in the OpenAI platform (OpenAI API documentation) checked October 9, 2026

Start with a workload. Build the environment around it.

Tell us what you need to deploy, whose environment it must run in, and what it needs to connect to.