hyperlakeDiscuss a deployment ↗
Free tool · AI inference costs

LLM API vs self-hosted GPU cost calculator

Enter your own request volume, API prices and GPU costs to see what each route costs a month, where they break even and how busy your GPUs must be. It runs in your browser and sends nothing.

Updated
The short answer

This free calculator compares paying per token for a hosted LLM API with running your own GPUs. You enter your request volume, API prices, GPU costs and measured throughput. It returns each route's monthly cost, the break-even volume and the GPU utilization needed to beat the API. It has no built-in prices, and nothing you enter leaves your browser.

Hosted LLM APIs charge for each token you send and receive. Your own GPUs cost about the same whether they are busy or idle. Which route is cheaper depends on how much work you can keep the GPUs doing, and you can work that out with your own numbers.

This calculator does it. You enter your volume, the API's prices, what the GPUs cost and how many tokens a GPU serves per second. It returns the monthly cost of each route, the volume at which they break even, and the utilization your GPUs need to reach. Nothing is pre-filled with a price, so nothing can be out of date, and every formula is written out below.

Run the calculator

1. Your workload

Count every call to the model. An agent that calls the model ten times for one task makes ten requests.

Everything sent in: instructions, history, retrieved text and tool results.

What the model writes back, including any reasoning tokens your provider bills as output.

2. The hosted API you would pay per token

From your provider's price list or invoice. If you use caching or discounts, enter the average you actually pay.

Output tokens are usually priced higher than input tokens. Use the same currency as every other price here.

3. Your own GPUs
How are the GPUs paid for?

From your cloud's price list or contract. If one copy of the model spans several GPUs, treat the whole group as one GPU: enter its price and its throughput.

Include the share of the server, networking and install that each GPU needs.

The purchase price is spread evenly over this many years.

Electricity, rack space or colocation, and bandwidth, per GPU.

730 means always on: 8,760 hours in a year divided by 12.

Measured total (input plus output) tokens per second on your model, serving software, request mix and concurrency. Avoid peak figures from a spec sheet.

The share of that measured throughput you plan to use on average. Lower leaves more room for traffic peaks and response-time targets.

Extra GPUs kept for failures and maintenance. Enter 0 for none.

4. Everything else

Engineering and on-call time, monitoring, storage, networking and software licences. Enter 0 only if there is truly none.

Only changes how results are written. Enter every price in one currency.

How does the calculator work?

It follows ordinary cost arithmetic, one step at a time, using only what you entered.

Step What it calculates Formula
1 Tokens a month requests × input tokens per request; requests × output tokens per request
2 Hosted API cost (input tokens ÷ 1,000,000) × input price + (output tokens ÷ 1,000,000) × output price
3 API price per token API cost ÷ total tokens, which blends your two prices at your input and output mix
4 One GPU a month, rented price per GPU-hour × hours per month
5 One GPU a month, bought purchase price ÷ (years × 12) + power and hosting per month
6 Tokens one GPU serves a month tokens per second × target utilization × 3,600 × hours per month
7 GPUs needed total tokens ÷ tokens one GPU serves, rounded up to a whole GPU (at least 1), plus spare GPUs
8 Owned cost GPUs × cost of one GPU + other monthly cost
9 Break-even utilization owned cost ÷ (API price per token × GPUs × tokens per second × 3,600 × hours per month)

Utilization. Break-even utilization is the share of your fleet's measured peak throughput that you must actually use for owned capacity to cost the same as the API. Owned capacity costs less when your use is above it, and more when it is below.

Break-even volume. GPUs come in whole units, so the owned cost is a staircase while the API cost is a straight line. Let g be the cost of one GPU for a month, c the tokens one GPU serves in a month at your target utilization, p the API price per token, and F the cost of your spare GPUs plus your other monthly cost. Owned capacity can only win if p × c is more than g; otherwise even a fully used GPU costs more than the API tokens it replaces, and the calculator reports that the API is cheaper at every volume. When it can win:

  • The first break-even is at (F + k × g) ÷ p tokens a month, where k = F ÷ (p × c − g) rounded up, and at least 1.
  • Owned capacity stays cheaper from (F + q × g) ÷ p tokens a month, where q = (F + g) ÷ (p × c − g) rounded up.

The two differ because the owned cost jumps each time another GPU is needed, and just after a jump the API can be cheaper again for a while. These formulas are tested against a brute-force scan of both cost curves.

A worked example

These round numbers show the arithmetic. They are not prices, benchmarks or recommendations; the "Fill in example numbers" button loads them into the calculator.

The inputs: 20,000,000 requests a month, with 800 input and 200 output tokens each; an API that charges $1 per 1M input tokens and $4 per 1M output tokens; GPUs rented at $2.50 an hour for 730 hours; a measured throughput of 2,000 tokens a second per GPU; a 60% target utilization; one spare GPU; and $5,000 a month of other costs.

Step Working Result
Tokens a month 20,000,000 × 800 input; 20,000,000 × 200 output 16 billion in, 4 billion out, 20 billion in all
Hosted API 16,000 × $1 + 4,000 × $4 (millions of tokens) $32,000 a month, or $1.60 per 1M tokens
One GPU a month $2.50 × 730 hours $1,825
Tokens one GPU serves 2,000 × 0.60 × 3,600 × 730 3,153,600,000 a month
GPUs needed 20,000,000,000 ÷ 3,153,600,000 = 6.34, rounded up to 7, plus 1 spare 8 GPUs
Owned cost 8 × $1,825 + $5,000 $19,600 a month, or $0.98 per 1M tokens
Difference $32,000 − $19,600 $12,400 a month, 38.75% of the API cost
Utilization fleet peak is 8 × 5,256,000,000 = 42,048,000,000 tokens; the load is 20,000,000,000 47.6% used; owned breaks even at 29.1% (19,600 ÷ 67,276.8)
Break-even volume p × c − g = 5,045.76 − 1,825 = 3,220.76; F = 1,825 + 5,000 = 6,825; k = 6,825 ÷ 3,220.76 = 2.12, rounded up to 3 (6,825 + 3 × 1,825) ÷ 0.0000016 = 7,687,500,000 tokens, or 7,687,500 requests a month

At this volume the owned route is cheaper. At a tenth of it, 2,000,000 requests a month, the API costs $3,200 and the owned route still costs 2 × $1,825 + $5,000 = $8,650. Both halves of that picture matter, and the calculator shows both.

What does the calculator leave out?

It compares the cost of a token, not the value of an answer. Before you rely on a result, think about what it does not cover:

  • Throughput is one number. Real throughput depends on the model, its precision, context length, batching and serving software. Measure it on your own setup and request mix; don't copy a figure from another one.
  • Averages hide peaks. The calculator sizes a fleet for average load at your target utilization. Bursty traffic and response-time targets may need a lower utilization, which means more GPUs.
  • Whole GPUs. If one copy of your model spans several GPUs, treat the whole group as one unit: enter its price and its throughput, and read "GPUs" as groups.
  • The models are not the same. A hosted model and an open model on your GPUs may differ in quality, context length and features. Quality, evaluation and migration work are not priced.
  • Prices change and have terms. Volume discounts, committed-use contracts, reserved capacity, free tiers, caching and batch discounts exist on both sides. Enter effective averages, and re-run the numbers when a price changes.
  • Costs that aren't entered. Data transfer, taxes, resale value, cost of capital and the time spent setting up are not included unless you put them in "other monthly cost".
  • A steady month. The calculator describes one month at a steady volume. It does not model growth, ramp-up or seasonality; try other volumes to see them.
  • Non-cost reasons. Data residency, security rules and control over model versions are not priced. See the next section.

Where do you find each input?

Input Where it comes from
Requests, input and output tokens Your provider's usage reports, the token counts returned with each response, or your own application logs
API prices Your provider's price list or invoice. Anthropic and Google, for example, publish prices per million tokens with separate input and output rates (Anthropic, Google Gemini)
GPU price A cloud provider's price sheet, such as Google Cloud's Compute Engine pricing, a contract, or a hardware quote for bought GPUs
Throughput per GPU A load test of your own model and serving software, for example with vLLM's benchmark tool, using prompts like your real ones
Target utilization Your choice. Lower values leave room for peaks and response-time targets; higher values use the GPUs harder
Other monthly cost Your own budget: engineering and on-call time, monitoring, storage, networking and licences

When is cost not the only reason to own capacity?

Some teams run inference in their own environment for reasons a cost calculator can't price: data that must stay in a region or on a network they control, a need to run inside a customer's environment, control over which model version serves each request, or capacity that is predictable. Those reasons count even when the API is cheaper, and they don't change the arithmetic above. They change what the answer is worth to you.

Hyperlake builds for teams in that position. The use cases page explains why per-use charges and owned capacity can cross in either direction, and the FAQ covers what Hyperlake does and does not include.

See it live. A live demonstration: real deployed models analyze recorded robot feeds, with every result tied to evidence. Open the Physical AI reliability and safety operations demonstration.

The calculator is free to link to, quote and cite. Copy whichever version suits where you are writing.

If you share a result, say which inputs you used. The "Copy a link to these numbers" button puts them in the address so others can see the same scenario.

Frequently asked questions

When is self-hosting an LLM cheaper than paying per token?

When your volume keeps the GPUs busy enough that each token you serve costs less than the API's average price per token, after your fixed costs. Break-even utilization is the share of the GPUs' measured throughput you must use for that to be true. The calculator works it out for your inputs, and it can also show that the API is cheaper at every volume.

Does the calculator include prices for any provider or GPU?

No. It has no built-in prices, so nothing on it can go out of date or favor a vendor. You enter prices from your own invoices, price lists or quotes. The example button fills in round numbers only to show the arithmetic.

Is anything I enter stored or sent?

No. The calculation runs in your browser and the page makes no request with your numbers. The link button writes your numbers after the

How do I measure tokens per second per GPU?

Run a load test against your own model and serving software, with prompts like your real ones and at the concurrency you expect, then read the total token throughput (input plus output). The vLLM benchmark, for example, reports total token throughput and output token throughput separately. Throughput changes with model size, precision, context length and batching, so a figure from another setup can mislead.

How does the calculator decide how many GPUs I need?

It divides your monthly tokens by what one GPU serves in a month at your target utilization, rounds up to a whole GPU and adds the spare GPUs you chose. A lower target utilization leaves more room for traffic peaks and needs more GPUs.

Why can the owned cost be higher again just above the break-even volume?

GPUs come in whole units, so the owned cost jumps each time another GPU is needed, while the API cost rises smoothly. Just after a jump the API can be cheaper for a while. The result shows both the first break-even and the volume above which owned capacity stays cheaper.

Can I link to or cite this calculator?

Yes, there is no need to ask. Use the copyable link or citation on this page. Link to this page's address rather than to a result, unless you want to share a specific scenario.

Sources

Sources for the facts on this page, last checked October 5, 2026.

  1. Anthropic API pricing (prices per million tokens, input and output) checked October 5, 2026
  2. Gemini API pricing (prices per 1M tokens, input and output) checked October 5, 2026
  3. Compute Engine pricing, Google Cloud (GPU machine families) checked October 5, 2026
  4. vLLM Benchmark CLI (measuring serving throughput) checked October 5, 2026
  5. Hyperlake use cases (the agent scale effect) checked October 5, 2026

Start with a workload. Build the environment around it.

Tell us what you need to deploy, whose environment it must run in, and what it needs to connect to.