# Self-hosting an LLM vs paying for an API: break-even calculator

> Monthly cost of serving your token volume on rented GPUs versus an API, the volume where they cross, and the throughput self-hosting would need to match the API. API and GPU list prices are linked to their sources and every input is editable.

By GPUCostLab. Updated 7 October 2026. Canonical URL: https://gpucostlab.com/self-host-vs-api-cost/

No paid links on this page: vendor links go straight to the vendor.

> **Interactive calculator.** Enter monthly input and output tokens, an API price, a GPU hourly price and your measured throughput on the HTML version of this page (https://gpucostlab.com/self-host-vs-api-cost/). It returns the monthly cost both ways, the break-even volume, the throughput needed to match the API, and how the answer moves if throughput halves or doubles. The API prices it uses are below; GPU prices are at https://gpucostlab.com/gpu-cloud-prices/.

### API list prices, USD per million tokens (read 2026-10-07)

| Provider | Model | Input | Output | Cached input | Source |
|---|---|---|---|---|---|
| OpenAI | gpt-5.4-mini | 0.75 | 4.5 | 0.075 | https://developers.openai.com/api/docs/pricing |
| OpenAI | gpt-5.4-nano | 0.2 | 1.25 | 0.02 | https://developers.openai.com/api/docs/pricing |
| OpenAI | gpt-6-astra | 10.0 | 50.0 | 1.0 | https://developers.openai.com/api/docs/pricing |
| OpenAI | gpt-6-luna | 0.1 | 0.5 | 0.01 | https://developers.openai.com/api/docs/pricing |
| OpenAI | gpt-6.1-sol | 2.0 | 10.0 | 0.1 | https://developers.openai.com/api/docs/pricing |
| Anthropic | Claude Fable 5.1 | 10.0 | 50.0 | 0.25 | https://platform.claude.com/docs/en/about-claude/pricing |
| Anthropic | Claude Haiku 4.5 | 1.0 | 5.0 | 0.1 | https://platform.claude.com/docs/en/about-claude/pricing |
| Anthropic | Claude Opus 5.5 | 4.0 | 20.0 | 0.2 | https://platform.claude.com/docs/en/about-claude/pricing |
| Anthropic | Claude Sonnet 5.5 | 2.0 | 10.0 | 0.2 | https://platform.claude.com/docs/en/about-claude/pricing |
| Google | Gemini 2.5 Pro (gemini-2.5-pro) | 1.25 | 10.0 | 0.125 | https://ai.google.dev/gemini-api/docs/pricing?hl=en |
| Google | Gemini 3.1 Pro Preview (gemini-3.1-pro-preview) | 2.0 | 12.0 | 0.2 | https://ai.google.dev/gemini-api/docs/pricing?hl=en |
| Google | Gemini 3.5 Flash (gemini-3.5-flash) | 1.5 | 9.0 | 0.15 | https://ai.google.dev/gemini-api/docs/pricing?hl=en |
| Google | Gemini 3.5 Flash-Lite (gemini-3.5-flash-lite) | 0.3 | 2.5 | 0.03 | https://ai.google.dev/gemini-api/docs/pricing?hl=en |
| Google | Gemini 3.8 Flash (gemini-3.8-flash) | 0.75 | 3.75 | 0.075 | https://ai.google.dev/gemini-api/docs/pricing?hl=en |
| DeepSeek | deepseek-flash (DeepSeek-V4.1-Flash) | 0.3 | 1.2 | 0.006 | https://api-docs.deepseek.com/quick_start/pricing/ |
| DeepSeek | deepseek-v4-pro (DeepSeek-V4-Pro-0813) | 1.32 | 3.96 | 0.044 | https://api-docs.deepseek.com/quick_start/pricing/ |
| Mistral | Mistral Large 3 | 0.5 | 1.5 | 0.05 | https://docs.mistral.ai/inference/pricing |
| Mistral | Mistral Large 4 | 0.68 | 2.09 | 0.07 | https://docs.mistral.ai/inference/pricing |
| Mistral | Mistral Medium 3.5 | 1.5 | 7.5 | 0.15 | https://docs.mistral.ai/inference/pricing |
| Mistral | Mistral Small 4 | 0.15 | 0.6 | 0.015 | https://docs.mistral.ai/inference/pricing |
| Together AI | DeepSeek V4 Flash 0731 | 0.14 | 0.28 | 0.03 | https://www.together.ai/pricing |
| Together AI | DeepSeek V4 Pro 0813 | 1.32 | 3.96 | 0.13 | https://www.together.ai/pricing |
| Together AI | DeepSeek V4.1 Flash | 0.3 | 1.2 | 0.006 | https://www.together.ai/pricing |
| Together AI | Gemma 4 31B | 0.39 | 0.97 | Not listed | https://www.together.ai/pricing |
| Together AI | GLM-5.3 | 1.4 | 4.4 | 0.26 | https://www.together.ai/pricing |
| Together AI | GLM-5.3-Flash | 0.15 | 0.5 | 0.03 | https://www.together.ai/pricing |
| Together AI | gpt-oss-120B | 0.15 | 0.6 | Not listed | https://www.together.ai/pricing |
| Together AI | Kimi K3 | 2.7 | 13.5 | 0.27 | https://www.together.ai/pricing |
| Together AI | Llama 3 8B Instruct Lite | 0.14 | 0.14 | Not listed | https://www.together.ai/pricing |
| Together AI | Llama 3.3 70B | 1.04 | 1.04 | Not listed | https://www.together.ai/pricing |
| Together AI | MiniMax M2.7 | 0.3 | 1.2 | 0.06 | https://www.together.ai/pricing |
| Together AI | MiniMax M3 | 0.3 | 1.2 | 0.06 | https://www.together.ai/pricing |
| Together AI | Qwen3 235B A22B Instruct 2507 FP8 Throughput | 0.2 | 0.6 | Not listed | https://www.together.ai/pricing |
| Together AI | Qwen3.5-397B-A17B | 0.6 | 3.6 | 0.35 | https://www.together.ai/pricing |
| Together AI | Qwen3.5 9B | 0.17 | 0.25 | Not listed | https://www.together.ai/pricing |
| Together AI | Qwen3.6-Plus | 0.5 | 3.0 | Not listed | https://www.together.ai/pricing |
| Together AI | Qwen3.8-2.4T-A95B | 2.0 | 6.0 | 0.25 | https://www.together.ai/pricing |
| Together AI | Qwen3.8 Flash | 0.15 | 0.47 | Not listed | https://www.together.ai/pricing |
| Fireworks AI | Other base models: 4B to 16B parameters | 0.2 | 0.2 | Not listed | https://docs.fireworks.ai/serverless/pricing |
| Fireworks AI | Other base models: more than 16B parameters | 0.9 | 0.9 | Not listed | https://docs.fireworks.ai/serverless/pricing |
| Fireworks AI | Other base models: less than 4B parameters | 0.1 | 0.1 | Not listed | https://docs.fireworks.ai/serverless/pricing |
| Fireworks AI | Other base models: MoE 56.1B to 176B parameters (e.g. DBRX, Mixtral 8x22B) | 1.2 | 1.2 | Not listed | https://docs.fireworks.ai/serverless/pricing |
| Fireworks AI | Other base models: MoE up to 56B parameters (e.g. Mixtral 8x7B) | 0.5 | 0.5 | Not listed | https://docs.fireworks.ai/serverless/pricing |
| Fireworks AI | DeepSeek V4.1 Flash | 0.3 | 1.2 | 0.006 | https://docs.fireworks.ai/serverless/pricing |
| Fireworks AI | GLM 5.3 | 1.4 | 4.4 | 0.26 | https://docs.fireworks.ai/serverless/pricing |
| Fireworks AI | GLM 5.3 Flash | 0.15 | 0.5 | 0.03 | https://docs.fireworks.ai/serverless/pricing |
| Fireworks AI | OpenAI GPT OSS 120B | 0.15 | 0.6 | 0.015 | https://docs.fireworks.ai/serverless/pricing |
| Fireworks AI | Kimi K3 | 3.0 | 15.0 | 0.3 | https://docs.fireworks.ai/serverless/pricing |
| Fireworks AI | MiniMax M3 | 0.3 | 1.2 | 0.06 | https://docs.fireworks.ai/serverless/pricing |
| Fireworks AI | Qwen 3.8 Max | 2.0 | 6.0 | 0.25 | https://docs.fireworks.ai/serverless/pricing |
| DeepInfra | DeepSeek-V3.2 | 0.26 | 0.38 | 0.13 | https://deepinfra.com/pricing |
| DeepInfra | DeepSeek-V4-Flash | 0.09 | 0.18 | 0.018 | https://deepinfra.com/pricing |
| DeepInfra | DeepSeek-V4-Flash-0731 | 0.06 | 0.18 | 0.015 | https://deepinfra.com/pricing |
| DeepInfra | DeepSeek-V4-Pro | 1.3 | 2.6 | 0.1 | https://deepinfra.com/pricing |
| DeepInfra | gemma-3-27b-it | 0.08 | 0.16 | Not listed | https://deepinfra.com/pricing |
| DeepInfra | gemma-4-26B-A4B-it | 0.07 | 0.34 | Not listed | https://deepinfra.com/pricing |
| DeepInfra | gemma-4-31B-it | 0.2 | 0.4 | Not listed | https://deepinfra.com/pricing |
| DeepInfra | Kimi-K2.6 | 0.75 | 3.5 | 0.15 | https://deepinfra.com/pricing |
| DeepInfra | Kimi-K3 | 2.85 | 14.25 | 0.285 | https://deepinfra.com/pricing |
| DeepInfra | Meta-Llama-3.1-8B-Instruct-Turbo | 0.02 | 0.04 | Not listed | https://deepinfra.com/pricing |
| DeepInfra | Llama-3.3-70B-Instruct-Turbo | 0.1 | 0.32 | Not listed | https://deepinfra.com/pricing |
| DeepInfra | Mistral-Small-3.2-24B-Instruct-2506 | 0.075 | 0.2 | Not listed | https://deepinfra.com/pricing |
| DeepInfra | Qwen3-14B | 0.12 | 0.24 | Not listed | https://deepinfra.com/pricing |
| DeepInfra | Qwen3-235B-A22B-Instruct-2507 | 0.09 | 0.55 | Not listed | https://deepinfra.com/pricing |
| DeepInfra | Qwen3-30B-A3B | 0.12 | 0.5 | Not listed | https://deepinfra.com/pricing |
| DeepInfra | Qwen3-32B | 0.08 | 0.28 | Not listed | https://deepinfra.com/pricing |
| DeepInfra | Qwen3.5-27B | 0.26 | 2.6 | Not listed | https://deepinfra.com/pricing |
| DeepInfra | Qwen3.5-35B-A3B | 0.14 | 1.0 | 0.05 | https://deepinfra.com/pricing |
| DeepInfra | Qwen3.5-397B-A17B | 0.45 | 3.0 | 0.22 | https://deepinfra.com/pricing |
| DeepInfra | Qwen3.5-9B | 0.1 | 0.15 | Not listed | https://deepinfra.com/pricing |
| DeepInfra | Qwen3.6-27B | 0.32 | 3.2 | Not listed | https://deepinfra.com/pricing |
| DeepInfra | Qwen3.6-35B-A3B | 0.1 | 0.95 | 0.1 | https://deepinfra.com/pricing |
| DeepInfra | Qwen3.8-Max | 1.65 | 4.951 | 0.206 | https://deepinfra.com/pricing |
| Groq | GPT OSS 120B (openai/gpt-oss-120b) | 0.15 | 0.6 | Not listed | https://console.groq.com/docs/models |
| Groq | GPT OSS 20B (openai/gpt-oss-20b) | 0.075 | 0.3 | Not listed | https://console.groq.com/docs/models |
| Groq | Llama 3.1 8B (llama-3.1-8b-instant) | Check provider | Check provider | Not listed | https://console.groq.com/docs/models |
| Groq | Llama 3.3 70B (llama-3.3-70b-versatile) | Check provider | Check provider | Not listed | https://console.groq.com/docs/models |
| Groq | MiniMax M2.7 (minimaxai/minimax-m2.7) | Check provider | Check provider | Not listed | https://console.groq.com/docs/models |
| Groq | Qwen/Qwen3.8-27B (qwen/qwen3.8-27b) | 0.8 | 4.0 | Not listed | https://console.groq.com/docs/models |
| Hyperstack | OpenAI gpt-oss-120b | 0.1 | 0.4 | Not listed | https://www.hyperstack.cloud/gpu-pricing |
| Hyperstack | Llama 3.1 8B | 0.2 | 0.2 | Not listed | https://www.hyperstack.cloud/gpu-pricing |
| Hyperstack | Llama 3.3 70B | 0.8 | 0.8 | Not listed | https://www.hyperstack.cloud/gpu-pricing |
| Crusoe | Gemma 4 31B-it | 0.14 | 0.4 | 0.14 | https://www.crusoe.ai/cloud/pricing |
| Crusoe | GLM 5.3 | 1.4 | 4.4 | 0.26 | https://www.crusoe.ai/cloud/pricing |
| Crusoe | GPT-OSS 120B | 0.05 | 0.2 | 0.05 | https://www.crusoe.ai/cloud/pricing |

Raw data: https://gpucostlab.com/tools/self-host-vs-api-break-even/api-prices.json

## How the costs are calculated

```
API cost per month       = input tokens x input price + output tokens x output price
                           (prompt-cached input at its own price, if the API has one)
GPU busy hours per month = input tokens / prefill tokens per second
                         + output tokens / decode tokens per second        (per replica)
Dedicated                = replicas x GPUs per replica x 730 hours x price per GPU-hour
                           replicas = busy hours / (730 x utilization), rounded up, at least 1
Hourly                   = busy hours / utilization x GPUs per replica x price per GPU-hour
Break-even volume        = the monthly volume, at your input/output mix, where the two costs are equal
```

A replica is one running copy of the model. If the model needs two GPUs, as Llama 3.3 70B at BF16 does on 80 GB cards, then a replica is two GPUs.

## Throughput is the assumption that decides this

Every price on the page is a list price you can check. Throughput is the one number nobody can look up for you. It depends on the model, the precision, the GPU, the engine and its version, how many requests run at once, and how long your prompts and answers are. Change it by a factor of two and the answer can flip. That is why the result always shows the cost at a quarter, half, double and four times your figure.

Two throughputs, not one:

- **Prefill** processes the prompt, many tokens per forward pass. It is mostly compute-bound and fast per token.
- **Decode** generates the answer, one token per sequence per forward pass. It is mostly bound by memory bandwidth, and it is where batching matters. A single stream is slow, and many concurrent streams share each read of the weights.

With continuous batching, both run on the same GPUs and compete for the same time, so the calculator adds their busy hours. Use the **total** decode rate across all concurrent streams at the concurrency you will actually run, not the speed of one stream.

To measure it, run your engine's serving benchmark (vLLM ships one as `vllm bench serve`) with your real distribution of prompt and output lengths, at the concurrency you plan to serve. Read input and output tokens per second. Measured throughput figures on this site are [planned](/methodology/); none are published yet.

## Utilization: you pay for idle GPUs

An API bills for tokens. A dedicated GPU bills for hours, busy or not. Traffic is uneven: if your peak hour carries five times your average load and you size for the peak, your replicas average about 20% busy. The utilization input is the share of paid GPU time spent working at the throughput you entered. The busy-time figure in the result shows what your inputs imply.

**Hourly** mode assumes you can stop paying the moment there is no traffic, for example with serverless GPUs billed per second. It ignores cold starts, minimum billing periods and capacity that is not there when you ask for it, so treat it as the best case.

## A worked example

These are the calculator's defaults. The throughput figures are placeholders, so the example shows the method, not the market.

- **Volume:** 300 million input and 60 million output tokens a month.
- **API:** Llama 3.3 70B from Together AI at $1.04 per million tokens in and out, which comes to $374 a month.
- **Self-hosted:** the same model at BF16 on two H100 SXM GPUs from RunPod Secure Cloud at $3.49 per GPU-hour. One replica running all month costs $5,095.
- **Busy time:** at 20,000 prefill and 2,000 decode tokens per second, the replica is busy 12.5 hours a month, 1.7% of the time it is paid for.

Self-hosting breaks even at about 4.9 billion tokens a month, 13.6 times this volume. Below that, no throughput can help, because one idle replica already costs more than the whole API bill.

At twenty times the volume (7.2 billion tokens a month), the same replica is busy a third of the time and costs $5,095 against an API bill of $7,488. Self-hosting wins, but only just: if real throughput is half the placeholder, it takes two replicas ($10,191) and the API wins again. That is the throughput point in practice.

## When self-hosting is worth it anyway

Cost is not the only reason. Self-host when:

- Data must stay in your own account or region.
- You run a fine-tuned model, or one that no API serves.
- You need control of latency, batching or the engine.
- Your volume is steady and high enough that the sensitivity table favors you at half your measured throughput.

Use an API when traffic is spiky or small, when an API serves the same open model at a price no single GPU can match, or when you have nobody to keep a serving stack up at night.

## What is not included

- Engineering and on-call time. Put it in "other self-hosting cost". It is often the largest line.
- Storage, egress, and CPU or RAM that some GPU providers bill separately.
- Quality differences between the API model and the one you host. Compare like with like: the price list marks open-weights models, and some hosts serve quantized versions.
- API rate limits, batch-API discounts (several providers list batch at half price; see the notes under the price table), and committed-use discounts on either side.

## Renting the GPUs

For testing a model before you commit, by-the-hour providers are the cheapest way to measure real throughput: [RunPod](https://www.runpod.io/) and the [Vast.ai](https://vast.ai/) marketplace for single cards, and [Lambda](https://lambda.ai/) and [DigitalOcean](https://www.digitalocean.com/) for H100, H200 and B200 instances. Compare current prices in the [GPU price table](/gpu-cloud-prices/), and size the replica with the [VRAM calculator](/llm-vram-calculator/).
