# What it costs to run an LLM, with the arithmetic shown

> Free calculators for LLM GPU memory, self-hosting versus API cost, fine-tuning cost and running models on your own GPU or Mac, and a GPU price table read from each provider's own page, AWS, Google Cloud and Azure included. Written by a vLLM contributor. Every number has a source, and nothing is called a benchmark until it has been measured.

By GPUCostLab. Updated 7 October 2026. Canonical URL: https://gpucostlab.com/

## Three numbers decide what self-hosting costs

1. **Memory: does the model fit?** Weights plus a KV cache that grows with every token of every running sequence. The [VRAM calculator](/llm-vram-calculator/) works it out from each model's `config.json`, including the cases most calculators get wrong: grouped-query attention, sliding windows, linear attention, latent attention (MLA), and tensor parallelism that does not split the cache.
2. **Price: what does an hour cost?** The [GPU price table](/gpu-cloud-prices/) lists on-demand prices per GPU-hour from eleven GPU clouds' own pricing pages, next to AWS, Google Cloud and Azure instances, with the date each was read and a link to the source.
3. **Throughput and utilization: how much work does that hour do?** This is the number that decides whether self-hosting beats an API, and no price list can tell you. The [break-even calculator](/self-host-vs-api-cost/) shows how the answer moves when it changes.

Training has the same three numbers. The [fine-tuning cost estimator](/llm-fine-tuning-cost/) works out memory for full fine-tuning, LoRA and QLoRA, whether it fits on the GPUs you pick, and the GPU-hours and cost of your tokens and epochs, with the throughput assumption shown and editable.

## Running models at home

The same arithmetic applies to your own graphics card or Mac, with two twists: when a model does not fit in VRAM, the part in system RAM sets the speed, and on a Mac the GPU may use only part of unified memory. The [Run AI at home](/run-ai-at-home/) section covers both. The [Can I run it?](/can-i-run-this-llm/) calculator takes any of 94 GPUs, Macs and mini-PCs and any preset model at GGUF, AWQ, FP8 or MXFP4 sizes, and the [electricity calculator](/local-llm-electricity-cost/) prices the result at your state's EIA electricity rate against an API. Guides cover [how much VRAM each model needs](/how-much-vram-to-run-llms-locally/), [GPUs by budget tier](/best-gpu-for-local-llms/) and [Apple Silicon against NVIDIA](/apple-silicon-vs-nvidia-for-local-ai/).

## How the site works

**Sources, not estimates.** Prices come from provider pages, GPU specifications from vendor datasheets, and model shapes from the models' published configs. Where a page doesn't state a value, the site says so instead of filling one in. The [methodology](/methodology/) lists every source and assumption.

**Formulas on the page.** Each calculator shows its arithmetic with your numbers, so you can check it or redo it in a spreadsheet.

**Measured benchmarks are planned, not published.** Throughput and cost-per-token measurements on real GPUs will appear with their data and scripts once they have been run. Until then, nothing here is presented as a measurement.

**Paid links are labeled.** The site may earn referral commissions from some GPU providers and, on the home-hardware guides, from retailers such as Amazon, labeled "(paid link)" wherever they appear. None are active today. Providers that pay nothing are listed exactly the same way. See the [affiliate disclosure](/affiliate-disclosure/).
