What it costs to run an LLM, with the arithmetic shown
Free calculators for LLM GPU memory, self-hosting versus API cost, fine-tuning cost and running models on your own GPU or Mac, and a GPU price table read from each provider's own page, AWS, Google Cloud and Azure included. Written by a vLLM contributor. Every number has a source, and nothing is called a benchmark until it has been measured.
Free tools
-
Can my computer run this model?
Your GPU, Mac or mini-PC against any preset model at any GGUF, AWQ, FP8 or MXFP4 size. Fits, partial CPU offload, or no, with the memory-bandwidth ceiling on tokens per second.
-
Electricity cost of AI at home
Rated or measured watts, hours and utilization, and your state's EIA electricity price, turned into kWh, dollars per month and per million tokens, next to an API's price for the same tokens.
-
LLM VRAM calculator
Weights plus KV cache for your context and batch, worked out from each model's config.json, and which GPUs hold it. Handles GQA, MLA, sliding-window and linear-attention models.
-
Self-host vs API break-even
Your monthly tokens on rented GPUs versus an API, using list prices from both sides, with the break-even volume and the throughput you would need to win.
-
GPU cloud price table
On-demand $/GPU-hour for H100, H200, B200, A100, L40S, RTX 4090 and more across eleven GPU clouds plus AWS, Google Cloud and Azure, each linked to its source. Sortable, filterable, downloadable.
-
Fine-tuning cost estimator
Full fine-tune, LoRA or QLoRA. Memory per GPU and whether it fits, GPU-hours for your tokens and epochs, and the cost at listed GPU prices. Throughput is a labeled assumption you can replace with your own.
Three numbers decide what self-hosting costs
- Memory: does the model fit? Weights plus a KV cache that grows with every token of every running sequence. The VRAM calculator works it out from each model's
config.json, including the cases most calculators get wrong: grouped-query attention, sliding windows, linear attention, latent attention (MLA), and tensor parallelism that does not split the cache. - Price: what does an hour cost? The GPU price table lists on-demand prices per GPU-hour from eleven GPU clouds' own pricing pages, next to AWS, Google Cloud and Azure instances, with the date each was read and a link to the source.
- Throughput and utilization: how much work does that hour do? This is the number that decides whether self-hosting beats an API, and no price list can tell you. The break-even calculator shows how the answer moves when it changes.
Training has the same three numbers. The fine-tuning cost estimator works out memory for full fine-tuning, LoRA and QLoRA, whether it fits on the GPUs you pick, and the GPU-hours and cost of your tokens and epochs, with the throughput assumption shown and editable.
Running models at home
The same arithmetic applies to your own graphics card or Mac, with two twists: when a model does not fit in VRAM, the part in system RAM sets the speed, and on a Mac the GPU may use only part of unified memory. The Run AI at home section covers both. The Can I run it? calculator takes any of 94 GPUs, Macs and mini-PCs and any preset model at GGUF, AWQ, FP8 or MXFP4 sizes, and the electricity calculator prices the result at your state's EIA electricity rate against an API. Guides cover how much VRAM each model needs, GPUs by budget tier and Apple Silicon against NVIDIA.
How the site works
Sources, not estimates. Prices come from provider pages, GPU specifications from vendor datasheets, and model shapes from the models' published configs. Where a page doesn't state a value, the site says so instead of filling one in. The methodology lists every source and assumption.
Formulas on the page. Each calculator shows its arithmetic with your numbers, so you can check it or redo it in a spreadsheet.
Measured benchmarks are planned, not published. Throughput and cost-per-token measurements on real GPUs will appear with their data and scripts once they have been run. Until then, nothing here is presented as a measurement.
Paid links are labeled. The site may earn referral commissions from some GPU providers and, on the home-hardware guides, from retailers such as Amazon, labeled "(paid link)" wherever they appear. None are active today. Providers that pay nothing are listed exactly the same way. See the affiliate disclosure.