Rules

  1. Every number has a source and a date. Prices, specifications and model shapes link to the page they were read from, with the date they were read.
  2. Unknown stays unknown. When a page does not state a value, the site shows "Not stated", "Check provider" or "n/a". Values are never filled in from memory, from third-party aggregators, or by guessing a form factor from a product name.
  3. Formulas are on the page. The calculators show their arithmetic with your numbers substituted, and their code is tested before every deploy.
  4. Assumptions are inputs. Anything that is not arithmetic or a sourced fact, such as runtime memory overhead, throughput or utilization, is an editable field labeled as an assumption.
  5. Measured means measured. No throughput, latency or cost-per-token result appears on this site until it has been measured with the method below and published with its data and scripts.

Data sources

GPU prices

prices.json holds on-demand list prices per GPU-hour from the public pricing pages of RunPod, Lambda, DigitalOcean (GPU Droplets and Paperspace), Vast.ai, CoreWeave, Hyperstack, Nebius, Modal, Crusoe, JarvisLabs and Voltage Park, and from the public price files of AWS, Google Cloud and Azure.

AWS, Google Cloud and Azure

The hyperscalers publish machine-readable prices that need no account, so the script reads those instead of scraping their pages. One representative US region each, on-demand Linux prices only, for a fixed list of GPU instance types (the single-GPU size where one exists and the full-machine size):

Provider Region Source read GPU names from
AWS us-east-1 The AWS Price List file for EC2 in us-east-1 (CSV, about 300 MB), streamed once per run and reduced to the rows for the listed instance types. Each row links to the exact file version read. AWS's instance-type pages (P6, P5, P4, G6e, G6, G5)
Google Cloud us-central1 The accelerator-optimized pricing page, whose tables show its default region, us-central1; the script checks the region name on every read. Google's Cloud Billing Catalog API needs an API key from a Google Cloud project, so it is not used. Google's accelerator-optimized machine documentation
Azure eastus The Azure Retail Prices API (prices.azure.com), Linux pay-as-you-go rows only. A size the API lists no eastus price for is read in eastus2, and its row says so (ND H200 v5 in October 2026). Each row links to the API query for that size and region. Azure's GPU size-series documentation

API prices

api-prices.json holds standard-tier list prices per million tokens from the pricing pages or pricing docs of OpenAI, Anthropic, Google, DeepSeek, Mistral, Together AI, Fireworks AI, DeepInfra, Groq, Hyperstack and Crusoe.

GPU specifications

gpus.json holds GPU specifications from NVIDIA's and AMD's datasheets, product pages and architecture whitepapers.

Model shapes

models.json holds model shapes copied by script from each model's config.json on Hugging Face.

Home hardware (Run AI at home)

consumer-hardware.json holds 94 graphics cards, Apple Silicon chips and unified-memory PCs: NVIDIA GeForce RTX 30, 40 and 50 series and RTX workstation cards, AMD Radeon RX 7000 and 9000 series and Radeon PRO and AI PRO cards, Intel Arc B-series, Apple M1 to M6 chips, AMD Ryzen AI Max and NVIDIA DGX Spark and RTX Spark.

Electricity prices

electricity-prices.json holds the average residential electricity price for each state, DC and the U.S. from the U.S. Energy Information Administration's Electric Power Monthly, Tables 5.6.A (latest month) and 5.6.B (year to date), with the U.S. figure checked against Table 5.3. EIA marks these values as preliminary. They are re-read when a new Electric Power Monthly is released (monthly).

Quantization formats

local-ai-assumptions.json holds the bits per weight of each GGUF type from llama.cpp's quantization README (measured there on Llama 3.1 8B), the block layouts behind llama.cpp's KV-cache types, AutoAWQ's and GPTQModel's default group size, and the RAM bandwidth presets, each with its source and commit.

Calculator assumptions

Assumption Value Why
Share of GPU memory the engine may use 0.92 vLLM's default gpu_memory_utilization
Runtime allowance per GPU 2 GiB A round number for the CUDA context, CUDA graphs, NCCL buffers and peak activations. Editable; vLLM logs the real figures at startup. Per-model measurements are planned.
GPU memory Marketed GB, treated as GiB The driver can report slightly less, for example with ECC enabled
Quantized weights Embeddings and LM head at 16-bit; group scales counted How common AWQ, GPTQ, FP8 and MXFP4 checkpoints are stored
Sliding-window layers Cache at most the window vLLM V1's hybrid KV-cache manager
Linear-attention state FP32 recurrent state, 16-bit convolution state The models' mamba_ssm_dtype setting
Hours in a month 730 8,760 hours a year / 12
Throughput (break-even calculator) No default you should keep Placeholders only; measure your own

Run AI at home calculators

Assumption Value Why
GGUF weight size Parameters x llama.cpp's effective bits per weight for that type The mixes keep some tensors at higher precision; figures measured on Llama 3.1 8B, so other architectures differ by a few percent
Token embeddings In system RAM, read one row per token (GGUF) llama.cpp always keeps the input layer on the CPU
Partial offload LM head first, then whole layers, each with its share of the cache llama.cpp's -ngl placement
Runtime allowance 1 GiB per GPU A round assumption for llama.cpp's compute buffers and the driver context; llama.cpp prints its buffer sizes at load. Editable
RAM kept for the OS and apps 4 GiB An assumption. Editable
System RAM bandwidth DDR5-5600, two channels: 89.6 GB/s peak Transfer rate x 8 bytes x channels; the top memory speed AMD lists for the Ryzen 9 9950X. Editable, with presets
Decode ceiling Bandwidth / (active weights + whole KV cache) per token An upper bound for one conversation; not a measurement
Several identical GPUs Memory adds; ceiling as one GPU reading all the bytes llama.cpp's default layer split runs the cards one after another
Rest of system 100 W around a graphics card, 30 W around an APU chip figure, 0 W for whole-computer figures An assumption; a plug-in meter replaces it
Days in a month 365 / 12 The same 730-hour month as the break-even calculator

Fine-tuning cost estimator

Assumption Value Why
Full fine-tune memory 2 + 2 + 12 bytes per parameter with mixed-precision AdamW; 8 or 2 bytes of optimizer state for the other optimizer choices ZeRO (K = 12), 8-bit optimizers
LoRA adapters Rank r on q, k, v and o of every layer; FP32 weights, gradients and AdamW moments (16 bytes per adapter parameter) LoRA; PEFT keeps adapters in FP32 by default. The trainable count is editable.
QLoRA frozen weights 4.127 bits per parameter; embeddings and LM head at 16-bit NF4 plus double-quantized constants, QLoRA
Activations 34 x s x b x h bytes per layer, or 2 x s x b x h per layer plus one recomputed layer with gradient checkpointing; FP32 logits Korthikanti et al. 2022, attention-score term dropped for FlashAttention-style kernels
Training FLOPs per token 6N + 12 x L x H x Q x T; 4N for LoRA and QLoRA PaLM appendix B; adapters skip the frozen weights' gradients
MFU 30% of peak dense BF16, editable A rule of thumb below the 38-46% reported for large, heavily tuned pre-training runs (PaLM, Llama 3). Not a measurement; enter your own tokens per second.
Framework overhead per GPU 2 GiB Same round allowance as the VRAM calculator. Editable.
Scaling across GPUs MFU unchanged as GPUs are added A simplification; the result warns when GPUs have no NVLink

Benchmarks (planned, no results yet)

Status: planned. Nothing below has been measured. The method is published first so it cannot be bent to fit the results.

Every benchmark will publish its scripts, raw results and run dates, and will be re-run when engines or prices change enough to matter.

Conflicts of interest

This site may earn referral or affiliate commissions from RunPod, Vast.ai, DigitalOcean and Lambda, and from Amazon on the home-hardware guides (see the affiliate disclosure; none are active today). Hardware is ranked by memory and bandwidth from vendor specifications, never by where it is sold or whether a retailer pays. Every provider in the price table, whether it pays a commission or not, is read the same way, and tables are ordered by price or by the column you choose.

Corrections

Corrections are dated and described on the page they affect. If a price, specification or model shape here is wrong, the contact details are on the about page.

Also available as Markdown.