Free estimator
LLM fine-tuning cost estimator: LoRA, QLoRA or full fine-tune, memory and GPU-hours
Memory per GPU for full fine-tuning, LoRA and QLoRA, whether it fits on the GPUs you pick, the GPU-hours your tokens and epochs take at a stated throughput assumption, and the cost at on-demand list prices, including the cheapest listing for that GPU. Formulas and sources are on the page.
No paid links on this page: vendor links go straight to the vendor. Affiliate policy.
Fine-tuning cost estimator
Runs in your browser. Model shapes from each model's config.json, GPU peak throughput from vendor datasheets, and on-demand prices read 7 October 2026. Throughput is an assumption you set, not a measurement.
GPU compute and what it costs at list price
Peak dense BF16 throughput from each vendor's datasheet (7 October 2026) and the lowest on-demand price listed for that exact GPU in the price table (7 October 2026). The last two columns are arithmetic, not measurements: the price of 1018 FLOPs (one exaFLOP) at 100% of peak, and at 30% MFU, the estimator's default. A training job needs about 6 x parameters x tokens FLOPs for a full fine-tune and 4 x parameters x tokens for LoRA, plus attention.
| GPU | Memory | Peak BF16 TFLOPS (dense) | Lowest $/GPU-hour | Where | $ per exaFLOP at peak | $ per exaFLOP at 30% MFU |
|---|---|---|---|---|---|---|
| NVIDIA GeForce RTX 4090 | 24 GB | 165.2 | $0.340 | RunPod (Community Cloud) | $0.57 | $1.91 |
| NVIDIA GeForce RTX 5090 | 32 GB | 209.5 | $0.690 | RunPod (Community Cloud) | $0.91 | $3.05 |
| NVIDIA RTX A6000 | 48 GB | 154.8 | $0.330 | RunPod (Community Cloud) | $0.59 | $1.97 |
| NVIDIA RTX 6000 Ada Generation | 48 GB | 364 | $0.740 | RunPod (Community Cloud) | $0.56 | $1.88 |
| NVIDIA RTX PRO 6000 Blackwell Server Edition | 96 GB | Not published | $1.69 | RunPod (Community Cloud) | n/a | n/a |
| NVIDIA L4 | 24 GB | 121 | $0.440 | RunPod (Community Cloud) | $1.01 | $3.37 |
| NVIDIA A10 | 24 GB | 125 | $1.10 | Modal | $2.45 | $8.16 |
| NVIDIA A40 | 48 GB | 149.7 | $0.350 | RunPod (Community Cloud) | $0.65 | $2.16 |
| NVIDIA L40S | 48 GB | 362.05 | $0.790 | RunPod (Community Cloud) | $0.61 | $2.02 |
| NVIDIA A100 40GB SXM | 40 GB | 312 | $1.99 | Lambda | $1.77 | $5.91 |
| NVIDIA A100 80GB PCIe | 80 GB | 312 | $1.19 | RunPod (Community Cloud) | $1.06 | $3.53 |
| NVIDIA A100 80GB SXM | 80 GB | 312 | $1.39 | RunPod (Community Cloud) | $1.24 | $4.13 |
| NVIDIA H100 PCIe | 80 GB | 756 | $1.99 | RunPod (Community Cloud) | $0.73 | $2.44 |
| NVIDIA H100 NVL | 94 GB | 835.5 | $2.59 | RunPod (Community Cloud) | $0.86 | $2.87 |
| NVIDIA H100 SXM | 80 GB | 989.4 | starting at $1.99 | Voltage Park | $0.56 | $1.86 |
| NVIDIA H200 NVL | 141 GB | 835.5 | Not listed | n/a | n/a | |
| NVIDIA H200 SXM | 141 GB | 989.5 | $3.99 | Hyperstack | $1.12 | $3.73 |
| NVIDIA B200 (HGX B200) | 180 GB | 2250 | $5.98 | RunPod (Community Cloud) | $0.74 | $2.46 |
| NVIDIA B300 (HGX B300, Blackwell Ultra) | 270 GB | 2250 | $6.94 | RunPod (Community Cloud) | $0.86 | $2.86 |
| AMD Instinct MI300X | 192 GB | 1307.4 | $2.59 | DigitalOcean GPU Droplets | $0.55 | $1.83 |
| AMD Instinct MI325X | 256 GB | 1307.4 | $3.80 | DigitalOcean GPU Droplets | $0.81 | $2.69 |
| AMD Instinct MI355X | 288 GB | 2516.6 | Not listed | n/a | n/a |
Memory per parameter, by method
| Method | Weights | Gradients | Optimizer state (AdamW) | Source |
|---|---|---|---|---|
| Full fine-tune, mixed precision | 2 bytes (BF16) | 2 bytes | 12 bytes: FP32 master copy + two FP32 moments | ZeRO, section 3.1 |
| Full fine-tune, no master copy | 2 bytes | 2 bytes | 8 bytes: two FP32 moments | 8-bit optimizers, section 1.1 |
| Full fine-tune, 8-bit AdamW | 2 bytes | 2 bytes | 2 bytes: two 8-bit moments | 8-bit optimizers |
| LoRA | 2 bytes frozen; 4 bytes per adapter parameter (FP32) | 4 bytes per adapter parameter | 8 bytes per adapter parameter | LoRA |
| QLoRA | 0.516 bytes frozen (NF4 + double-quantized constants); embeddings and LM head 2 bytes | 4 bytes per adapter parameter | 8 bytes per adapter parameter | QLoRA, section 3 |
How the estimate is calculated
Memory per GPU = (weights + gradients + optimizer state) / GPUs, when sharded with FSDP or ZeRO-3
+ activations + logits + framework overhead (per GPU, never sharded)
Full fine-tune = 2 + 2 + 12 bytes per parameter (BF16 weights and gradients, AdamW
with an FP32 master copy and two moments)
LoRA = 2 bytes per frozen parameter + 16 bytes per adapter parameter
QLoRA = 4.127 bits per frozen parameter (NF4 + quantization constants) + 16 bytes per adapter parameter
Activations = 34 x sequence x micro-batch x hidden bytes per layer
or, with gradient checkpointing, 2 x sequence x micro-batch x hidden per layer + one layer's 34
Logits = 4 bytes x sequence x micro-batch x vocabulary
FLOPs per token = 6 x parameters used per token + 12 x layers x heads x head dim x sequence length
(4 x parameters for LoRA and QLoRA: no weight gradients for the frozen weights)
Tokens/s per GPU = MFU x peak dense BF16 FLOP/s / FLOPs per token (or your measured figure)
GPU-hours = dataset tokens x epochs / tokens/s per GPU / 3,600
Cost = GPU-hours x price per GPU-hour
Every constant comes from a paper you can check:
- 16 bytes per parameter for full fine-tuning. Mixed-precision Adam keeps 16-bit weights and gradients (2 + 2 bytes), plus an FP32 master copy of the weights and two FP32 moments (4 + 4 + 4 = 12 bytes). That accounting is from the ZeRO paper (Rajbhandari et al., 2020), which writes it as 2Ψ + 2Ψ + KΨ with K = 12. If your optimizer keeps no FP32 master copy, the state is 8 bytes. With 8-bit AdamW it is 2 bytes (Dettmers et al., 2022).
- LoRA adapters. LoRA trains two small matrices of rank r beside each adapted weight, r x (inputs + outputs) parameters per matrix. The estimator puts adapters on the four attention projections (q, k, v and o) of every layer. At rank 16 that is 13.6 million parameters for Llama 3.1 8B, 0.17% of the model. Adapters are kept in FP32, which is PEFT's default, so each costs 4 bytes of weight, 4 of gradient and 8 of AdamW moments. Adding the MLP projections roughly triples the adapter count, which is still small next to the frozen weights. If you know your count, enter it.
- QLoRA's 4.127 bits. The QLoRA paper stores frozen weights as 4-bit NormalFloat in blocks of 64 and quantizes the block constants again ("double quantization"), which brings them from 0.5 to 0.127 bits per parameter. The usual Transformers and bitsandbytes setup leaves the embeddings and the LM head unquantized, so they stay at 2 bytes, the same treatment as in the VRAM calculator.
- Activations. Korthikanti et al. (2022) count 34 x s x b x h bytes per transformer layer for 16-bit activations, plus 5 x a x s² x b for the attention scores. FlashAttention-style kernels do not store the attention-score matrices, so that second term is dropped. Checkpointing every layer ("full activation recomputation") keeps only each layer's 2 x s x b x h input and recomputes the rest, one layer at a time. Their constant assumes a GPT-style layer with a 4x MLP; models with wider MLPs or many experts store somewhat more.
- FLOPs per token. From the PaLM paper's appendix B (Chowdhery et al., 2022): 2N for the forward pass, 4N for the backward pass, plus 12 x L x H x Q x T for the attention matmuls. The backward pass does two matmuls for each forward one, one for the input gradients and one for the weight gradients. LoRA and QLoRA skip the weight gradients of the frozen weights, so the estimator uses 4N for them. For mixture-of-experts models, N is the parameters active per token, while memory holds all of them.
Throughput is the assumption that decides the cost
Memory is arithmetic. Throughput is not, and nobody can look it up for your job. It depends on the GPU, the model, sequence length, micro-batch, kernels, the trainer, and how many GPUs talk to each other over what link. The estimator expresses it as model FLOPs utilization (MFU): the share of the GPU's peak dense BF16 throughput that goes into the model's own FLOPs.
The default is 30%, and it is a rule of thumb, not a measurement. For scale: Google reported 46.2% MFU for PaLM 540B, and by the same accounting Megatron-Turing NLG 530B reached 29.7% (PaLM, section 4.1 and appendix B). Meta reported 38-43% BF16 MFU for Llama 3 405B pre-training on up to 16,384 H100s (Llama 3 paper, table 4). Those are pre-training runs on heavily tuned stacks. A fine-tuning job on one to eight GPUs with off-the-shelf tooling, short sequences and a micro-batch of one usually gets less. Gradient checkpointing costs time too: Korthikanti et al. measured 30-40% extra execution time for full recomputation, which under the MFU definition shows up as a lower MFU.
To replace the assumption, run a few hundred steps of your real job and read the tokens per second from your trainer's log (or samples per second times sequence length). Divide by the number of GPUs and enter it as measured throughput. The result then says "your measurement" instead of "assumed". The sensitivity table shows the cost if the real figure is half or double.
Measured fine-tuning throughput on the GPUs in the price table is planned. None is published yet.
A worked example
These are the estimator's defaults.
- Job: Llama 3.1 8B, LoRA at rank 16 on q, k, v and o (13.6 million trainable parameters), 50 million tokens for 3 epochs (150 million training tokens) at 4,096 tokens per sequence, one sequence per step, gradient checkpointing on, AdamW.
- Memory on one H100 SXM: frozen BF16 weights 14.96 GiB, adapters with their gradients and optimizer state 0.20 GiB, stored activations 1.00 GiB, one recomputed layer 0.53 GiB, FP32 logits 1.96 GiB (a 128,256-token vocabulary is expensive at the loss), and a 2 GiB overhead allowance: 20.6 GiB of 80. It fits with room to raise the micro-batch.
- Compute: 4 x 8.03 billion = 32.1 GFLOP per token for the weights, plus 12 x 32 layers x 32 heads x 128 x 4,096 = 6.4 GFLOP for attention, 38.6 GFLOP in all. At 30% of the H100 SXM's 989.4 dense BF16 TFLOPS, that is about 7,700 tokens per second, so 150 million tokens take 5.4 GPU-hours.
- Cost: at the lowest H100 SXM listing in the price table, Voltage Park's $1.99 per GPU-hour (a "starting at" price), about $11; at Google Cloud's on-demand a3-highgpu-8g rate of $11.06 per GPU, about $60 (prices read 7 October 2026).
The same job as a full fine-tune needs 16 bytes for each of 8.03 billion parameters: 125 GiB per GPU on one H100, which does not fit. Sharded across two or more H100s it does (20.4 GiB each across eight). Its compute is 54.6 GFLOP per token, 41% more than LoRA, so 7.7 GPU-hours at the same MFU. For comparison, the LoRA paper reports 43.1 tokens per second per V100 for LoRA against 32.5 for full fine-tuning on GPT-3 175B, which it calls a 25% speedup. That is a smaller gain than the FLOP counts alone predict (about 1.4 to 1.5 times), one more reason to measure your own.
Llama 3.3 70B with QLoRA fits on one 80 GB GPU at about 48 GiB, with 37 GiB of that the 4-bit weights. It takes about 44 GPU-hours for the same 150 million tokens at 30% MFU, before the dequantization overhead that QLoRA adds.
All of this is arithmetic from the stated assumptions, not a measurement.
LoRA, QLoRA or full fine-tuning: what each costs you
- Full fine-tuning changes every weight. Memory is 16 bytes per parameter before activations, so anything past a few billion parameters needs several GPUs with sharded optimizer state, and compute is the full 6N per token.
- LoRA freezes the model and trains adapters. Memory drops to the 16-bit weights plus a little, and compute drops by about a third. It is the usual choice when the model fits in 16-bit on the GPUs you have.
- QLoRA also stores the frozen weights in 4 bits, which is what puts a 70B model on one 80 GB GPU or an 8B model on a 24 GB card. The weights are dequantized for every matrix multiply, so it is slower than LoRA at the same FLOP count; the FLOP formula does not include that.
Whether LoRA matches full fine-tuning on quality depends on the task and the data, and is not something a cost estimator can tell you. Measure it on your own evaluation set.
Hyperscaler or GPU cloud
The price list includes AWS, Google Cloud and Azure on-demand instance prices alongside the GPU clouds. Per GPU-hour, hyperscaler on-demand prices are usually several times higher, but the comparison is not like for like. A hyperscaler instance includes its CPUs, memory, local storage and fast networking, and most of that capacity is bought with reservations or committed-use discounts that list prices do not show. Some of these instances come only as 8-GPU machines; the estimator says when that means paying for GPUs your job does not use. The GPU price table puts them side by side.
What is not included
- Failed and repeated runs, hyperparameter sweeps and evaluation. In practice these can cost more than the final run.
- Data preparation, storage, egress and checkpoint saving.
- Padding. If you do not pack sequences, padding tokens cost the same compute as real ones. Count them in the dataset tokens.
- Per-expert activation memory in mixture-of-experts models, attention FLOPs for latent-attention (MLA) and linear-attention layers, and memory spikes from optimizer steps or long-sequence outliers. The estimator says when a model is affected.
- Communication time across GPUs. The estimate assumes MFU holds as you add GPUs. Over PCIe without NVLink it usually does not.
Renting the GPUs
For a short measurement run to replace the throughput assumption, by-the-hour providers are the cheapest way in: RunPod and the Vast.ai marketplace for single cards, and Lambda and DigitalOcean for H100, H200 and B200 instances. The estimator's price list is the same data as the GPU price table, including the providers that pay this site nothing, and it is always sorted by price.
Also available as Markdown.