LLM VRAM calculator

Runs in your browser. Model shapes from each model's config.json on Hugging Face and GPU specs from vendor datasheets, read 7 October 2026. Both tables are below the calculator.

Model and precision
Workload
Engine assumptions (vLLM defaults)

gpu_memory_utilization is the share of each GPU's memory one vLLM instance may use; 0.92 is vLLM's current default. The runtime allowance covers the CUDA context, CUDA graphs, NCCL buffers and the activation peak of one forward pass. 2 GiB is a round assumption, not a measurement: vLLM logs the real figures at startup, so replace it with yours.

Model presets

Shapes copied from each model's config.json; parameter counts from the safetensors metadata on Hugging Face (what the checkpoint holds, including any vision encoder). Read 7 October 2026. Raw data: models.json.

ModelTotal paramsActiveAttention layers (KV heads x head dim)Max contextPublished asSource
Llama 3.1 8B Instruct 8.0B all 32 full (8 KV x 128) 131,072 BF16 config.json
Llama 3.3 70B Instruct 70.6B all 80 full (8 KV x 128) 131,072 BF16 config.json
Llama 3.1 405B Instruct 405.9B all 126 full (8 KV x 128) 131,072 BF16 config.json
Qwen3 8B 8.2B all 36 full (8 KV x 128) 40,960 BF16 config.json
Qwen3 14B 14.8B all 40 full (8 KV x 128) 40,960 BF16 config.json
Qwen3 32B 32.8B all 64 full (8 KV x 128) 40,960 BF16 config.json
Qwen2.5 7B Instruct 7.6B all 28 full (4 KV x 128) 32,768 BF16 config.json
Qwen2.5 72B Instruct 72.7B all 80 full (8 KV x 128) 32,768 BF16 config.json
Qwen3.8 27B 27.8B all 16 full (4 KV x 256) + 48 linear (fixed state) 262,144 BF16 config.json
Mistral 7B Instruct v0.3 7.2B all 32 full (8 KV x 128) 32,768 BF16 config.json
Mistral Small 3.2 24B Instruct (2506) 24.0B all 40 full (8 KV x 128) 131,072 BF16 config.json
Gemma 4 31B IT 31.3B all 50 sliding 1024 (16 KV x 256) + 10 full (4 KV x 512) 262,144 BF16 config.json
Qwen3 30B-A3B Instruct (2507) 30.5B 3.3B 48 full (4 KV x 128) 262,144 BF16 config.json
Qwen3 235B-A22B Instruct (2507) 235.1B 22.0B 94 full (4 KV x 128) 262,144 BF16 config.json
Qwen3.6 35B-A3B 36.0B 3.0B 10 full (2 KV x 256) + 30 linear (fixed state) 262,144 BF16 config.json
Qwen3.5 122B-A10B 125.1B 10.0B 12 full (2 KV x 256) + 36 linear (fixed state) 262,144 BF16 config.json
Gemma 4 26B-A4B IT 25.8B 3.8B 25 sliding 1024 (8 KV x 256) + 5 full (2 KV x 512) 262,144 BF16 config.json
gpt-oss-20b 20.9B 3.6B 12 sliding 128 (8 KV x 64) + 12 full (8 KV x 64) 131,072 MXFP4 config.json
gpt-oss-120b 116.8B 5.1B 18 sliding 128 (8 KV x 64) + 18 full (8 KV x 64) 131,072 MXFP4 config.json
GLM-4.5-Air 110.5B 12.0B 46 full (8 KV x 128) 131,072 BF16 config.json
MiniMax-M2.7 228.7B 11.0B 62 full (8 KV x 128) 204,800 FP8 config.json
DeepSeek-V3.2 685.4B 41.0B 61 MLA (latent 576) + 61 indexer 163,840 FP8 config.json
GLM-5.3 753.3B 41.8B 78 MLA (latent 576) + 21 indexer 1,048,576 FP8 config.json
Kimi K2 Instruct (0905) 1026.5B 32.0B 61 MLA (latent 576) 262,144 FP8 config.json
Notes on each preset (hybrid attention, gated repositories, what the counts include)
Llama 3.1 8B Instruct
meta-llama/Llama-3.1-8B-Instruct is gated on Hugging Face, so its config.json was read from unsloth/Llama-3.1-8B-Instruct, an ungated copy. The parameter count comes from meta-llama/Llama-3.1-8B-Instruct's own public safetensors metadata.
Llama 3.3 70B Instruct
meta-llama/Llama-3.3-70B-Instruct is gated on Hugging Face, so its config.json was read from unsloth/Llama-3.3-70B-Instruct, an ungated copy. The parameter count comes from meta-llama/Llama-3.3-70B-Instruct's own public safetensors metadata.
Llama 3.1 405B Instruct
Architecture fields read from unsloth's ungated copy of the config (the file also carries a bitsandbytes quantization block, which does not change the shapes). head_dim is not stated in this config; it is hidden_size / num_attention_heads = 128. meta-llama/Llama-3.1-405B-Instruct is gated on Hugging Face, so its config.json was read from unsloth/Meta-Llama-3.1-405B-Instruct-bnb-4bit, an ungated copy. The parameter count comes from meta-llama/Llama-3.1-405B-Instruct's own public safetensors metadata.
Qwen3 8B
Standard grouped-query attention on every layer.
Qwen3 14B
Standard grouped-query attention on every layer.
Qwen3 32B
Standard grouped-query attention on every layer.
Qwen2.5 7B Instruct
config.json sets sliding_window but use_sliding_window is false, so every layer is full attention.
Qwen2.5 72B Instruct
config.json sets sliding_window but use_sliding_window is false, so every layer is full attention.
Qwen3.8 27B
Hybrid attention: 3 of every 4 layers are Gated DeltaNet linear attention with a fixed-size state per sequence; only the full-attention layers keep a KV cache. Parameter count includes the vision encoder.
Mistral 7B Instruct v0.3
Standard grouped-query attention on every layer.
Mistral Small 3.2 24B Instruct (2506)
Parameter count includes the vision encoder.
Gemma 4 31B IT
5 of every 6 layers use 1,024-token sliding-window attention; the global layers use 4 KV heads of dimension 512 with keys and values from one projection (attention_k_eq_v). We count both K and V as cached, the conservative reading. Parameter count includes the vision encoder.
Qwen3 30B-A3B Instruct (2507)
Active parameters: model card: '30.5B in total and 3.3B activated'.
Qwen3 235B-A22B Instruct (2507)
Active parameters: model card: '235B in total and 22B activated'.
Qwen3.6 35B-A3B
Hybrid attention (Gated DeltaNet + full attention every 4th layer). Parameter count includes the vision encoder. Active parameters: model card: '35B in total and 3B activated'.
Qwen3.5 122B-A10B
Hybrid attention (Gated DeltaNet + full attention every 4th layer). Parameter count includes the vision encoder. Active parameters: model card: '122B in total and 10B activated'.
Gemma 4 26B-A4B IT
Sliding-window (1,024 tokens) on 5 of every 6 layers; global layers use 2 KV heads of dimension 512 (attention_k_eq_v; both K and V counted). Parameter count includes the vision encoder. Active parameters: model card: 'Active Parameters 3.8B'.
gpt-oss-20b
Published with MoE weights in MXFP4; attention, router, embeddings and LM head stay BF16. Alternating 128-token sliding-window and full-attention layers. Active parameters: model card: '21B parameters with 3.6B active parameters'.
gpt-oss-120b
Published with MoE weights in MXFP4; attention, router, embeddings and LM head stay BF16. Alternating 128-token sliding-window and full-attention layers. Active parameters: model card: '117B parameters with 5.1B active parameters'.
GLM-4.5-Air
Parameter count from the checkpoint (110.5B) includes the multi-token-prediction layer; the model card's 106B does not. Active parameters: model card: '106 billion total parameters and 12 billion active parameters'.
MiniMax-M2.7
Published as an FP8 (block-scaled) checkpoint. Active parameters: computed from config.json: total parameters minus the routed experts a token does not use, (experts - experts_per_token) x 3 x hidden_size x moe_intermediate_size per MoE layer; the publisher's card does not state it. Includes embeddings..
DeepSeek-V3.2
Multi-head latent attention (MLA) caches one 576-dim latent per token per layer, shared by all heads, plus a small FP8 key for the sparse-attention indexer. Published as FP8. Parameter count includes the 1-layer multi-token-prediction module, which is loaded only for speculative decoding. Active parameters: computed from config.json: total parameters minus the routed experts a token does not use, (experts - experts_per_token) x 3 x hidden_size x moe_intermediate_size per MoE layer; the publisher's card does not state it. Includes embeddings and the MTP layer..
GLM-5.3
DeepSeek-style MLA with a sparse-attention indexer; config.json marks 21 of 78 layers as having their own indexer ('full') and 57 as reusing one ('shared'), so the indexer cache is counted on 21 layers. FP8 KV layout assumed to match DeepSeek-V3.2's in vLLM. Parameter count includes the MTP layer. Active parameters: computed from config.json: total parameters minus the routed experts a token does not use, (experts - experts_per_token) x 3 x hidden_size x moe_intermediate_size per MoE layer; the publisher's card does not state it. Includes embeddings and the MTP layer..
Kimi K2 Instruct (0905)
DeepSeek-V3 architecture (MLA). Published as FP8. Active parameters: model card: '32 billion activated parameters and a total of 1 trillion parameters'.

GPU specifications

From each vendor's datasheet, product page or architecture whitepaper. Compute is dense tensor-core TFLOPS (no sparsity). "n/a" means the vendor does not publish it or the GPU lacks that data type. Raw data: gpus.json.

GPUMemoryBandwidthFP16/BF16FP8FP4GPU-to-GPUSource
NVIDIA GeForce RTX 4090 24 GB GDDR6X 1.01 TB/s 165.2 330.3 n/a PCIe Gen4, no NVLink NVIDIA
NVIDIA GeForce RTX 5090 32 GB GDDR7 1.79 TB/s 209.5 419 1676 PCIe Gen5, no NVLink NVIDIA
NVIDIA RTX A6000 48 GB GDDR6 0.77 TB/s 154.8 n/a n/a NVLink bridge (2 GPUs) 112.5 GB/s bidirectional; PCIe 4.0 x16 NVIDIA
NVIDIA RTX 6000 Ada Generation 48 GB GDDR6 0.96 TB/s 364 728.5 n/a PCIe 4.0 x16, no NVLink NVIDIA
NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB GDDR7 1.60 TB/s n/a n/a n/a PCIe Gen5 x16, no NVLink NVIDIA
NVIDIA L4 24 GB GDDR6 0.30 TB/s 121 242 n/a PCIe Gen4 x16 64 GB/s, no NVLink listed NVIDIA
NVIDIA A10 24 GB GDDR6 0.60 TB/s 125 n/a n/a PCIe Gen4 64 GB/s, no NVLink listed NVIDIA
NVIDIA A40 48 GB GDDR6 0.70 TB/s 149.7 n/a n/a NVLink bridge (2-way) 112.5 GB/s bidirectional; PCIe Gen4 NVIDIA
NVIDIA L40S 48 GB GDDR6 0.86 TB/s 362.05 733 n/a PCIe Gen4 x16 64 GB/s, no NVLink NVIDIA
NVIDIA A100 40GB SXM 40 GB HBM2 1.55 TB/s 312 n/a n/a NVLink 600 GB/s; PCIe Gen4 64 GB/s NVIDIA
NVIDIA A100 80GB PCIe 80 GB HBM2e 1.94 TB/s 312 n/a n/a NVLink bridge for 2 GPUs 600 GB/s; PCIe Gen4 64 GB/s NVIDIA
NVIDIA A100 80GB SXM 80 GB HBM2e 2.04 TB/s 312 n/a n/a NVLink 600 GB/s; PCIe Gen4 64 GB/s NVIDIA
NVIDIA H100 PCIe 80 GB HBM2e 2.00 TB/s 756 1513 n/a NVLink bridge (2 GPUs) 600 GB/s; PCIe Gen5 x16 NVIDIA
NVIDIA H100 NVL 94 GB HBM3 3.90 TB/s 835.5 1670.5 n/a NVLink bridge 600 GB/s; PCIe Gen5 128 GB/s NVIDIA
NVIDIA H100 SXM 80 GB HBM3 3.35 TB/s 989.4 1978.9 n/a NVLink 900 GB/s; PCIe Gen5 128 GB/s NVIDIA
NVIDIA H200 NVL 141 GB HBM3e 4.80 TB/s 835.5 1670.5 n/a 2- or 4-way NVLink bridge 900 GB/s per GPU; PCIe Gen5 128 GB/s NVIDIA
NVIDIA H200 SXM 141 GB HBM3e 4.80 TB/s 989.5 1979 n/a NVLink 900 GB/s; PCIe Gen5 128 GB/s NVIDIA
NVIDIA B200 (HGX B200) 180 GB HBM3E 8.00 TB/s 2250 4500 9000 NVLink 5 1.8 TB/s GPU-to-GPU (NVLink Switch) NVIDIA
NVIDIA B300 (HGX B300, Blackwell Ultra) 270 GB HBM3E 7.70 TB/s 2250 4500 14000 NVLink 5 1.8 TB/s; PCIe Gen6 256 GB/s NVIDIA
AMD Instinct MI300X 192 GB HBM3 5.30 TB/s 1307.4 2614.9 n/a AMD Infinity Fabric: 8 links, 128 GB/s peak link bandwidth; PCIe 5.0 x16 AMD
AMD Instinct MI325X 256 GB HBM3E 6.00 TB/s 1307.4 2614.9 n/a AMD Infinity Fabric: 8 links, 128 GB/s peak link bandwidth; PCIe 5.0 x16 AMD
AMD Instinct MI355X 288 GB HBM3E 8.00 TB/s 2516.6 5033.2 10066.3 AMD Infinity Fabric: 7 x 153.6 GB/s scale-up links; PCIe Gen5 x16 128 GB/s AMD
How each GPU's figures were read (sparsity, accumulate rate, conflicting sources)
NVIDIA GeForce RTX 4090
Product page lists only '1321 AI TOPS', 24 GB GDDR6X, 'Total Graphics Power (W) 450', 'NVLink (SLI-Ready) No'. Ada whitepaper V2.02 Appendix A: 'Peak FP16 Tensor TFLOPS with FP32 Accumulate 165.2/330.4' (used) vs 'with FP16 Accumulate 330.3/660.6'; FP8 'with FP32 Accumulate 330.3/660.6' (used, matches FP32-accumulate choice) vs 'with FP16 Accumulate 660.6/1321.2'; second figure = sparsity. Bandwidth '1008 GB/sec' from whitepaper. Second source: https://images.nvidia.com/aem-dam/Solutions/Data-Center/l....
NVIDIA GeForce RTX 5090
Product page lists only '3352 AI TOPS', '32 GB GDDR7', 'Total Graphics Power (W) 575', 'NVLink (SLI-Ready) No'. RTX Blackwell whitepaper v1.1 Table 3: 'Peak FP16 Tensor TFLOPS with FP32 Accumulate 209.5/419' (used) vs 'with FP16 Accumulate 419/838'; FP8 'with FP32 Accumulate 419/838' (used) vs 'with FP16 Accumulate 838/1676'; 'Peak FP4 Tensor TFLOPS with FP32 Accumulate (FP4 AI TOPS) 1676/3352'; second figure = sparsity. Bandwidth '1792 GB/sec' from whitepaper. Second source: https://images.nvidia.com/aem-dam/Solutions/geforce/black....
NVIDIA RTX A6000
Datasheet: 'Tensor performance 309.7 TFLOPS' with footnote 'Effective teraFLOPS (TFLOPS) using the new sparsity feature'. RTX PRO Blackwell whitepaper v1.1 Table 4 gives dense explicitly: 'Peak FP16 Tensor TFLOPS with FP32 Accumulate 154.8/309.6' (same as FP16 accumulate for this pro card). FP8 'N/A' (Ampere has no FP8 tensor cores). Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/design....
NVIDIA RTX 6000 Ada Generation
Datasheet: 'Tensor performance 1457.0 TFLOPS' = 'Effective FP8 teraFLOPS (TFLOPS) using the new sparsity feature'; 'NVIDIA NVLink No'. RTX PRO Blackwell whitepaper v1.1 Table 4: 'Peak FP16 Tensor TFLOPS with FP32 Accumulate 364/728' and 'Peak FP8 Tensor TFLOPS with FP32 Accumulate 728.5/1457' (FP16-accumulate rates identical on this pro card). Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/design....
NVIDIA RTX PRO 6000 Blackwell Server Edition
Compute left null: product page lists 'FP4 Tensor Core 4 PFLOPS', 'FP8 Tensor Core 2 PFLOPS', 'FP16 | BF16 Tensor Core 1 PFLOP' with NO sparsity marker, and the datasheet only gives 'Peak FP4 AI PFLOPS 4 PFLOPS'. NVIDIA's RTX PRO Blackwell whitepaper v1.1 lists the same-GPU Workstation Edition (GB202, 752 Tensor Cores, 126 TFLOPS FP32 vs 120 for Server Edition) at FP16 503.8/1007.6, FP8 1007.6/2015.2, FP4 2015.2/4030.4 TFLOPS (dense/sparse), so the page figures appear to be sparse; halving would give about 500/1000/2000 TFLOPS, but NVIDIA publishes no dense figure for the Server Edition. Product brief: 'Memory type GDDR7', 'Peak memory bandwidth 1,597 GB/s', 600 W, 'NVIDIA NVLink Not supported'. Second source: https://dam-cdn.nvd.orangelogic.com/AssetLink/3km2720jiy7....
NVIDIA L4
Product page: 'FP16 Tensor Core 242 teraFLOPS*', 'FP8 Tensor Core 485 teraFLOPs*', '* Shown with sparsity. Specifications are one-half lower without sparsity.' Ada whitepaper V2.02 Table 5 lists dense explicitly: 'FP16 Tensor Core Performance 121 | 242 TFLOPS', 'FP8 ... 242 | 485 TFLOPS', '24GB GDDR6 w/ ECC'. Neither source mentions NVLink; interconnect listed as PCIe only. Second source: https://images.nvidia.com/aem-dam/Solutions/Data-Center/l....
NVIDIA A10
Product page: 'FP16 Tensor Core 125 teraFLOPS | 250 teraFLOPS*', '*With Sparsity' (dense explicit). No FP8 tensor cores (Ampere). Interconnect listed only as 'PCIe Gen4 64GB/s'.
NVIDIA A40
Datasheet: 'Peak FP16 Tensor TFLOPS with FP16 Accumulate 149.7 | 299.4*' and 'Peak BF16 Tensor TFLOPS with FP32 Accumulate 149.7 | 299.4*', '* Structural sparsity enabled'. No FP8 tensor cores (Ampere). Datasheet lists PCIe Gen4 as 31.5 GB/s, product page as 64GB/s. Second source: https://www.nvidia.com/en-us/data-center/a40/.
NVIDIA L40S
Product page full spec table: 'FP16 Tensor Core 362.05 I 733*', 'FP8 Tensor Core 733 I 1,466*', '*With Sparsity' (dense listed explicitly; note 362.05 is not exactly half of 733). '48GB GDDR6 with ECC', 'Max Power Consumption 350W', 'NVIDIA NVLink Support: No'.
NVIDIA A100 40GB SXM
A100 datasheet (r4, 2021): 'FP16 Tensor Core 312 TFLOPS | 624 TFLOPS*', '* With sparsity'; A100 40GB SXM column: '40GB HBM2', '1,555GB/s', TDP 400W. No FP8 tensor cores (Ampere). The current A100 product page lists only the 80GB variants.
NVIDIA A100 80GB PCIe
Product page: 'FP16 Tensor Core 312 TFLOPS | 624 TFLOPS*', '* With sparsity' (dense explicit). No FP8 tensor cores (Ampere). NVLink only via bridge pairing two cards.
NVIDIA A100 80GB SXM
Product page: 'FP16 Tensor Core 312 TFLOPS | 624 TFLOPS*', '* With sparsity' (dense listed explicitly). No FP8 tensor cores (Ampere). '400W TDP for standard configuration. HGX A100-80GB custom thermal solution (CTS) SKU can support TDPs up to 500W'.
NVIDIA H100 PCIe
Hopper whitepaper v1.04 ('Includes final GPU / memory clocks and final TFLOPS performance specs') Table 3, H100 PCIe: 'Peak FP16 Tensor TFLOPS with FP32 Accumulate 756/1513' and 'Peak FP8 Tensor TFLOPS 1513/3026' (second figure = sparsity). Product brief PB-11133 v02: 'Memory type HBM2e', 'Peak memory bandwidth 2,000 GB/s', 350 W max, 'Total maximum NVLink bandwidth 600 Gbytes per second' (its overview text says 900 GB/s; table value used). Whitepaper lists bandwidth as 2039 GB/sec. NVIDIA's current H100 product page no longer lists H100 PCIe; an older 2022 H100 datasheet listed preliminary rounded figures (1,600 TFLOPS* FP16, 2TB/s), not used. Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/gtcs22....
NVIDIA H100 NVL
Product page: 'FP16 Tensor Core* 1,671 teraFLOPS', 'FP8 Tensor Core* 3,341 teraFLOPS', '* With sparsity'; dense values are half. TDP '350-400W (configurable)' (400 W recorded). Product brief PB-11773: 'Memory type HBM3', 'Memory size 94 GB', 'Peak memory bandwidth 3,938 GB/s' (page says 3.9TB/s). Second source: https://www.nvidia.com/content/dam/en-zz/Solutions/Data-C....
NVIDIA H100 SXM
Product page: 'FP16 Tensor Core* 1,979 teraFLOPS', 'FP8 Tensor Core* 3,958 teraFLOPS', '* With sparsity'; Hopper whitepaper v1.04 Table 3 gives dense explicitly: 'Peak FP16 Tensor TFLOPS with FP32 Accumulate 989.4/1978.9', FP8 1978.9/3957.8 (sparse after slash), and '80 GB HBM3'. TDP listed as 'Up to 700W (configurable)'. Second source: https://dam-cdn.nvd.orangelogic.com/AssetLink/705n6ur546g....
NVIDIA H200 NVL
Product page: 'FP16 Tensor Core² 1,671 TFLOPS', 'FP8 Tensor Core² 3,341 TFLOPS', '² With sparsity'; dense = half. Marked '¹ Preliminary specifications'. TDP 'Up to 600W (configurable)'. HBM3e from page text about H200 memory.
NVIDIA H200 SXM
Product page: 'FP16 Tensor Core² 1,979 TFLOPS', 'FP8 Tensor Core² 3,958 TFLOPS', '² With sparsity'; dense = half. Page footnote '¹ Preliminary specifications. May be subject to change.' Page text: '141 gigabytes (GB) of HBM3e memory at 4.8 terabytes per second'. TDP 'Up to 700W (configurable)'.
NVIDIA B200 (HGX B200)
No per-GPU B200 datasheet found; compute derived by dividing NVIDIA's 8-GPU HGX B200 totals by 8. HGX page: 'FP4 Tensor Core 144 PFLOPS | 72 PFLOPS' ('Sparse | Dense'), 'FP8/FP6 Tensor Core 72 PFLOPS', 'FP16/BF16 Tensor Core 36 PFLOPS' ('Specification in Sparse. Dense is 1/2 sparse spec shown.'). PCF summary: 'eight NVIDIA Blackwell B200 GPUs, each with 180 GB of HBM3E', 'Per individual GPU: Configurable up to 1000 W', total bandwidth 'Up to 62 TB/s'. Per-GPU bandwidth 'Up to 8TB/s' from https://docs.nvidia.com/enterprise-reference-architectures/hgx-ai-factory/latest/components.html (62/8 = 7.75 TB/s would follow from the PCF total). Second source: https://images.nvidia.com/aem-dam/Solutions/documents/HGX....
NVIDIA B300 (HGX B300, Blackwell Ultra)
NVIDIA Blackwell Ultra Datasheet, 'Individual Blackwell Ultra GPU Specifications', HGX B300 column: 'FP4 Tensor Core 18 PFLOPS | 14 PFLOPS' (Sparse | Dense), 'FP8/FP6 9 PFLOPS', 'FP16/BF16 4.5 PFLOPS' ('Specification in sparse. Dense is 1/2 sparse spec shown'), '270 GB HBM3E | 7.7 TB/s', 'Configurable up to 1,100 W' (GB300 NVL72 variant: 279 GB, 8 TB/s, 1,400 W). Conflicts: HGX page 8-GPU FP4 dense 108 PFLOPS (=13.5/GPU) and docs.nvidia.com reference architecture says B300 SXM 288GB, up to 8TB/s; datasheet per-GPU values used. Second source: https://www.nvidia.com/en-us/data-center/hgx/.
AMD Instinct MI300X
Product page: 'Peak Half Precision (FP16) Performance 1.3 PFLOPs' and 'with Structured Sparsity 2.61 PFLOPs'; FP8 '2.61 PFLOPs' / sparsity '5.22 PFLOPs'. Page footnote gives precise dense values: '1307.4 TFLOPS peak theoretical half precision (FP16)', '2614.9 TFLOPS peak theoretical 8-bit precision (FP8)'. TBP '750W Peak'. No FP4 listed.
AMD Instinct MI325X
Product page: FP16 '1.3 PFLOPs' (sparsity '2.61 PFLOPs'), FP8 '2.61 PFLOPs' (sparsity '5.22 PFLOPs'); footnote MI325-002: '1307.4 TFLOPS peak theoretical half precision (FP16)... 2614.9 TFLOPS peak theoretical 8-bit precision (FP8)'. TBP '1000W Peak'. No FP4 listed.
AMD Instinct MI355X
Product page: 'Peak Half Precision Matrix (FP16) Performance 2.5 PFLOPs' (sparsity '5 PFLOPs'), OCP-FP8 '5 PFLOPs' (sparsity '10.1 PFLOPs'), 'MXFP4 Performance 10.1 PFLOPs' (no sparsity figure), TBP '1400W'. Brochure gives precise values: FP16 matrix 2.5166 PFLOPS (5.0332 w/ sparsity), OCP-FP8 5.0332 (10.0664 w/ sparsity), MXFP4 10.0663 (sparsity N/A). FP4 is MXFP4 (microscaling). Second source: https://www.amd.com/content/dam/amd/en/documents/instinct....

The formula

Serving memory has three parts: the weights, the KV cache, and a runtime allowance. Only the KV cache depends on your traffic.

weights            = parameters x bytes per parameter
                     (quantized formats keep embeddings and the LM head at 16-bit)
KV cache per token = 2 x layers x KV heads x head dimension x bytes per value
KV cache in use    = KV per token x tokens per sequence x concurrent sequences
fits when          weights / TP + allowance + KV cache per GPU <= gpu_memory_utilization x GPU memory

The 2 counts keys and values. TP is the tensor-parallel size, the number of GPUs one copy of the model is split across.

Worked example. Llama 3.1 8B has 32 layers, 8 KV heads and a head dimension of 128, so at BF16 the cache is 2 x 32 x 8 x 128 x 2 bytes = 131,072 bytes (128 KiB) per token. A sequence at 8,192 tokens holds exactly 1 GiB. The weights are 8,030,261,248 parameters x 2 bytes = 14.96 GiB. An RTX 4090 has 24 GB; vLLM's default gpu_memory_utilization of 0.92 makes 22.08 GiB usable. Take away the weights and a 2 GiB allowance and 5.1 GiB is left, which is room for five concurrent 8k sequences. That is arithmetic from the formula, not a measurement.

KV heads, not attention heads

The cache stores keys and values for KV heads, and most current models have far fewer of those than query heads. Llama 3.1 8B has 32 query heads sharing 8 KV heads (grouped-query attention), which makes its cache four times smaller than if every head kept its own keys and values. Llama 3.3 70B also has only 8 KV heads, across 80 layers: 320 KiB per token. Qwen3 32B has 8 KV heads across 64 layers: 256 KiB per token. Calculators that use num_attention_heads overstate the cache by the grouping factor.

What changes the answer

Tensor parallelism does not always split the cache

With tensor parallelism, vLLM splits KV heads across GPUs, but a GPU never holds less than one head. When the tensor-parallel size reaches the number of KV heads, heads are replicated. Qwen3.5 122B-A10B has only 2 KV heads on its full-attention layers, so going from 2 to 8 GPUs barely shrinks the per-GPU cache (457 MiB to 402 MiB per 32k-token sequence; the part that still shrinks is the linear-attention state).

MLA is the extreme case. The latent is shared by all heads, so every tensor-parallel GPU keeps the full cache. At 32k tokens, a DeepSeek-V3.2 sequence takes 2.4 GiB on each of 8 GPUs. A Llama 3.3 70B sequence of the same length takes 10 GiB in total, which splits to 1.25 GiB per GPU. This is why large MLA deployments use data-parallel attention. The calculator models tensor parallelism only, and says which GPU counts a model's heads allow.

The runtime allowance

When vLLM starts, it loads the weights. Then it runs a profiling forward pass to measure peak activation memory and accounts for memory outside PyTorch's allocator (the CUDA context, NCCL buffers) and for CUDA graphs. Whatever is left of gpu_memory_utilization x GPU memory becomes KV cache. The calculator stands in for the middle part with one number per GPU. The default of 2 GiB is a stated round number, not a measurement. vLLM logs the real figures at startup, so replace it with yours.

GPU memory is the marketed figure treated as GiB. The capacity the driver reports can be a few percent lower, for example with ECC enabled, so leave headroom if you are within a few percent of the limit.

The decode ceiling column

Generating one token means reading every active weight and the sequence's whole cache from GPU memory. Memory bandwidth divided by those bytes is therefore an upper bound on single-stream decode speed. For Llama 3.1 8B at BF16 on an H100 SXM (3,350 GB/s) at 8k context, that is 3,350 GB/s / 17.1 GB = 196 tokens per second. Real engines land below it. Batching reads the weights once for many sequences, so total throughput goes far above one stream. Use the ceiling to compare GPUs, and the break-even calculator with your own measured throughput to compare costs. Measured numbers on this site are planned, not published.

What the calculator does not model

Where to run it

If the model fits a 24 to 32 GB card, hourly rentals of RTX 4090 and RTX 5090 cards are the cheapest way to test it: RunPod lists both, and Vast.ai is a marketplace where individual hosts set the price. For 70B-class models and long contexts you need 80 to 180 GB data-center GPUs from providers such as Lambda or DigitalOcean. The GPU price table compares on-demand prices from eleven GPU clouds and from AWS, Google Cloud and Azure, most of which pay this site nothing.

Before you rent anything, run the break-even calculator. For many open models, an API serving the same model costs less than one idle GPU.

Also available as Markdown.