# Can my computer run this LLM? Memory, CPU offload and speed for GPUs, Macs and mini-PCs

> Pick your graphics card, Mac or mini-PC, a model and a quantization (GGUF Q2_K to Q8_0, AWQ, FP8 or MXFP4). See whether it fits in GPU memory, how many layers spill into system RAM, and the most tokens per second your memory bandwidth allows. Specs come from the vendors' own pages, and the arithmetic is on the page.

By GPUCostLab. Updated 7 October 2026. Canonical URL: https://gpucostlab.com/can-i-run-this-llm/

> **Interactive calculator.** Pick a GPU, Mac or mini-PC (or enter memory and bandwidth), a model, a quantization, a context length and a KV-cache type on the HTML version of this page (https://gpucostlab.com/can-i-run-this-llm/). It returns weights, KV cache and total memory; whether the model fits in GPU memory, how many layers spill to system RAM, or whether it does not fit; and the decode-speed ceiling (memory bandwidth / bytes read per token), an upper bound and never a measurement. The data it uses is below.

### Hardware (vendor specs, read 2026-10-07)

| Device | Memory | Bandwidth (GB/s) | Rated power | Source |
|---|---|---|---|---|
| NVIDIA GeForce RTX 4060 | 8 GB | 272 | 115 W (Total Graphics Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 4060 Ti (8 GB) | 8 GB | 288 | 160 W (Total Graphics Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 5050 | 8 GB | 320 | 130 W (Total Graphics Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 3060 Ti | 8 GB | 448 | 200 W (Graphics Card Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 3070 | 8 GB | 448 | 220 W (Graphics Card Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 5060 | 8 GB | 448 | 145 W (Total Graphics Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 5060 Ti (8 GB) | 8 GB | 448 | 180 W (Total Graphics Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 3080 (10 GB) | 10 GB | 760 | 320 W (Graphics Card Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 3080 (12 GB) | 12 GB | n/a | 350 W (Graphics Card Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 3060 (12 GB) | 12 GB | 360 | 170 W (Graphics Card Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 4070 | 12 GB | 504 | 200 W (Total Graphics Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 4070 SUPER | 12 GB | 504 | 220 W (Total Graphics Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 4070 Ti | 12 GB | 504 | 285 W (Total Graphics Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 5070 | 12 GB | 672 | 250 W (Total Graphics Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 3080 Ti | 12 GB | 912 | 350 W (Graphics Card Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 4060 Ti (16 GB) | 16 GB | 288 | 165 W (Total Graphics Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 5060 Ti (16 GB) | 16 GB | 448 | 180 W (Total Graphics Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 4070 Ti SUPER | 16 GB | 672 | 285 W (Total Graphics Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 4080 | 16 GB | 716.8 | 320 W (Total Graphics Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 4080 SUPER | 16 GB | 736 | 320 W (Total Graphics Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 5070 Ti | 16 GB | 896 | 300 W (Total Graphics Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 5080 | 16 GB | 960 | 360 W (Total Graphics Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 3090 | 24 GB | 936 | 350 W (Graphics Card Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 3090 Ti | 24 GB | 1008 | 450 W (Graphics Card Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 4090 | 24 GB | 1008 | 450 W (Total Graphics Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA GeForce RTX 5090 | 32 GB | 1792 | 575 W (Total Graphics Power) | https://www.nvidia.com/en-us/geforce/graphics-cards/compare/ |
| NVIDIA RTX A4000 | 16 GB | 448 | 140 W (Max Power Consumption) | https://www.nvidia.com/en-us/products/workstations/rtx-a4000/ |
| NVIDIA RTX 4000 Ada Generation | 20 GB | 360 | 130 W (Max Power Consumption) | https://www.nvidia.com/en-us/products/workstations/rtx-4000/ |
| NVIDIA RTX 4500 Ada Generation | 24 GB | 432 | 210 W (Max Power Consumption) | https://www.nvidia.com/en-us/products/workstations/rtx-4500/ |
| NVIDIA RTX PRO 4000 Blackwell | 24 GB | 672 | 145 W (Max Power Consumption) | https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-4000/ |
| NVIDIA RTX A5000 | 24 GB | 768 | 230 W (Max Power Consumption) | https://www.nvidia.com/en-us/products/workstations/rtx-a5000/ |
| NVIDIA RTX 5000 Ada Generation | 32 GB | 576 | 250 W (Max Power Consumption) | https://www.nvidia.com/en-us/products/workstations/rtx-5000/ |
| NVIDIA RTX PRO 4500 Blackwell Workstation Edition | 32 GB | 896 | 200 W (Max Power Consumption) | https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-4500/ |
| NVIDIA RTX A6000 | 48 GB | 768 | 300 W (Max Power Consumption) | https://www.nvidia.com/en-us/products/workstations/rtx-a6000/ |
| NVIDIA RTX 6000 Ada Generation | 48 GB | 960 | 300 W (Max Power Consumption) | https://www.nvidia.com/en-us/products/workstations/rtx-6000/ |
| NVIDIA RTX PRO 5000 Blackwell (48 GB) | 48 GB | 1344 | 300 W (Max Power Consumption) | https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-5000/ |
| NVIDIA RTX PRO 5000 72GB Blackwell | 72 GB | 1344 | 300 W (Max Power Consumption) | https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-5000/ |
| NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition | 96 GB | 1792 | 300 W (Max Power Consumption) | https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-6000-max-q/ |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 96 GB | 1792 | 600 W (Max Power Consumption) | https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-6000/ |
| AMD Radeon RX 7600 | 8 GB | 288 | 165 W (Typical Board Power (Desktop)) | https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7600.html |
| AMD Radeon RX 9050 | 8 GB | 288 | 92 W (Typical Board Power (Desktop)) | https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9050.html |
| AMD Radeon RX 9060 | 8 GB | 288 | 132 W (Typical Board Power (Desktop)) | https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9060.html |
| AMD Radeon RX 9060 XT (8GB) | 8 GB | 320 | 150 W (Typical Board Power (Desktop)) | https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9060xt-8gb.html |
| AMD Radeon RX 7700 XT | 12 GB | 432 | 245 W (Typical Board Power (Desktop)) | https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7700-xt.html |
| AMD Radeon RX 9070 GRE | 12 GB | 432 | 220 W (Typical Board Power (Desktop)) | https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9070-gre.html |
| AMD Radeon RX 7600 XT | 16 GB | 288 | 190 W (Typical Board Power (Desktop)) | https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7600-xt.html |
| AMD Radeon RX 9060 XT (16GB) | 16 GB | 320 | 160 W (Typical Board Power (Desktop)) | https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9060xt.html |
| AMD Radeon RX 9060 XT LP | 16 GB | 320 | 140 W (Typical Board Power (Desktop)) | https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9060xt-lp.html |
| AMD Radeon RX 7900 GRE | 16 GB | 576 | 260 W (Typical Board Power (Desktop)) | https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7900-gre.html |
| AMD Radeon RX 7700 | 16 GB | 624 | 263 W (Typical Board Power (Desktop)) | https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7700.html |
| AMD Radeon RX 7800 XT | 16 GB | 624 | 263 W (Typical Board Power (Desktop)) | https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7800-xt.html |
| AMD Radeon RX 9070 | 16 GB | 640 | 220 W (Typical Board Power (Desktop)) | https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9070.html |
| AMD Radeon RX 9070 XT | 16 GB | 640 | 304 W (Typical Board Power (Desktop)) | https://www.amd.com/en/products/graphics/desktops/radeon/9000-series/amd-radeon-rx-9070xt.html |
| AMD Radeon RX 7900 XT | 20 GB | 800 | 315 W (Typical Board Power (Desktop)) | https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7900xt.html |
| AMD Radeon RX 7900 XTX | 24 GB | 960 | 355 W (Typical Board Power (Desktop)) | https://www.amd.com/en/products/graphics/desktops/radeon/7000-series/amd-radeon-rx-7900xtx.html |
| AMD Radeon PRO W7800 | 32 GB | 576 | 260 W (Total Board Power (TBP)) | https://www.amd.com/en/products/graphics/workstations/radeon-pro/w7800.html |
| AMD Radeon AI PRO R9600 | 32 GB | 640 | 150 W (Total Board Power (TBP)) | https://www.amd.com/en/products/graphics/workstations/radeon-ai-pro/ai-9000-series/amd-radeon-ai-pro-r9600.html |
| AMD Radeon AI PRO R9600D | 32 GB | 640 | 150 W (Total Board Power (TBP)) | https://www.amd.com/en/products/graphics/workstations/radeon-ai-pro/ai-9000-series/amd-radeon-ai-pro-r9600d.html |
| AMD Radeon AI PRO R9700 | 32 GB | 640 | 300 W (Total Board Power (TBP)) | https://www.amd.com/en/products/graphics/workstations/radeon-ai-pro/ai-9000-series/amd-radeon-ai-pro-r9700.html |
| AMD Radeon AI PRO R9700S | 32 GB | 640 | 300 W (Total Board Power (TBP)) | https://www.amd.com/en/products/graphics/workstations/radeon-ai-pro/ai-9000-series/amd-radeon-ai-pro-r9700s.html |
| AMD Radeon PRO W7800 48GB | 48 GB | 864 | 260 W (Total Board Power (TBP)) | https://www.amd.com/en/products/graphics/workstations/radeon-pro/w7800-48gb.html |
| AMD Radeon PRO W7900 | 48 GB | 864 | 295 W (Total Board Power (TBP)) | https://www.amd.com/en/products/graphics/workstations/radeon-pro/w7900.html |
| Intel Arc B570 Graphics | 10 GB | 380 | 150 W (Total Board Power (TBP)) | https://www.intel.com/content/www/us/en/products/sku/241676/intel-arc-b570-graphics/specifications.html |
| Intel Arc B580 Graphics | 12 GB | 456 | 190 W (Total Board Power (TBP)) | https://www.intel.com/content/www/us/en/products/sku/241598/intel-arc-b580-graphics/specifications.html |
| Intel Arc Pro B50 Graphics | 16 GB | 224 | 70 W (Total Board Power (TBP)) | https://www.intel.com/content/www/us/en/products/sku/242615/intel-arc-pro-b50-graphics/specifications.html |
| Intel Arc Pro B60 Graphics | 24 GB | 456 | 200 W (Total Board Power (TBP)) | https://www.intel.com/content/www/us/en/products/sku/243916/intel-arc-pro-b60-graphics/specifications.html |
| Intel Arc Pro B65 Graphics | 32 GB | 608 | 200 W (Total Board Power (TBP)) | https://www.intel.com/content/www/us/en/products/sku/245796/intel-arc-pro-b65-graphics/specifications.html |
| Intel Arc Pro B70 Graphics | 32 GB | 608 | 230 W (Total Board Power (TBP)) | https://www.intel.com/content/www/us/en/products/sku/245797/intel-arc-pro-b70-graphics/specifications.html |
| Apple M1 Pro | 16/32 GB unified | 200 | n/a | https://support.apple.com/en-us/111902 |
| Apple M1 Max | 32/64 GB unified | 400 | 115 W (Mac Studio (2022): Apple's maximum power for the whole computer, at the wall) | https://support.apple.com/en-us/111900 |
| Apple M1 Ultra | 64/128 GB unified | 800 | 215 W (Mac Studio (2022): Apple's maximum power for the whole computer, at the wall) | https://support.apple.com/en-us/111900 |
| Apple M2 Pro | 16/32 GB unified | 200 | 100 W (Mac mini (2023): Apple's maximum power for the whole computer, at the wall) | https://support.apple.com/en-us/111837 |
| Apple M3 Pro | 18/36 GB unified | 150 | n/a | https://support.apple.com/en-us/117736 |
| Apple M3 Max (14-core CPU, 30-core GPU) | 36/96 GB unified | 300 | n/a | https://support.apple.com/en-us/117736 |
| Apple M2 Max | 32/64/96 GB unified | 400 | 145 W (Mac Studio (2023): Apple's maximum power for the whole computer, at the wall) | https://support.apple.com/en-us/111835 |
| Apple M3 Max (16-core CPU, 40-core GPU) | 48/64/128 GB unified | 400 | n/a | https://support.apple.com/en-us/117736 |
| Apple M2 Ultra | 64/128/192 GB unified | 800 | 295 W (Mac Studio (2023): Apple's maximum power for the whole computer, at the wall) | https://support.apple.com/en-us/111835 |
| Apple M4 | 16/24/32 GB unified | 120 | 65 W (Mac mini (2024): Apple's maximum power for the whole computer, at the wall) | https://support.apple.com/en-us/121555 |
| Apple M4 Max (14-core CPU, 32-core GPU) | 36 GB unified | 410 | 145 W (Mac Studio (2025): Apple's maximum power for the whole computer, at the wall) | https://support.apple.com/en-us/122211 |
| Apple M4 Pro | 24/48/64 GB unified | 273 | 140 W (Mac mini (2024): Apple's maximum power for the whole computer, at the wall) | https://support.apple.com/en-us/121553 |
| Apple M4 Max (16-core CPU, 40-core GPU) | 48/64/128 GB unified | 546 | n/a | https://support.apple.com/en-us/122211 |
| Apple M5 | 16/24/32 GB unified | 153 | n/a | https://support.apple.com/en-us/125405 |
| Apple M3 Ultra | 96/256/512 GB unified | 819 | 270 W (Mac Studio (2025): Apple's maximum power for the whole computer, at the wall) | https://support.apple.com/en-us/122211 |
| Apple M6 | 16/24/32 GB unified | 170 | 70 W (Mac mini (M6, 2026): Apple's maximum power for the whole computer, at the wall) | https://support.apple.com/en-us/128108 |
| Apple M5 Max (18-core CPU, 32-core GPU) | 36 GB unified | 460 | 200 W (Mac Studio (M5 Max, 2026): Apple's maximum power for the whole computer, at the wall) | https://support.apple.com/en-us/128107 |
| Apple M5 Pro | 24/48/64 GB unified | 307 | 145 W (Mac mini (M5 Pro, 2026): Apple's maximum power for the whole computer, at the wall) | https://support.apple.com/en-us/126318 |
| Apple M5 Max (18-core CPU, 40-core GPU) | 48/64/128 GB unified | 614 | n/a | https://support.apple.com/en-us/128107 |
| Apple M5 Ultra | 96/256/512 GB unified | 1200 | 385 W (Mac Studio (M5 Ultra, 2026): Apple's maximum power for the whole computer, at the wall) | https://support.apple.com/en-us/128107 |
| NVIDIA RTX Spark (N1X) desktop PCs | 128 GB unified | n/a | 140 W (Chip TDP as NVIDIA lists it (the chip only, not the whole PC)) | https://www.nvidia.com/en-us/products/rtx-spark/ |
| AMD Ryzen AI Max 385 (Radeon 8050S) | 128 GB unified | 256 (derived) | 120 W (Top of AMD's configurable TDP range (45-120 W; default 55 W): the chip only, and the computer maker sets it) | https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-max-385.html |
| AMD Ryzen AI Max 390 (Radeon 8050S) | 128 GB unified | 256 (derived) | 120 W (Top of AMD's configurable TDP range (45-120 W; default 55 W): the chip only, and the computer maker sets it) | https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-max-390.html |
| AMD Ryzen AI Max+ 395 (Radeon 8060S) | 128 GB unified | 256 | 120 W (Top of AMD's configurable TDP range (45-120 W; default 55 W): the chip only, and the computer maker sets it) | https://www.amd.com/en/products/processors/laptop/ryzen/ai-300-series/amd-ryzen-ai-max-plus-395.html |
| NVIDIA DGX Spark | 64/128 GB unified | 273 | 240 W (Power supply / power consumption of the whole DGX Spark, as NVIDIA lists it) | https://www.nvidia.com/en-us/products/workstations/dgx-spark/ |
| AMD Ryzen AI Max+ PRO 495 (Radeon 8065S) | 192 GB unified | 273.1 (derived) | 120 W (Top of AMD's configurable TDP range (45-120 W; default 55 W): the chip only, and the computer maker sets it) | https://www.amd.com/en/products/processors/laptop/ryzen-pro/ai-max-pro-400-series/amd-ryzen-ai-max-plus-pro-495.html |

### Weight formats

| Format | Used by | Bits per weight | Basis | Source |
|---|---|---|---|---|
| Q2_K | GGUF (llama.cpp, Ollama, LM Studio) | 3.159 | Effective bits per weight of a whole Q2_K file of Llama 3.1 8B in llama.cpp's table (2.95 GiB GiB). Mixes keep some tensors at higher precision, so other architectures differ by a few percent. | https://github.com/ggml-org/llama.cpp/blob/9c2e0e491a822adae1f0b1c831adb4160057d24f/tools/quantize/README.md#L141-L176 |
| Q3_K_S | GGUF (llama.cpp, Ollama, LM Studio) | 3.643 | Effective bits per weight of a whole Q3_K_S file of Llama 3.1 8B in llama.cpp's table (3.41 GiB GiB). Mixes keep some tensors at higher precision, so other architectures differ by a few percent. | https://github.com/ggml-org/llama.cpp/blob/9c2e0e491a822adae1f0b1c831adb4160057d24f/tools/quantize/README.md#L141-L176 |
| Q3_K_M | GGUF (llama.cpp, Ollama, LM Studio) | 3.996 | Effective bits per weight of a whole Q3_K_M file of Llama 3.1 8B in llama.cpp's table (3.74 GiB GiB). Mixes keep some tensors at higher precision, so other architectures differ by a few percent. | https://github.com/ggml-org/llama.cpp/blob/9c2e0e491a822adae1f0b1c831adb4160057d24f/tools/quantize/README.md#L141-L176 |
| IQ4_XS | GGUF (llama.cpp, Ollama, LM Studio) | 4.46 | Effective bits per weight of a whole IQ4_XS file of Llama 3.1 8B in llama.cpp's table (4.17 GiB GiB). Mixes keep some tensors at higher precision, so other architectures differ by a few percent. | https://github.com/ggml-org/llama.cpp/blob/9c2e0e491a822adae1f0b1c831adb4160057d24f/tools/quantize/README.md#L141-L176 |
| Q4_K_S | GGUF (llama.cpp, Ollama, LM Studio) | 4.667 | Effective bits per weight of a whole Q4_K_S file of Llama 3.1 8B in llama.cpp's table (4.36 GiB GiB). Mixes keep some tensors at higher precision, so other architectures differ by a few percent. | https://github.com/ggml-org/llama.cpp/blob/9c2e0e491a822adae1f0b1c831adb4160057d24f/tools/quantize/README.md#L141-L176 |
| Q4_K_M | GGUF (llama.cpp, Ollama, LM Studio) | 4.894 | Effective bits per weight of a whole Q4_K_M file of Llama 3.1 8B in llama.cpp's table (4.58 GiB GiB). Mixes keep some tensors at higher precision, so other architectures differ by a few percent. Q4_K itself is 4.5 bits (144 bytes per 256 weights); Q4_K_M stores half of the attention V and FFN down tensors, and the output tensor, in Q6_K. | https://github.com/ggml-org/llama.cpp/blob/9c2e0e491a822adae1f0b1c831adb4160057d24f/tools/quantize/README.md#L141-L176 |
| Q5_K_S | GGUF (llama.cpp, Ollama, LM Studio) | 5.57 | Effective bits per weight of a whole Q5_K_S file of Llama 3.1 8B in llama.cpp's table (5.21 GiB GiB). Mixes keep some tensors at higher precision, so other architectures differ by a few percent. | https://github.com/ggml-org/llama.cpp/blob/9c2e0e491a822adae1f0b1c831adb4160057d24f/tools/quantize/README.md#L141-L176 |
| Q5_K_M | GGUF (llama.cpp, Ollama, LM Studio) | 5.704 | Effective bits per weight of a whole Q5_K_M file of Llama 3.1 8B in llama.cpp's table (5.33 GiB GiB). Mixes keep some tensors at higher precision, so other architectures differ by a few percent. | https://github.com/ggml-org/llama.cpp/blob/9c2e0e491a822adae1f0b1c831adb4160057d24f/tools/quantize/README.md#L141-L176 |
| Q6_K | GGUF (llama.cpp, Ollama, LM Studio) | 6.563 | Effective bits per weight of a whole Q6_K file of Llama 3.1 8B in llama.cpp's table (6.14 GiB GiB). Mixes keep some tensors at higher precision, so other architectures differ by a few percent. | https://github.com/ggml-org/llama.cpp/blob/9c2e0e491a822adae1f0b1c831adb4160057d24f/tools/quantize/README.md#L141-L176 |
| Q8_0 | GGUF (llama.cpp, Ollama, LM Studio) | 8.501 | Effective bits per weight of a whole Q8_0 file of Llama 3.1 8B in llama.cpp's table (7.95 GiB GiB). Mixes keep some tensors at higher precision, so other architectures differ by a few percent. Q8_0 itself is 8.5 bits (34 bytes per 32 weights). | https://github.com/ggml-org/llama.cpp/blob/9c2e0e491a822adae1f0b1c831adb4160057d24f/tools/quantize/README.md#L141-L176 |
| AWQ / GPTQ 4-bit (group size 128) | GPU engines (vLLM, SGLang) | 4.156 | 4-bit weights plus a 16-bit scale and a 4-bit zero point per group of 128 (4.156 bits); embeddings and LM head kept at 16-bit, as the VRAM calculator counts them. Group size 128 is AutoAWQ's and GPTQModel's default. | https://github.com/casper-hansen/AutoAWQ/blob/bcaa8a3689e6ae2e84ec6b57e995ee8a7904a19e/awq/models/_config.py#L11-L13 |
| FP8 (8-bit float) | GPU engines (vLLM, SGLang) | 8 | 1 byte per weight; embeddings and LM head kept at 16-bit. Fast only on GPUs with FP8 tensor cores (NVIDIA Ada and newer). | https://docs.vllm.ai/en/latest/features/quantization/fp8/ |
| MXFP4 (how gpt-oss is published) | As published | 4.25 | 4-bit values plus one 8-bit shared scale per block of 32 (17 bytes per 32 weights = 4.25 bits); the tensors the checkpoint keeps in BF16 stay at 16-bit. | https://github.com/ggml-org/llama.cpp/blob/9c2e0e491a822adae1f0b1c831adb4160057d24f/ggml/src/ggml-common.h#L219 |
| F16 / BF16 (unquantized) | Unquantized | 16 | 2 bytes per weight. llama.cpp's table lists F16 at 16.0005 bits for the whole file. | https://github.com/ggml-org/llama.cpp/blob/9c2e0e491a822adae1f0b1c831adb4160057d24f/tools/quantize/README.md#L141-L176 |

### KV-cache types

| Type | Bytes per value | Basis | Source |
|---|---|---|---|
| F16 (llama.cpp default) / BF16 | 2 | 16-bit keys and values; llama.cpp's default for --cache-type-k and --cache-type-v. | https://github.com/ggml-org/llama.cpp/blob/9c2e0e491a822adae1f0b1c831adb4160057d24f/tools/server/README.md#L71 |
| q8_0 (llama.cpp) | 1.062 | 32 int8 values and one 16-bit scale per block: 34 bytes per 32 values. A quantized V cache needs flash attention, which llama.cpp turns on by default (-fa auto). | https://github.com/ggml-org/llama.cpp/blob/9c2e0e491a822adae1f0b1c831adb4160057d24f/ggml/src/ggml-common.h#L256 |
| q4_0 (llama.cpp) | 0.5625 | 32 4-bit values and one 16-bit scale per block: 18 bytes per 32 values. Needs flash attention for the V cache, as above. | https://github.com/ggml-org/llama.cpp/blob/9c2e0e491a822adae1f0b1c831adb4160057d24f/ggml/src/ggml-common.h#L199 |
| FP8 (vLLM --kv-cache-dtype fp8) | 1 | 1 byte per value; the VRAM calculator's layout, including the FP8 layout vLLM uses for latent-attention (MLA) caches. | https://docs.vllm.ai/en/latest/features/quantization/quantized_kvcache/ |

### Unified memory: how much the GPU may use

- **Apple Silicon.** macOS lets the GPU use only part of unified memory by default, and Apple documents no fixed rule. Metal reports the limit at run time as recommendedMaxWorkingSetSize, "an approximation of how much memory, in bytes, this GPU device can allocate without affecting its runtime performance", and llama.cpp and Ollama read it as the Mac's GPU memory. Apple's 2021 tech talk on Metal compute gives two examples: an M1 Pro or M1 Max with 32 GB lets the GPU access 21 GB, and an M1 Max with 64 GB, 48 GB. The calculator follows those examples: two thirds of memory up to 32 GB, three quarters above; enter your own share if your Mac reports a different limit. Apple's MLX project documents raising the limit with sudo sysctl iogpu.wired_limit_mb=<megabytes>, to a value larger than the model but smaller than the machine's memory. Community reports say the setting does not survive a reboot, and an MLX maintainer linked a kernel panic to a limit set too high, so leave macOS several gigabytes. Sources: https://developer.apple.com/documentation/metal/mtldevice/recommendedmaxworkingsetsize, https://developer.apple.com/videos/play/tech-talks/10580/, https://github.com/ml-explore/mlx-lm#large-models, https://ml-explore.github.io/mlx/build/html/python/_autosummary/mlx.core.set_wired_limit.html, https://github.com/ml-explore/mlx-lm/issues/883, https://github.com/ggml-org/llama.cpp/discussions/2182
- **AMD Ryzen AI Max.** AMD says up to 96 GB of a 128 GB Ryzen AI Max system can be assigned to graphics with Variable Graphics Memory, and up to 160 GB of 192 GB on the Ryzen AI Max+ PRO 495; the calculator uses those shares (75% and 83%). On Linux the GPU can also map system memory through GTT, which AMD's ROCm guide says defaults to about half of RAM and can be raised, so the usable share depends on how the machine is set up. Enter your own share if you have changed it. Sources: https://ir.amd.com/news-events/press-releases/detail/1232/amd-announces-expanded-consumer-and-commercial-ai-pc, https://www.amd.com/en/blogs/2025/amd-ryzen-ai-max-395-processor-breakthrough-ai-.html, https://www.amd.com/en/blogs/2026/amd-powers-next-generation-agent-computers-with-new-ryzen-ai-hal.html, https://rocm.docs.amd.com/en/latest/reference/system-optimization/rdna3-5.html
- **NVIDIA DGX Spark and RTX Spark PCs.** NVIDIA documents no separate GPU limit for the DGX Spark's unified LPDDR5X (128 GB, and 64 GB in versions sold through other makers), so the calculator lets the GPU use all of it except the RAM kept for the operating system. NVIDIA publishes no memory bandwidth for RTX Spark PCs, so the calculator shows no speed ceiling for them. Sources: https://www.nvidia.com/en-us/products/workstations/dgx-spark/, https://www.nvidia.com/en-us/products/rtx-spark/

Raw data: https://gpucostlab.com/tools/local-ai-can-i-run/consumer-hardware.json and https://gpucostlab.com/tools/local-ai-can-i-run/local-ai-assumptions.json

## How the answer is worked out

```
weights          = parameters x bits per weight / 8
KV cache         = 2 x layers x KV heads x head dim x bytes per value x context tokens
total            = weights + KV cache + runtime allowance
fits on the GPU  when total - token embeddings <= GPU memory
                 (unified memory: memory x the share the GPU may use)
partial offload  layers on GPU = floor((GPU memory - allowance - LM head) / (one layer's weights + its cache))
decode ceiling   = memory bandwidth / bytes read per generated token
bytes per token  = weights x (active / total parameters) + the whole KV cache
offload ceiling  = 1 / (bytes in VRAM / GPU bandwidth + bytes in RAM / RAM bandwidth)
```

The memory arithmetic for the cache is the [VRAM calculator](/llm-vram-calculator/)'s own code, imported rather than copied, so grouped-query attention, sliding windows, linear attention and latent attention (MLA) are counted the same way on both pages. Model shapes come from each model's `config.json`.

Two details follow llama.cpp's source code. It keeps the token-embedding table in system RAM ("there is very little benefit to offloading the input layer, so always keep it on the CPU"), so that table never counts against VRAM and only one row of it is read per token. And when you offload with `-ngl`, the output layer (the LM head) goes to the GPU first, then whole layers. The [sources are listed in the data file](/tools/local-ai-can-i-run/local-ai-assumptions.json).

## Bits per weight: what a quantization costs

A GGUF file is not "parameters x 4 bits". The k-quant mixes keep some tensors at higher precision, and every block stores scales. llama.cpp's own quantization README lists the bits per weight of each file type, measured on Llama 3.1 8B, and the calculator uses those figures for every model:

| Format | Bits per weight | Llama 3.1 8B weights | Llama 3.3 70B weights |
|---|---|---|---|
| Q2_K | 3.16 | 2.95 GiB | 25.95 GiB |
| Q3_K_M | 4.00 | 3.74 GiB | 32.82 GiB |
| Q4_K_M | 4.89 | 4.58 GiB | 40.2 GiB |
| Q5_K_M | 5.70 | 5.33 GiB | 46.85 GiB |
| Q6_K | 6.56 | 6.14 GiB | 53.91 GiB |
| Q8_0 | 8.50 | 7.95 GiB | 69.82 GiB |
| F16 / BF16 (unquantized) | 16.00 | 14.96 GiB | 131.42 GiB |

So Q4_K_M costs 4.89 bits per weight, not 4. Other architectures land within a few percent of these figures, because the mixes are decided tensor by tensor. AWQ and GPTQ 4-bit checkpoints (group size 128) and FP8 are sized like the VRAM calculator sizes them: the quantized weights plus their scales, with the embeddings and LM head left at 16-bit. gpt-oss is published in MXFP4, 4.25 bits per weight for the expert weights. The format table under the calculator has the basis and source for each.

The KV cache has its own precision. llama.cpp's default is F16. Its `q8_0` cache type takes 34 bytes per 32 values and `q4_0` 18 bytes per 32, and a quantized V cache needs flash attention, which current llama.cpp turns on by default (`-fa auto`). Whether a quantized cache hurts answers depends on the model; check it on your own prompts.

## Partial offload: why the speed falls off a cliff

When a model is bigger than VRAM, llama.cpp can keep some layers on the GPU and run the rest on the CPU from system RAM. That works, but every generated token still has to read every layer once, and system RAM is far slower than VRAM. Two-channel desktop memory peaks at 89.6 GB/s with DDR5-5600, against 1008 GB/s on an RTX 4090.

Llama 3.3 70B at Q4_K_M needs 43.7 GiB at an 8k context. On one 24 GB card, 43 of its 80 layers fit. The other 37 run from system RAM, and reading them takes 90% of each token's time. The ceiling moves with the RAM, not the GPU:

| System RAM | Peak bandwidth | Ceiling, Llama 3.3 70B Q4_K_M, 8k context |
|---|---|---|
| DDR4-3200, two channels | 51.2 GB/s | 2.3 tokens/s |
| DDR5-4800, two channels | 76.8 GB/s | 3.4 tokens/s |
| DDR5-5600, two channels | 89.6 GB/s | 3.9 tokens/s |
| DDR5-6400, two channels | 102.4 GB/s | 4.4 tokens/s |
| DDR5-7200, two channels | 115.2 GB/s | 4.9 tokens/s |

With two 24 GB cards the same model fits entirely in VRAM (44.7 GiB of 48), and the ceiling on two RTX 3090s rises to 20.7 tokens per second. A second card adds memory, not bandwidth: llama.cpp's default layer split passes each token through the cards one after another, so the ceiling is that of one card reading all the bytes. These figures are peaks. Crucial's table of effective DDR5-5600 bandwidth gives 69.21 GB/s, about 77% of the peak, and inference engines lose more on top.

Mixture-of-experts models are the exception worth knowing. Each token reads only the active experts, so a model like gpt-oss-20b or Qwen3 30B-A3B stays usable with part of it in RAM, and llama.cpp can keep just the expert weights in RAM (`--n-cpu-moe` or `--override-tensor`) while attention stays on the GPU. That usually beats offloading whole layers. The calculator models whole-layer offload only, so for MoE models it is the cautious estimate.

## Unified memory: Macs, Ryzen AI Max and DGX Spark

On a Mac, a Ryzen AI Max mini-PC or a DGX Spark, the CPU and GPU share one pool of memory, so there is no separate VRAM to overflow. The question becomes how much of that pool the GPU may use.


- **Apple Silicon.** macOS lets the GPU use only part of unified memory by default, and Apple documents no fixed rule. Metal reports the limit at run time as recommendedMaxWorkingSetSize, "an approximation of how much memory, in bytes, this GPU device can allocate without affecting its runtime performance", and llama.cpp and Ollama read it as the Mac's GPU memory. Apple's 2021 tech talk on Metal compute gives two examples: an M1 Pro or M1 Max with 32 GB lets the GPU access 21 GB, and an M1 Max with 64 GB, 48 GB. The calculator follows those examples: two thirds of memory up to 32 GB, three quarters above; enter your own share if your Mac reports a different limit. Apple's MLX project documents raising the limit with sudo sysctl iogpu.wired_limit_mb=<megabytes>, to a value larger than the model but smaller than the machine's memory. Community reports say the setting does not survive a reboot, and an MLX maintainer linked a kernel panic to a limit set too high, so leave macOS several gigabytes.
- **AMD Ryzen AI Max.** AMD says up to 96 GB of a 128 GB Ryzen AI Max system can be assigned to graphics with Variable Graphics Memory, and up to 160 GB of 192 GB on the Ryzen AI Max+ PRO 495; the calculator uses those shares (75% and 83%). On Linux the GPU can also map system memory through GTT, which AMD's ROCm guide says defaults to about half of RAM and can be raised, so the usable share depends on how the machine is set up. Enter your own share if you have changed it.
- **NVIDIA DGX Spark and RTX Spark PCs.** NVIDIA documents no separate GPU limit for the DGX Spark's unified LPDDR5X (128 GB, and 64 GB in versions sold through other makers), so the calculator lets the GPU use all of it except the RAM kept for the operating system. NVIDIA publishes no memory bandwidth for RTX Spark PCs, so the calculator shows no speed ceiling for them.

When a model fits in memory but not in the GPU's share, the calculator says "over the default GPU limit" instead of "partial offload", and shows the ceiling you would get after raising the limit.

## A worked example

Llama 3.1 8B at Q4_K_M with an 8k context on an RTX 4090:

- **Weights:** 8,030,261,248 parameters x 4.8944 bits / 8 = 4.58 GiB.
- **KV cache:** 2 x 32 layers x 8 KV heads x 128 x 2 bytes = 128 KiB per token, x 8,192 tokens = 1 GiB.
- **Total with the 1 GiB allowance:** 6.58 GiB, comfortably inside 24 GB.
- **Ceiling:** each token reads the weights except the embedding table, plus the cache, about 5.67 GB. 1008 GB/s divided by that is **178 tokens per second**, at most, for one conversation at a full 8k context.

That is arithmetic, not a measurement. Real engines stay below the bandwidth ceiling, and how far below depends on the engine, the kernels and the GPU. Prompt processing (prefill) is limited by compute rather than bandwidth and is not estimated here. Measured speeds on real hardware are [planned](/methodology/#benchmarks-planned-no-results-yet), with the scripts and raw results to be published alongside them.

## What the calculator does not model

- Prompt-processing speed, batching several conversations at once, and speculative decoding.
- Vision encoders' activations for image input (their weights are in the multimodal presets' parameter counts).
- Expert-only offload for MoE models, and tensor (row) split across GPUs.
- Memory the operating system or a display takes from VRAM beyond the runtime allowance. If your card also drives your monitors, raise the allowance.
- Quality. Smaller quantizations lose some accuracy, by an amount that depends on the model and the task; llama.cpp publishes perplexity changes for Llama 3.1 8B, and nothing here measures it for other models.

## Where the numbers come from

Memory sizes, bandwidth and power are read from each vendor's own spec pages, datasheets and whitepapers, linked row by row in the hardware table above. Where a vendor does not publish a value, the table says "n/a" and the calculator shows no ceiling rather than a guess: NVIDIA publishes no memory bandwidth for the RTX 3080 12 GB or for RTX Spark PCs. Model shapes come from Hugging Face `config.json` files, and the bits-per-weight figures from llama.cpp's repository. The [methodology](/methodology/) lists every source and assumption, and the [VRAM guide](/how-much-vram-to-run-llms-locally/) applies the calculator to the models people ask about most.
