Local AI
Run AI at home: will the model fit, how fast can it go, and what will it cost?
Free calculators and guides for running language models on your own GPU, Mac or mini-PC. Whether a model fits, how many layers spill into system RAM, the speed ceiling your memory bandwidth allows, and what the electricity costs against an API. Built from vendor specifications and the models' own configs, with nothing presented as a benchmark.
Three questions decide whether a model runs well at home
- Does it fit? The weights at your quantization, plus a KV cache that grows with your context length, plus the software's buffers, have to fit in GPU memory. If they do not, llama.cpp can run the rest of the layers from system RAM, at a large cost in speed. On a Mac or a unified-memory mini-PC, the GPU may use only part of the memory by default.
- How fast can it go? Every generated token reads all the active weights and the whole cache once, so memory bandwidth divided by those bytes is the most tokens per second any software can reach. That ceiling is arithmetic, and real speed is lower.
- What does it cost? Electricity at your state's price, the hardware spread over the months you will use it, against an API serving the same open model.
The calculators
- Can my computer run this model? Pick one of 94 graphics cards, Macs and mini-PCs (or enter your own memory and bandwidth), a model and a GGUF, AWQ, FP8 or MXFP4 quantization. It answers fits, partial offload with the number of layers on the GPU, or no, and gives the speed ceiling, with the arithmetic shown.
- Electricity cost of running AI at home: rated or measured watts, hours and utilization, and the EIA's average residential price for your state (July 2026 data), turned into kWh, dollars a month and dollars per million tokens, next to an API's list price for the same tokens.
The guides
- How much VRAM do you need to run Llama, Qwen or gpt-oss locally? Memory for 14 popular open models at every common quantization and context length, and the smallest GPU or Mac memory that holds each.
- Best GPU for running LLMs at home, at each budget tier: cards grouped by memory size and ranked by bandwidth, with what each tier can run.
- Apple Silicon vs NVIDIA for local AI: capacity against bandwidth, for every M1 to M6 chip against the RTX 4090, RTX 5090, RTX PRO 6000, DGX Spark and Ryzen AI Max.
What fits in common memory sizes
The largest models from the VRAM guide that fit at Q4_K_M with an 8k context and a 1 GiB allowance (gpt-oss in its published MXFP4). On a Mac, use the memory the GPU may use, about three quarters of the total, not the total itself.
| Memory | Fits at Q4_K_M, 8k context |
|---|---|
| 8 GB | Qwen3 8B, Llama 3.1 8B Instruct |
| 12 GB | Qwen3 14B, Qwen3 8B, Llama 3.1 8B Instruct |
| 16 GB | Mistral Small 3.2 24B Instruct (2506), gpt-oss-20b, Qwen3 14B, and 2 smaller |
| 24 GB | Qwen3 32B, Gemma 4 31B IT, Qwen3 30B-A3B Instruct (2507), and 6 smaller |
| 32 GB | Qwen3 32B, Gemma 4 31B IT, Qwen3 30B-A3B Instruct (2507), and 6 smaller |
| 48 GB | Qwen2.5 72B Instruct, Llama 3.3 70B Instruct, Qwen3 32B, and 8 smaller |
| 64 GB | gpt-oss-120b, Qwen2.5 72B Instruct, Llama 3.3 70B Instruct, and 9 smaller |
| 96 GB | GLM-4.5-Air, gpt-oss-120b, Qwen2.5 72B Instruct, and 10 smaller |
| 128 GB | GLM-4.5-Air, gpt-oss-120b, Qwen2.5 72B Instruct, and 10 smaller |
| 192 GB | Qwen3 235B-A22B Instruct (2507), GLM-4.5-Air, gpt-oss-120b, and 11 smaller |
| 256 GB | Qwen3 235B-A22B Instruct (2507), GLM-4.5-Air, gpt-oss-120b, and 11 smaller |
| 512 GB | Qwen3 235B-A22B Instruct (2507), GLM-4.5-Air, gpt-oss-120b, and 11 smaller |
How this section works
Specifications from the vendors. Memory sizes, memory bandwidth and rated power come from NVIDIA's, AMD's, Intel's and Apple's own spec pages, datasheets and whitepapers, with the link on every row of the hardware table. When a vendor does not publish a figure, the table says so and the calculator does not guess.
The models' own shapes. Layer counts, KV heads and head sizes come from each model's config.json, the same data as the VRAM calculator, and quantization sizes from llama.cpp's own tables.
Ceilings, not benchmarks. The speeds on these pages are upper bounds from memory bandwidth. Measured tokens per second on real home hardware are planned, and will be published with their scripts and raw results.
No prices from retailers. Hardware prices change too often to print. The electricity calculator takes what you paid, and the guides group hardware by what it can do. Some pages link to retailers; any paid link is labeled "(paid link)", and the affiliate disclosure lists every relationship.
Also available as Markdown.