# Apple Silicon vs NVIDIA for running LLMs at home: memory against bandwidth

> A Mac holds far larger models than any consumer graphics card, and a high-end NVIDIA card generates faster on the models it can hold. This guide works out where the line falls for every M1 to M6 chip against the RTX 4090, RTX 5090, RTX PRO 6000, DGX Spark and Ryzen AI Max, from vendor specs and the calculator's arithmetic.

By GPUCostLab. Updated 7 October 2026. Canonical URL: https://gpucostlab.com/apple-silicon-vs-nvidia-for-local-ai/

## The trade in one paragraph

Generating text is limited by memory: how much there is, which decides what fits, and how fast it is, which decides how many tokens per second are possible. An Apple Silicon Mac shares one large pool of memory between CPU and GPU, up to 512 GB on the M5 Ultra and M3 Ultra. An NVIDIA GeForce card has at most 32 GB (RTX 5090), but its memory is faster: 1,792 GB/s, against 1,200 GB/s on the fastest Mac. So for a model that fits on the card, the card has the higher ceiling, and for a model that does not, the Mac can run it at all.

Every speed below is the ceiling from the [Can I run it?](/can-i-run-this-llm/) calculator: memory bandwidth divided by the bytes read per token, for one conversation at an 8,192-token context. Real software stays below it on both platforms. Nothing here is a benchmark.

## Every Apple chip at its largest memory

Each Mac is shown at the largest memory Apple offers for that chip, with the GPU's default share of it (see the limit section below). Ceilings in tokens per second:

| Chip | Memory (GPU share) | Bandwidth | Llama 3.1 8B Q4_K_M | Qwen3 32B Q4_K_M | Llama 3.3 70B Q4_K_M | gpt-oss-120b MXFP4 | Qwen3 235B-A22B Q4_K_M |
|---|---|---|---|---|---|---|---|
| M5 Ultra | 512 GB (75%) | 1200 GB/s | 212 | 55 | 26 | 380 | 80 |
| M3 Ultra | 512 GB (75%) | 819 GB/s | 145 | 38 | 18 | 260 | 55 |
| M1 Ultra | 128 GB (75%) | 800 GB/s | 141 | 37 | 18 | 254 | no |
| M2 Ultra | 192 GB (75%) | 800 GB/s | 141 | 37 | 18 | 254 | 53 |
| M5 Max (18-core CPU, 40-core GPU) | 128 GB (75%) | 614 GB/s | 108 | 28 | 14 | 195 | no |
| M4 Max (16-core CPU, 40-core GPU) | 128 GB (75%) | 546 GB/s | 96 | 25 | 12 | 173 | no |
| M5 Max (18-core CPU, 32-core GPU) | 36 GB (75%) | 460 GB/s | 81 | 21 | no | no | no |
| M4 Max (14-core CPU, 32-core GPU) | 36 GB (75%) | 410 GB/s | 72 | 19 | no | no | no |
| M1 Max | 64 GB (75%) | 400 GB/s | 71 | 18 | 9 | no | no |
| M2 Max | 96 GB (75%) | 400 GB/s | 71 | 18 | 9 | 127 | no |
| M3 Max (16-core CPU, 40-core GPU) | 128 GB (75%) | 400 GB/s | 71 | 18 | 9 | 127 | no |
| M5 Pro | 64 GB (75%) | 307 GB/s | 54 | 14 | 7 | no | no |
| M3 Max (14-core CPU, 30-core GPU) | 96 GB (75%) | 300 GB/s | 53 | 14 | 7 | 95 | no |
| M4 Pro | 64 GB (75%) | 273 GB/s | 48 | 13 | 6 | no | no |
| M1 Pro | 32 GB (67%) | 200 GB/s | 35 | 9 | no | no | no |
| M2 Pro | 32 GB (67%) | 200 GB/s | 35 | 9 | no | no | no |
| M6 | 32 GB (67%) | 170 GB/s | 30 | 8 | no | no | no |
| M5 | 32 GB (67%) | 153 GB/s | 27 | 7 | no | no | no |
| M3 Pro | 36 GB (75%) | 150 GB/s | 26 | 7 | no | no | no |
| M4 | 32 GB (67%) | 120 GB/s | 21 | 6 | no | no | no |

Bandwidth within a chip family depends on the version: the M4 Max with a 32-core GPU has 410 GB/s and only 36 GB, while the 40-core version has 546 GB/s and up to 128 GB. The Pro chips have far less bandwidth than the Max chips of the same generation.

## The same models on NVIDIA and the other unified-memory machines

| Hardware | Memory | Bandwidth | Llama 3.1 8B Q4_K_M | Qwen3 32B Q4_K_M | Llama 3.3 70B Q4_K_M | gpt-oss-120b MXFP4 | Qwen3 235B-A22B Q4_K_M |
|---|---|---|---|---|---|---|---|
| NVIDIA GeForce RTX 4090 | 24 GB | 1,008 GB/s | 178 | 46 | offload, 3.9 | no | no |
| NVIDIA GeForce RTX 5090 | 32 GB | 1,792 GB/s | 316 | 82 | offload, 6.4 | no | no |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 96 GB | 1,792 GB/s | 316 | 82 | 40 | 568 | no |
| NVIDIA DGX Spark | 128 GB unified (97%) | 273 GB/s | 48 | 13 | 6 | 86 | no |
| AMD Ryzen AI Max+ 395 (Radeon 8060S) | 128 GB unified (75%) | 256 GB/s | 45 | 12 | 6 | 81 | no |
| AMD Ryzen AI Max+ PRO 495 (Radeon 8065S) | 192 GB unified (83%) | 273 GB/s | 48 | 13 | 6 | 87 | 18 |

For the graphics cards, "offload" assumes 32 GB of DDR5-5600 system RAM. A desktop with more RAM runs the bigger models too, slowly. With 128 GB of DDR5-5600 next to an RTX 4090, the ceilings are 3.9 tokens per second for Llama 3.3 70B (37 of its 80 layers in RAM), 43.6 for gpt-oss-120b, which as a mixture-of-experts model reads only a few experts per token, and 7 for Qwen3 235B-A22B. llama.cpp's option to keep only the expert weights in RAM (`--n-cpu-moe`) can do better than the whole-layer split the calculator assumes.

## What the tables say

- **For models that fit on the card, the RTX 5090 has the highest ceiling of any machine here.** For Llama 3.1 8B it is 316 tokens per second. The M5 Ultra reaches 212, more than the RTX 4090's 178; an M4 Max with a 40-core GPU reaches 96.
- **At 70B, the Mac wins unless you buy a workstation card.** Llama 3.3 70B at Q4_K_M does not fit on any GeForce card, and offloading leaves an RTX 4090 at 3.9 tokens per second. A 128 GB M4 Max holds it all, with a ceiling of 12.1, and an M5 Max 13.6. The 96 GB RTX PRO 6000 holds it at 39.6, or two 24 GB GeForce cards (see the [GPU tiers guide](/best-gpu-for-local-llms/)).
- **Above 128 GB, few machines hold the model at all.** Qwen3 235B-A22B at Q4_K_M needs 136.42 GiB at 8k. Only these hold it entirely in GPU-usable memory: Apple M2 Ultra, Apple M3 Ultra, Apple M5 Ultra, AMD Ryzen AI Max+ PRO 495 (Radeon 8065S).
- **The DGX Spark and Ryzen AI Max sit between.** Both have 128 GB, like an M4 Max or M5 Max, at about half the bandwidth (273 and 256 GB/s), so their ceilings are half as high or less on the same models.

## Prompt processing is a different question

Reading a long prompt before the first token appears (prefill) is limited by the GPU's compute, not by memory bandwidth, so the ceilings above say nothing about it. Long documents, big codebases and retrieval-augmented prompts are where it matters. This site has not measured prefill on Macs or NVIDIA cards yet and makes no claim about it; [measured benchmarks are planned](/methodology/#benchmarks-planned-no-results-yet), with scripts and raw results.

## Software

- **NVIDIA** cards use CUDA, which nearly every inference engine supports: llama.cpp, Ollama, LM Studio, vLLM, SGLang and TensorRT-LLM among them. FP8 and NVFP4 checkpoints need recent NVIDIA tensor cores to run at full speed.
- **Macs** use Metal through llama.cpp (and the tools built on it, such as Ollama and LM Studio) and through Apple's own MLX framework. vLLM-style serving engines target CUDA first. For one person chatting with a model, llama.cpp and MLX cover most needs.

## Power

Apple publishes a maximum power figure for each desktop Mac, for the whole computer at the wall while running a compute-intensive test: 140 W for the Mac mini (2024) with M4 Pro, 200 W for the Mac Studio (M5 Max, 2026) with M5 Max (18-core CPU, 32-core GPU), 270 W for the Mac Studio (2025) with M3 Ultra, 385 W for the Mac Studio (M5 Ultra, 2026) with M5 Ultra. NVIDIA's Total Graphics Power is for the card alone: 450 W for the RTX 4090 and 575 W for the RTX 5090, before the rest of the PC. Both are maximums, not what a machine draws while generating text. The [electricity calculator](/local-llm-electricity-cost/) turns either into a monthly cost; at its defaults, a Mac Studio with M3 Ultra comes to $3.01 a month and a PC with an RTX 4090 to $9.47, both upper bounds.

## The GPU memory limit on Macs

macOS does not let the GPU use all of unified memory. Metal reports the limit as `recommendedMaxWorkingSetSize`, and llama.cpp prints it when it starts. Apple documents no formula, but its 2021 tech talk on Metal compute gives two examples, 21 GB of 32 GB and 48 GB of 64 GB, and the calculator follows them: two thirds of memory up to 32 GB, three quarters above. A 128 GB Mac therefore has about 96 GB for a model by default.

The limit can be raised with `sudo sysctl iogpu.wired_limit_mb=<megabytes>`, which Apple's MLX project documents: set it above the model's size and below the machine's memory. Community reports say the setting resets on reboot, and an MLX maintainer linked a kernel panic to a limit set too high, so leave macOS several gigabytes. The [calculator](/can-i-run-this-llm/) lets you enter your own share and marks models that fit only after raising it.

## Which to choose

- **Models up to about 30B parameters, at the highest speed:** the RTX 5090, with a ceiling of 82 tokens per second for Qwen3 32B at Q4_K_M. A 24 GB card such as the RTX 4090 holds the same model (21.67 GiB at 8k) with a ceiling of 46, against 38 on an M3 Ultra and 55 on an M5 Ultra.
- **70B dense models or 100B+ mixture-of-experts models on one machine:** a Max or Ultra Mac with 96 to 128 GB or more, or a 128 GB DGX Spark or Ryzen AI Max PC at lower bandwidth.
- **200B-class models:** an M2, M3 or M5 Ultra with 192 GB or more, or a desktop with a large GPU plus lots of system RAM, which runs them at a fraction of the speed.
- **Training or fine-tuning:** NVIDIA, for the software support. The [fine-tuning cost estimator](/llm-fine-tuning-cost/) sizes the job and prices it on rented GPUs, which avoids buying hardware for a one-off run.
