The trade in one paragraph

Generating text is limited by memory: how much there is, which decides what fits, and how fast it is, which decides how many tokens per second are possible. An Apple Silicon Mac shares one large pool of memory between CPU and GPU, up to 512 GB on the M5 Ultra and M3 Ultra. An NVIDIA GeForce card has at most 32 GB (RTX 5090), but its memory is faster: 1,792 GB/s, against 1,200 GB/s on the fastest Mac. So for a model that fits on the card, the card has the higher ceiling, and for a model that does not, the Mac can run it at all.

Every speed below is the ceiling from the Can I run it? calculator: memory bandwidth divided by the bytes read per token, for one conversation at an 8,192-token context. Real software stays below it on both platforms. Nothing here is a benchmark.

Every Apple chip at its largest memory

Each Mac is shown at the largest memory Apple offers for that chip, with the GPU's default share of it (see the limit section below). Ceilings in tokens per second:

Chip Memory (GPU share) Bandwidth Llama 3.1 8B Q4_K_M Qwen3 32B Q4_K_M Llama 3.3 70B Q4_K_M gpt-oss-120b MXFP4 Qwen3 235B-A22B Q4_K_M
M5 Ultra 512 GB (75%) 1200 GB/s 212 55 26 380 80
M3 Ultra 512 GB (75%) 819 GB/s 145 38 18 260 55
M1 Ultra 128 GB (75%) 800 GB/s 141 37 18 254 no
M2 Ultra 192 GB (75%) 800 GB/s 141 37 18 254 53
M5 Max (18-core CPU, 40-core GPU) 128 GB (75%) 614 GB/s 108 28 14 195 no
M4 Max (16-core CPU, 40-core GPU) 128 GB (75%) 546 GB/s 96 25 12 173 no
M5 Max (18-core CPU, 32-core GPU) 36 GB (75%) 460 GB/s 81 21 no no no
M4 Max (14-core CPU, 32-core GPU) 36 GB (75%) 410 GB/s 72 19 no no no
M1 Max 64 GB (75%) 400 GB/s 71 18 9 no no
M2 Max 96 GB (75%) 400 GB/s 71 18 9 127 no
M3 Max (16-core CPU, 40-core GPU) 128 GB (75%) 400 GB/s 71 18 9 127 no
M5 Pro 64 GB (75%) 307 GB/s 54 14 7 no no
M3 Max (14-core CPU, 30-core GPU) 96 GB (75%) 300 GB/s 53 14 7 95 no
M4 Pro 64 GB (75%) 273 GB/s 48 13 6 no no
M1 Pro 32 GB (67%) 200 GB/s 35 9 no no no
M2 Pro 32 GB (67%) 200 GB/s 35 9 no no no
M6 32 GB (67%) 170 GB/s 30 8 no no no
M5 32 GB (67%) 153 GB/s 27 7 no no no
M3 Pro 36 GB (75%) 150 GB/s 26 7 no no no
M4 32 GB (67%) 120 GB/s 21 6 no no no

Bandwidth within a chip family depends on the version: the M4 Max with a 32-core GPU has 410 GB/s and only 36 GB, while the 40-core version has 546 GB/s and up to 128 GB. The Pro chips have far less bandwidth than the Max chips of the same generation.

The same models on NVIDIA and the other unified-memory machines

Hardware Memory Bandwidth Llama 3.1 8B Q4_K_M Qwen3 32B Q4_K_M Llama 3.3 70B Q4_K_M gpt-oss-120b MXFP4 Qwen3 235B-A22B Q4_K_M
NVIDIA GeForce RTX 4090 24 GB 1,008 GB/s 178 46 offload, 3.9 no no
NVIDIA GeForce RTX 5090 32 GB 1,792 GB/s 316 82 offload, 6.4 no no
NVIDIA RTX PRO 6000 Blackwell Workstation Edition 96 GB 1,792 GB/s 316 82 40 568 no
NVIDIA DGX Spark 128 GB unified (97%) 273 GB/s 48 13 6 86 no
AMD Ryzen AI Max+ 395 (Radeon 8060S) 128 GB unified (75%) 256 GB/s 45 12 6 81 no
AMD Ryzen AI Max+ PRO 495 (Radeon 8065S) 192 GB unified (83%) 273 GB/s 48 13 6 87 18

For the graphics cards, "offload" assumes 32 GB of DDR5-5600 system RAM. A desktop with more RAM runs the bigger models too, slowly. With 128 GB of DDR5-5600 next to an RTX 4090, the ceilings are 3.9 tokens per second for Llama 3.3 70B (37 of its 80 layers in RAM), 43.6 for gpt-oss-120b, which as a mixture-of-experts model reads only a few experts per token, and 7 for Qwen3 235B-A22B. llama.cpp's option to keep only the expert weights in RAM (--n-cpu-moe) can do better than the whole-layer split the calculator assumes.

What the tables say

Prompt processing is a different question

Reading a long prompt before the first token appears (prefill) is limited by the GPU's compute, not by memory bandwidth, so the ceilings above say nothing about it. Long documents, big codebases and retrieval-augmented prompts are where it matters. This site has not measured prefill on Macs or NVIDIA cards yet and makes no claim about it; measured benchmarks are planned, with scripts and raw results.

Software

Power

Apple publishes a maximum power figure for each desktop Mac, for the whole computer at the wall while running a compute-intensive test: 140 W for the Mac mini (2024) with M4 Pro, 200 W for the Mac Studio (M5 Max, 2026) with M5 Max (18-core CPU, 32-core GPU), 270 W for the Mac Studio (2025) with M3 Ultra, 385 W for the Mac Studio (M5 Ultra, 2026) with M5 Ultra. NVIDIA's Total Graphics Power is for the card alone: 450 W for the RTX 4090 and 575 W for the RTX 5090, before the rest of the PC. Both are maximums, not what a machine draws while generating text. The electricity calculator turns either into a monthly cost; at its defaults, a Mac Studio with M3 Ultra comes to $3.01 a month and a PC with an RTX 4090 to $9.47, both upper bounds.

The GPU memory limit on Macs

macOS does not let the GPU use all of unified memory. Metal reports the limit as recommendedMaxWorkingSetSize, and llama.cpp prints it when it starts. Apple documents no formula, but its 2021 tech talk on Metal compute gives two examples, 21 GB of 32 GB and 48 GB of 64 GB, and the calculator follows them: two thirds of memory up to 32 GB, three quarters above. A 128 GB Mac therefore has about 96 GB for a model by default.

The limit can be raised with sudo sysctl iogpu.wired_limit_mb=<megabytes>, which Apple's MLX project documents: set it above the model's size and below the machine's memory. Community reports say the setting resets on reboot, and an MLX maintainer linked a kernel panic to a limit set too high, so leave macOS several gigabytes. The calculator lets you enter your own share and marks models that fit only after raising it.

Which to choose

Also available as Markdown.