Best local AI models for Apple M1 Max
22.4 GB usable of 32 GB unified memory. At a 4k context, 168 of the 233 models in our catalog with verified parameter counts fit fully, up to Qwen3-30B-A3B at 30B parameters. Computed for the 32 GB configuration; a larger memory configuration fits more.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 168 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Qwen3-30B-A3B | 30B | Q4_K_M | 22 GB |
| Qwen3-Coder 30B-A3B | 30B | Q4_K_M | 22 GB |
| Gemma 3 27B | 27B | Q4_K_M | 19.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | Q4_K_M | 19.8 GB |
| Wan 2.2 / 2.5 | 27B | Q4_K_M | 19.8 GB |
| Gemma 4 26B-A4B | 26B | Q5_K_M | 22.2 GB |
| Gemma 4 (all sizes) | 26B | Q5_K_M | 22.2 GB |
| Aria | 25B | Q5_K_M | 21.3 GB |
| Mistral Small 3.2 | 24B | Q5_K_M | 20.4 GB |
| Magistral Small | 24B | Q5_K_M | 20.4 GB |
| Devstral Small 1.1 | 24B | Q5_K_M | 20.4 GB |
| Solar Pro | 22B | Q6_K | 21.6 GB |
| Codestral 22B | 22B | Q6_K | 21.6 GB |
| gpt-oss-20b | 21B | Q6_K | 20.7 GB |
| Reka Flash 3 | 21B | Q6_K | 20.7 GB |
| Qwen-Image | 20B | Q6_K | 19.7 GB |
| Qwen-Image-Edit | 20B | Q6_K | 19.7 GB |
| CogVLM2 | 19B | Q6_K | 18.7 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q8_0 | 21.6 GB |
| Ling-Coder-Lite | 16.8B | Q8_0 | 21.4 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q8_0 | 20.4 GB |
| Kimi-VL A3B | 16B | Q8_0 | 20.4 GB |
| Apriel-1.5-15B-Thinker | 15B | Q8_0 | 19.1 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q8_0 | 19.1 GB |
| Qwen2.5 14B | 14.7B | Q8_0 | 19.5 GB |
| Phi-3 Medium | 14B | Q8_0 | 17.8 GB |
| Phi-4 | 14B | Q8_0 | 17.8 GB |
| Phi-4-reasoning / -plus | 14B | Q8_0 | 17.8 GB |
| Wan 2.2 T2I | 14B | Q8_0 | 17.8 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q8_0 | 17.8 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Qwen3 8B / 14B / 32B | 32B | 23.4 GB needed | 25.4 GB |
| Qwen3.5 (dense variants) | 32B | 23.4 GB needed | 25.4 GB |
| Aya Expanse 8B / 32B | 32B | 23.4 GB needed | 25.4 GB |
| Granite 4.0 Small/Tiny | 32B | 23.4 GB needed | 25.4 GB |
| Qwen2.5-Coder 0.5B to 32B | 32B | 23.4 GB needed | 25.4 GB |
| OTel 2.0 LLM 31B IT | 32.1B | 27.5 GB needed | 29.5 GB |
| DeepSeek-Coder 1.3B / 6.7B / 33B | 33B | 24.2 GB needed | 26.2 GB |
| WizardCoder 33B | 33B | 24.2 GB needed | 26.2 GB |
| Yi 1.5 9B / 34B | 34B | 24.9 GB needed | 26.9 GB |
| Granite Code 3B to 34B | 34B | 24.9 GB needed | 26.9 GB |
How to read this
The Apple M1 Max chip with a 32 GB unified memory pool provides 22.4 GB of usable memory for local AI models. This usable share is the actual space available for running models after the macOS operating system and basic system tasks take their portion. Keeping your model size within this 22.4 GB limit ensures that the model runs entirely on the fast graphics processor of your Mac.
The quantization column shows the best compression level for each model. Quantization reduces the size of a model so it can fit into your memory. For example Qwen3-30B-A3B and Qwen3-Coder 30B-A3B fit at the Q4_K_M quantization level using 22 GB of memory. Gemma 3 27B and Wan 2.2 / 2.5 fit at Q4_K_M using 19.8 GB of memory. Gemma 4 26B-A4B and Gemma 4 (all sizes) use 22.2 GB of memory at the Q5_K_M quantization level.
Smaller models can run at higher quality levels because they use less memory. Mistral Small 3.2 and Magistral Small fit at Q5_K_M using 20.4 GB of memory. Solar Pro and Codestral 22B run at Q6_K using 21.6 GB of memory. You can run Qwen2.5 14B at the Q8_0 quantization level using 19.5 GB of memory. Phi-4 and Wan 2.2 T2I also run at the Q8_0 quantization level using 17.8 GB of memory.
If a model exceeds the 22.4 GB limit you must offload parts of it to the system CPU. This offloading allows you to run larger models but it reduces your generation speed. For example Qwen3.5 (dense variants) and Qwen2.5-Coder 0.5B to 32B require 23.4 GB of memory at Q4_K_M which needs 25.4 GB of system RAM. DeepSeek-Coder 1.3B / 6.7B / 33B requires 24.2 GB of memory at Q4_K_M which needs 26.2 GB of system RAM.
Other offload cases include Yi 1.5 9B / 34B and Granite Code 3B to 34B which require 24.9 GB of memory at Q4_K_M and need 26.9 GB of system RAM. OTel 2.0 LLM 31B IT requires 27.5 GB of memory at Q4_K_M and needs 29.5 GB of system RAM. These models will load on a 32 GB system but the CPU processing will make them slower than models that fit entirely in the graphics memory.
You must also consider the context window when planning your memory use. The memory figures listed here assume a standard 4k context window. If you increase the context window to process longer documents the memory usage will grow. A larger context window might force you to use a lower quantization level or a smaller model to avoid slow CPU offloading.