Best local AI models for Apple M2 Max

22.4 GB usable of 32 GB unified memory. At a 4k context, 168 of the 233 models in our catalog with verified parameter counts fit fully, up to Qwen3-30B-A3B at 30B parameters. Computed for the 32 GB configuration; a larger memory configuration fits more.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 168 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Qwen3-30B-A3B30BQ4_K_M22 GB
Qwen3-Coder 30B-A3B30BQ4_K_M22 GB
Gemma 3 27B27BQ4_K_M19.8 GB
Gemma 3 4B/12B/27B (vision)27BQ4_K_M19.8 GB
Wan 2.2 / 2.527BQ4_K_M19.8 GB
Gemma 4 26B-A4B26BQ5_K_M22.2 GB
Gemma 4 (all sizes)26BQ5_K_M22.2 GB
Aria25BQ5_K_M21.3 GB
Mistral Small 3.224BQ5_K_M20.4 GB
Magistral Small24BQ5_K_M20.4 GB
Devstral Small 1.124BQ5_K_M20.4 GB
Solar Pro22BQ6_K21.6 GB
Codestral 22B22BQ6_K21.6 GB
gpt-oss-20b21BQ6_K20.7 GB
Reka Flash 321BQ6_K20.7 GB
Qwen-Image20BQ6_K19.7 GB
Qwen-Image-Edit20BQ6_K19.7 GB
CogVLM219BQ6_K18.7 GB
HunyuanImage 2.1 / 3.017BQ8_021.6 GB
Ling-Coder-Lite16.8BQ8_021.4 GB
DeepSeek-Coder-V2 16B / 236B16BQ8_020.4 GB
Kimi-VL A3B16BQ8_020.4 GB
Apriel-1.5-15B-Thinker15BQ8_019.1 GB
StarCoder2 3B / 7B / 15B15BQ8_019.1 GB
Qwen2.5 14B14.7BQ8_019.5 GB
Phi-3 Medium14BQ8_017.8 GB
Phi-414BQ8_017.8 GB
Phi-4-reasoning / -plus14BQ8_017.8 GB
Wan 2.2 T2I14BQ8_017.8 GB
Wan 2.1 (1.3B / 14B)14BQ8_017.8 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at Q4_K_MSystem RAM at 4k
Qwen3 8B / 14B / 32B32B23.4 GB needed25.4 GB
Qwen3.5 (dense variants)32B23.4 GB needed25.4 GB
Aya Expanse 8B / 32B32B23.4 GB needed25.4 GB
Granite 4.0 Small/Tiny32B23.4 GB needed25.4 GB
Qwen2.5-Coder 0.5B to 32B32B23.4 GB needed25.4 GB
OTel 2.0 LLM 31B IT32.1B27.5 GB needed29.5 GB
DeepSeek-Coder 1.3B / 6.7B / 33B33B24.2 GB needed26.2 GB
WizardCoder 33B33B24.2 GB needed26.2 GB
Yi 1.5 9B / 34B34B24.9 GB needed26.9 GB
Granite Code 3B to 34B34B24.9 GB needed26.9 GB

How to read this

The Apple M2 Max chip with a 32 GB unified memory pool provides 22.4 GB of usable memory for local AI models. This limit is critical because macOS reserves the remaining memory for the system and display. To run a model entirely on the fast graphics hardware of your Mac, the model files must fit within this 22.4 GB limit. Models that fit within this space run quickly and efficiently.

The quantization column shows the compression level used to fit these models into memory. Quants like Q4_K_M represent four bit quantization which allows larger models to fit but with a small loss in precision. Quants like Q5_K_M and Q6_K offer higher precision for medium models. The Q8_0 quant provides the highest precision for smaller models like the 14B and 15B variants because they easily fit inside the 22.4 GB budget.

The largest models you can run entirely in unified memory include Qwen3-30B-A3B and Qwen3-Coder 30B-A3B at Q4_K_M which use 22 GB. Gemma 4 26B-A4B and Gemma 4 (all sizes) fit at Q5_K_M using 22.2 GB. Gemma 3 27B and Wan 2.2 / 2.5 use 19.8 GB at Q4_K_M. Aria uses 21.3 GB at Q5_K_M. Mistral Small 3.2 and Devstral Small 1.1 use 20.4 GB at Q5_K_M. Solar Pro and Codestral 22B fit at Q6_K using 21.6 GB.

Other models that fit completely in memory include gpt-oss-20b at Q6_K using 20.7 GB and Qwen-Image at Q6_K using 19.7 GB. HunyuanImage 2.1 / 3.0 uses 21.6 GB at Q8_0. DeepSeek-Coder-V2 16B / 236B uses 20.4 GB at Q8_0. Apriel-1.5-15B-Thinker uses 19.1 GB at Q8_0. Qwen2.5 14B uses 19.5 GB at Q8_0. Phi-4 and Wan 2.2 T2I both use 17.8 GB at Q8_0.

When a model exceeds the 22.4 GB limit you must offload parts of it to the CPU. This offload process uses your 32 GB system RAM but slows down processing speed significantly. For example Qwen3 8B / 14B / 32B and Qwen2.5-Coder 0.5B to 32B need 23.4 GB at Q4_K_M and require 25.4 GB of system RAM. DeepSeek-Coder 1.3B / 6.7B / 33B needs 24.2 GB at Q4_K_M and requires 26.2 GB of system RAM. Yi 1.5 9B / 34B needs 24.9 GB at Q4_K_M and requires 26.9 GB of system RAM.

All memory calculations assume a standard 4k context window. If you increase the context window to process longer documents the memory usage will rise. This extra memory consumption can push a model that normally fits in unified memory into CPU offloading. Keep your context window at 4k to maintain maximum speed on your Apple M2 Max.