Best local AI models for Apple M2

11.2 GB usable of 16 GB unified memory. At a 4k context, 144 of the 233 models in our catalog with verified parameter counts fit fully, up to Apriel-1.5-15B-Thinker at 15B parameters. Computed for the 16 GB configuration; a larger memory configuration fits more.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 144 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Apriel-1.5-15B-Thinker15BQ4_K_M11 GB
StarCoder2 3B / 7B / 15B15BQ4_K_M11 GB
Phi-3 Medium14BQ4_K_M10.2 GB
Phi-414BQ4_K_M10.2 GB
Phi-4-reasoning / -plus14BQ4_K_M10.2 GB
Wan 2.2 T2I14BQ4_K_M10.2 GB
Wan 2.1 (1.3B / 14B)14BQ4_K_M10.2 GB
SkyReels V214BQ4_K_M10.2 GB
Vicuna 13B13BQ5_K_M11.1 GB
HunyuanVideo13BQ5_K_M11.1 GB
HunyuanVideo-Avatar13BQ5_K_M11.1 GB
LTX-Video / LTX-213BQ5_K_M11.1 GB
FramePack13BQ5_K_M11.1 GB
Gemma 3 12B12BQ5_K_M10.2 GB
Gemma 4 12B12BQ5_K_M10.2 GB
Mistral NeMo 12B12BQ5_K_M10.2 GB
Pixtral 12B12BQ5_K_M10.2 GB
FLUX.1 schnell12BQ5_K_M10.2 GB
FLUX.1 Kontext dev12BQ5_K_M10.2 GB
FLUX.1 Krea dev12BQ5_K_M10.2 GB
Open-Sora 2.011BQ6_K10.8 GB
Mochi 110BQ6_K9.8 GB
Gemma 2 9B9BQ6_K10.3 GB
Nemotron Nano 4B / 9B9BQ6_K8.9 GB
GLM-4 9B / GLM-4.5-Air9BQ6_K8.9 GB
Yi-Coder 1.5B / 9B9BQ6_K8.9 GB
GLM-4-9B-Chat / CodeGeeX49BQ6_K8.9 GB
GLM-4V-9B / GLM-4.1V-Thinking9BQ6_K8.9 GB
Chroma8.9BQ6_K8.8 GB
Llama 3.1 8B8BQ8_010.7 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
FLUX.1 dev12B14.4 GB needed16.4 GB
Qwen2.5 14B14.7B11.6 GB needed13.6 GB
DeepSeek-Coder-V2 16B / 236B16B11.7 GB needed13.7 GB
Kimi-VL A3B16B11.7 GB needed13.7 GB
Ling-Coder-Lite16.8B12.3 GB needed14.3 GB
HunyuanImage 2.1 / 3.017B12.4 GB needed14.4 GB
CogVLM219B13.9 GB needed15.9 GB
Qwen-Image20B14.6 GB needed16.6 GB
Qwen-Image-Edit20B14.6 GB needed16.6 GB
gpt-oss-20b21B15.4 GB needed17.4 GB

How to read this

The Apple M2 chip features a unified memory architecture that shares RAM between the processor and the graphics engine. On a system with a 16 GB unified memory pool, the maximum usable share allocated for local AI models is exactly 11.2 GB. This limit determines which models can run entirely on the graphics hardware for fast generation speeds.

To fit within this 11.2 GB limit, models use quantization to compress their weights. The quant column shows the best format that fits in memory. For example, the 15B StarCoder2 and Apriel-1.5-15B-Thinker models fit using the Q4_K_M quant which uses 11 GB. The 13B Vicuna, HunyuanVideo, HunyuanVideo-Avatar, LTX-Video, LTX-2, and FramePack models fit using the Q5_K_M quant which uses 11.1 GB. The 11B Open-Sora 2.0 fits using the Q6_K quant at 10.8 GB, while the Llama 3.1 8B model can run at the higher quality Q8_0 quant using 10.7 GB.

Other models fit comfortably within the memory limit at various quantization levels. The 14B models including Phi-3 Medium, Phi-4, Phi-4-reasoning, Phi-4-plus, Wan 2.2 T2I, Wan 2.1, SkyReels V2, and the 12B models including Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, FLUX.1 Kontext dev, and FLUX.1 Krea dev all use 10.2 GB. The Mochi 1 model uses 9.8 GB at Q6_K. The Gemma 2 9B model uses 10.3 GB at Q6_K. The 9B models including Nemotron Nano, Yi-Coder, GLM-4, GLM-4.5-Air, GLM-4-9B-Chat, CodeGeeX4, GLM-4V-9B, and GLM-4.1V-Thinking use 8.9 GB at Q6_K, while Chroma uses 8.8 GB.

When a model exceeds the 11.2 GB limit, the system must offload parts of the model to the CPU. This offload process requires a larger system RAM pool such as 32 GB. Offloading allows you to run larger models but it significantly reduces generation speeds because data must transfer between the CPU and GPU.

Several models require this CPU offload setup to run. The FLUX.1 dev model needs 14.4 GB at FP8 or optimized settings and requires 16.4 GB of system RAM. The Qwen2.5 14B model needs 11.6 GB at Q4_K_M and requires 13.6 GB of system RAM. The DeepSeek-Coder-V2 16B, DeepSeek-Coder-V2 236B, and Kimi-VL A3B models need 11.7 GB at Q4_K_M and require 13.7 GB of system RAM. The Ling-Coder-Lite model needs 12.3 GB at Q4_K_M and requires 14.3 GB of system RAM. The HunyuanImage 2.1 and HunyuanImage 3.0 models need 12.4 GB at Q4_K_M and require 14.4 GB of system RAM.

Larger offload models demand even more system memory. The CogVLM2 model needs 13.9 GB at Q4_K_M and requires 15.9 GB of system RAM. The Qwen-Image and Qwen-Image-Edit models need 14.6 GB at Q4_K_M and require 16.6 GB of system RAM. The gpt-oss-20b model needs 15.4 GB at Q4_K_M and requires 17.4 GB of system RAM.

Users must also consider the memory cost of context. The memory usage figures listed for these models are calculated using a basic 4k context window. If you increase the context window to process longer documents or chat histories, the system will require additional memory which can push a fitting model over the 11.2 GB limit.