Best local AI models for AMD Pro W5700X
16 GB GDDR6. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| gpt-oss-20b | 21B | Q4_K_M | 15.4 GB |
| Reka Flash 3 | 21B | Q4_K_M | 15.4 GB |
| Qwen-Image | 20B | Q4_K_M | 14.6 GB |
| Qwen-Image-Edit | 20B | Q4_K_M | 14.6 GB |
| CogVLM2 | 19B | Q4_K_M | 13.9 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q5_K_M | 14.5 GB |
| Ling-Coder-Lite | 16.8B | Q5_K_M | 14.3 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q6_K | 15.7 GB |
| Kimi-VL A3B | 16B | Q6_K | 15.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q6_K | 14.8 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q6_K | 14.8 GB |
| Qwen2.5 14B | 14.7B | Q6_K | 15.3 GB |
| Phi-3 Medium | 14B | Q6_K | 13.8 GB |
| Phi-4 | 14B | Q6_K | 13.8 GB |
| Phi-4-reasoning / -plus | 14B | Q6_K | 13.8 GB |
| Wan 2.2 T2I | 14B | Q6_K | 13.8 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q6_K | 13.8 GB |
| SkyReels V2 | 14B | Q6_K | 13.8 GB |
| Vicuna 13B | 13B | Q6_K | 12.8 GB |
| HunyuanVideo | 13B | Q6_K | 12.8 GB |
| HunyuanVideo-Avatar | 13B | Q6_K | 12.8 GB |
| LTX-Video / LTX-2 | 13B | Q6_K | 12.8 GB |
| FramePack | 13B | Q6_K | 12.8 GB |
| FLUX.1 dev | 12B | FP8 / optimized | 14.4 GB |
| Gemma 3 12B | 12B | Q8_0 | 15.3 GB |
| Gemma 4 12B | 12B | Q8_0 | 15.3 GB |
| Mistral NeMo 12B | 12B | Q8_0 | 15.3 GB |
| Pixtral 12B | 12B | Q8_0 | 15.3 GB |
| FLUX.1 schnell | 12B | Q8_0 | 15.3 GB |
| FLUX.1 Kontext dev | 12B | Q8_0 | 15.3 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
| Mistral Small 3.2 | 24B | 17.6 GB needed | 19.6 GB |
| Magistral Small | 24B | 17.6 GB needed | 19.6 GB |
| Devstral Small 1.1 | 24B | 17.6 GB needed | 19.6 GB |
| Aria | 25B | 18.3 GB needed | 20.3 GB |
| Gemma 4 26B-A4B | 26B | 19 GB needed | 21 GB |
| Gemma 4 (all sizes) | 26B | 19 GB needed | 21 GB |
| Gemma 3 27B | 27B | 19.8 GB needed | 21.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | 19.8 GB needed | 21.8 GB |
How to read this
The AMD Radeon Pro W5700X features 16 GB of GDDR6 VRAM. This memory capacity determines which local AI models you can run entirely on your graphics hardware. When a model fits completely within this 16 GB boundary, you get the fastest possible generation speeds. If a model exceeds this limit, your system must split the workload between your graphics card and your system memory.
The quantization column shows the compression level used to fit these models into memory. Quantization reduces the size of model weights to save space. For example, the 21B models gpt-oss-20b and Reka Flash 3 fit within 15.4 GB of VRAM when compressed to the Q4_K_M format. Larger models like DeepSeek-Coder-V2 16B / 236B and Kimi-VL A3B can run at a higher quality Q6_K quantization, which uses 15.7 GB of VRAM.
For 12B models, you can run at the highest quality levels. Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, and FLUX.1 Kontext dev all run at the Q8_0 quantization level using 15.3 GB of VRAM. The FLUX.1 dev model runs at an optimized FP8 quantization level using 14.4 GB of VRAM. Vision and video models like Qwen-Image at 14.6 GB and HunyuanVideo at 12.8 GB also fit comfortably.
When you want to run models that exceed 16 GB, you must use CPU offloading. This process shares the workload with your system RAM but slows down generation speeds. If you have 32 GB of system RAM, you can run Solar Pro or Codestral 22B at Q4_K_M, which requires 16.1 GB of memory and 18.1 GB of system RAM. Larger models like Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 require 17.6 GB of memory and 19.6 GB of system RAM.
The largest offloaded models you can run with 32 GB of system RAM include Gemma 3 27B and Gemma 3 4B/12B/27B (vision). These models require 19.8 GB of memory at Q4_K_M and 21.8 GB of system RAM. Gemma 4 (all sizes) at 26B requires 19 GB of memory and 21 GB of system RAM. Aria at 25B requires 18.3 GB of memory and 20.3 GB of system RAM.
All VRAM calculations assume a standard 4k context window. If you increase the context window to process longer documents or chat histories, the memory usage will rise. This extra memory demand may force you to use a lower quantization level or offload more layers to your system RAM.