Best local AI models for AMD Pro 5700 XT
16 GB GDDR6. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| gpt-oss-20b | 21B | Q4_K_M | 15.4 GB |
| Reka Flash 3 | 21B | Q4_K_M | 15.4 GB |
| Qwen-Image | 20B | Q4_K_M | 14.6 GB |
| Qwen-Image-Edit | 20B | Q4_K_M | 14.6 GB |
| CogVLM2 | 19B | Q4_K_M | 13.9 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q5_K_M | 14.5 GB |
| Ling-Coder-Lite | 16.8B | Q5_K_M | 14.3 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q6_K | 15.7 GB |
| Kimi-VL A3B | 16B | Q6_K | 15.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q6_K | 14.8 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q6_K | 14.8 GB |
| Qwen2.5 14B | 14.7B | Q6_K | 15.3 GB |
| Phi-3 Medium | 14B | Q6_K | 13.8 GB |
| Phi-4 | 14B | Q6_K | 13.8 GB |
| Phi-4-reasoning / -plus | 14B | Q6_K | 13.8 GB |
| Wan 2.2 T2I | 14B | Q6_K | 13.8 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q6_K | 13.8 GB |
| SkyReels V2 | 14B | Q6_K | 13.8 GB |
| Vicuna 13B | 13B | Q6_K | 12.8 GB |
| HunyuanVideo | 13B | Q6_K | 12.8 GB |
| HunyuanVideo-Avatar | 13B | Q6_K | 12.8 GB |
| LTX-Video / LTX-2 | 13B | Q6_K | 12.8 GB |
| FramePack | 13B | Q6_K | 12.8 GB |
| FLUX.1 dev | 12B | FP8 / optimized | 14.4 GB |
| Gemma 3 12B | 12B | Q8_0 | 15.3 GB |
| Gemma 4 12B | 12B | Q8_0 | 15.3 GB |
| Mistral NeMo 12B | 12B | Q8_0 | 15.3 GB |
| Pixtral 12B | 12B | Q8_0 | 15.3 GB |
| FLUX.1 schnell | 12B | Q8_0 | 15.3 GB |
| FLUX.1 Kontext dev | 12B | Q8_0 | 15.3 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
| Mistral Small 3.2 | 24B | 17.6 GB needed | 19.6 GB |
| Magistral Small | 24B | 17.6 GB needed | 19.6 GB |
| Devstral Small 1.1 | 24B | 17.6 GB needed | 19.6 GB |
| Aria | 25B | 18.3 GB needed | 20.3 GB |
| Gemma 4 26B-A4B | 26B | 19 GB needed | 21 GB |
| Gemma 4 (all sizes) | 26B | 19 GB needed | 21 GB |
| Gemma 3 27B | 27B | 19.8 GB needed | 21.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | 19.8 GB needed | 21.8 GB |
How to read this
The AMD Radeon Pro WX 5700 XT graphics card features 16 GB of GDDR6 memory. This dedicated video memory determines the size of the artificial intelligence models you can run locally. For optimal performance, the entire model must fit within this 16 GB limit. If a model exceeds this capacity, your system must use slower system memory, which decreases processing speeds.
Quantization is a method that compresses model weights to save memory. The quant column shows the best quantization level that fits your hardware. For example, you can run the 21B gpt-oss-20b or Reka Flash 3 at a Q4_K_M quantization using 15.4 GB of video memory. Smaller models like the 15B StarCoder2 or Apriel-1.5-15B-Thinker can run at a higher quality Q6_K quantization using 14.8 GB of memory.
For vision and image tasks, the 20B Qwen-Image and Qwen-Image-Edit models fit comfortably at Q4_K_M quantization using 14.6 GB of memory. The 19B CogVLM2 model uses 13.9 GB of memory at Q4_K_M. If you want to run the FLUX.1 dev model, you can use the FP8 optimized quantization which requires 14.4 GB of video memory.
You can also run highly accurate 12B models at the excellent Q8_0 quantization level. Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, and Pixtral 12B all require 15.3 GB of memory at Q8_0. The FLUX.1 schnell and FLUX.1 Kontext dev models also run at Q8_0 quantization using 15.3 GB of video memory.
When a model is too large for the 16 GB video memory, you can offload parts of it to your system RAM. If your computer has 32 GB of system RAM, you can run larger models with a performance penalty. For example, the 22B Solar Pro and Codestral 22B require 16.1 GB at Q4_K_M quantization and need 18.1 GB of system RAM. The 27B Gemma 3 model requires 19.8 GB at Q4_K_M quantization and needs 21.8 GB of system RAM.
All memory calculations assume a standard 4k context window. If you increase the context window to process longer documents, the model will require significantly more memory. You must choose a smaller model or a lower quantization level if you plan to use long context windows.