Best local AI models for AMD RX 9070 XT
16 GB GDDR6. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| gpt-oss-20b | 21B | Q4_K_M | 15.4 GB |
| Reka Flash 3 | 21B | Q4_K_M | 15.4 GB |
| Qwen-Image | 20B | Q4_K_M | 14.6 GB |
| Qwen-Image-Edit | 20B | Q4_K_M | 14.6 GB |
| CogVLM2 | 19B | Q4_K_M | 13.9 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q5_K_M | 14.5 GB |
| Ling-Coder-Lite | 16.8B | Q5_K_M | 14.3 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q6_K | 15.7 GB |
| Kimi-VL A3B | 16B | Q6_K | 15.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q6_K | 14.8 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q6_K | 14.8 GB |
| Qwen2.5 14B | 14.7B | Q6_K | 15.3 GB |
| Phi-3 Medium | 14B | Q6_K | 13.8 GB |
| Phi-4 | 14B | Q6_K | 13.8 GB |
| Phi-4-reasoning / -plus | 14B | Q6_K | 13.8 GB |
| Wan 2.2 T2I | 14B | Q6_K | 13.8 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q6_K | 13.8 GB |
| SkyReels V2 | 14B | Q6_K | 13.8 GB |
| Vicuna 13B | 13B | Q6_K | 12.8 GB |
| HunyuanVideo | 13B | Q6_K | 12.8 GB |
| HunyuanVideo-Avatar | 13B | Q6_K | 12.8 GB |
| LTX-Video / LTX-2 | 13B | Q6_K | 12.8 GB |
| FramePack | 13B | Q6_K | 12.8 GB |
| FLUX.1 dev | 12B | FP8 / optimized | 14.4 GB |
| Gemma 3 12B | 12B | Q8_0 | 15.3 GB |
| Gemma 4 12B | 12B | Q8_0 | 15.3 GB |
| Mistral NeMo 12B | 12B | Q8_0 | 15.3 GB |
| Pixtral 12B | 12B | Q8_0 | 15.3 GB |
| FLUX.1 schnell | 12B | Q8_0 | 15.3 GB |
| FLUX.1 Kontext dev | 12B | Q8_0 | 15.3 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
| Mistral Small 3.2 | 24B | 17.6 GB needed | 19.6 GB |
| Magistral Small | 24B | 17.6 GB needed | 19.6 GB |
| Devstral Small 1.1 | 24B | 17.6 GB needed | 19.6 GB |
| Aria | 25B | 18.3 GB needed | 20.3 GB |
| Gemma 4 26B-A4B | 26B | 19 GB needed | 21 GB |
| Gemma 4 (all sizes) | 26B | 19 GB needed | 21 GB |
| Gemma 3 27B | 27B | 19.8 GB needed | 21.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | 19.8 GB needed | 21.8 GB |
How to read this
The AMD RX 9070 XT graphics card features 16 GB of GDDR6 memory. This memory size determines which local AI models you can run entirely on your hardware. For the best performance, the entire model must fit inside this onboard memory. If a model exceeds this limit, your system must use slower system memory to process the remaining data.
The quantization column shows the compression level used to fit these models into memory. Quantization reduces the precision of model weights to save space. For example, the 21B gpt-oss-20b and Reka Flash 3 models fit within 15.4 GB of memory using the Q4_K_M quantization. Larger models like Qwen-Image and Qwen-Image-Edit use 14.6 GB of memory at the same Q4_K_M level. This compression allows you to run larger parameter sizes on consumer hardware.
As model sizes decrease, you can use higher precision levels for better output quality. The 17B HunyuanImage 2.1 / 3.0 uses 14.5 GB at Q5_K_M quantization. The 16B DeepSeek-Coder-V2 16B / 236B and Kimi-VL A3B models fit within 15.7 GB using a higher Q6_K quantization. You can also run the 14B Phi-4, Phi-4-reasoning / -plus, Wan 2.2 T2I, and SkyReels V2 models at Q6_K quantization using 13.8 GB of memory.
For maximum precision, you can run 12B models at Q8_0 quantization. The Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, and Pixtral 12B models use 15.3 GB of memory at this level. The FLUX.1 dev model can run at FP8 / optimized quantization using 14.4 GB of memory. Running these models at higher quantization levels preserves original model accuracy.
When a model is too large for the 16 GB GDDR6 memory, you can offload parts of it to your system RAM. This requires at least 32 GB of system RAM. For example, Solar Pro and Codestral 22B need 16.1 GB at Q4_K_M quantization and require 18.1 GB of system RAM. The Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 models need 17.6 GB at Q4_K_M and require 19.6 GB of system RAM. Offloading allows you to run the Gemma 3 27B model which needs 19.8 GB at Q4_K_M and requires 21.8 GB of system RAM. This process makes larger models run but reduces generation speed.
All listed memory figures assume a standard 4k context window. Generating longer responses or processing larger inputs increases memory usage. If you increase the context window beyond 4k tokens, the model will require more memory than the listed figures. You may need to use a lower quantization level to avoid running out of memory during long conversations.