Best local AI models for AMD RX 9070
16 GB GDDR6. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| gpt-oss-20b | 21B | Q4_K_M | 15.4 GB |
| Reka Flash 3 | 21B | Q4_K_M | 15.4 GB |
| Qwen-Image | 20B | Q4_K_M | 14.6 GB |
| Qwen-Image-Edit | 20B | Q4_K_M | 14.6 GB |
| CogVLM2 | 19B | Q4_K_M | 13.9 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q5_K_M | 14.5 GB |
| Ling-Coder-Lite | 16.8B | Q5_K_M | 14.3 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q6_K | 15.7 GB |
| Kimi-VL A3B | 16B | Q6_K | 15.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q6_K | 14.8 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q6_K | 14.8 GB |
| Qwen2.5 14B | 14.7B | Q6_K | 15.3 GB |
| Phi-3 Medium | 14B | Q6_K | 13.8 GB |
| Phi-4 | 14B | Q6_K | 13.8 GB |
| Phi-4-reasoning / -plus | 14B | Q6_K | 13.8 GB |
| Wan 2.2 T2I | 14B | Q6_K | 13.8 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q6_K | 13.8 GB |
| SkyReels V2 | 14B | Q6_K | 13.8 GB |
| Vicuna 13B | 13B | Q6_K | 12.8 GB |
| HunyuanVideo | 13B | Q6_K | 12.8 GB |
| HunyuanVideo-Avatar | 13B | Q6_K | 12.8 GB |
| LTX-Video / LTX-2 | 13B | Q6_K | 12.8 GB |
| FramePack | 13B | Q6_K | 12.8 GB |
| FLUX.1 dev | 12B | FP8 / optimized | 14.4 GB |
| Gemma 3 12B | 12B | Q8_0 | 15.3 GB |
| Gemma 4 12B | 12B | Q8_0 | 15.3 GB |
| Mistral NeMo 12B | 12B | Q8_0 | 15.3 GB |
| Pixtral 12B | 12B | Q8_0 | 15.3 GB |
| FLUX.1 schnell | 12B | Q8_0 | 15.3 GB |
| FLUX.1 Kontext dev | 12B | Q8_0 | 15.3 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
| Mistral Small 3.2 | 24B | 17.6 GB needed | 19.6 GB |
| Magistral Small | 24B | 17.6 GB needed | 19.6 GB |
| Devstral Small 1.1 | 24B | 17.6 GB needed | 19.6 GB |
| Aria | 25B | 18.3 GB needed | 20.3 GB |
| Gemma 4 26B-A4B | 26B | 19 GB needed | 21 GB |
| Gemma 4 (all sizes) | 26B | 19 GB needed | 21 GB |
| Gemma 3 27B | 27B | 19.8 GB needed | 21.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | 19.8 GB needed | 21.8 GB |
How to read this
The AMD RX 9070 graphics card features 16 GB of GDDR6 memory. This memory size determines which local AI models you can run entirely on your hardware. When a model fits completely within this 16 GB limit, it runs at maximum speed because the graphics processor has direct access to the data. If a model exceeds this limit, you must offload parts of it to your system memory.
To fit larger models into the 16 GB memory of the AMD RX 9070, developers use quantization. The quant column shows the specific compression level used for each model. For example, gpt-oss-20b and Reka Flash 3 are 21B models that fit using the Q4_K_M quant, which uses 15.4 GB of memory. Qwen-Image and Qwen-Image-Edit are 20B models that fit at Q4_K_M using 14.6 GB of memory. CogVLM2 is a 19B model that fits at Q4_K_M using 13.9 GB of memory.
Higher quality quants are available for slightly smaller models. HunyuanImage 2.1 / 3.0 is a 17B model that fits at Q5_K_M using 14.5 GB of memory. Ling-Coder-Lite is a 16.8B model that fits at Q5_K_M using 14.3 GB of memory. DeepSeek-Coder-V2 16B / 236B and Kimi-VL A3B are 16B models that fit at Q6_K using 15.7 GB of memory. Apriel-1.5-15B-Thinker and StarCoder2 3B / 7B / 15B are 15B models that fit at Q6_K using 14.8 GB of memory. Qwen2.5 14B fits at Q6_K using 15.3 GB of memory.
Several 14B and 13B models run comfortably at Q6_K. Phi-3 Medium, Phi-4, Phi-4-reasoning / -plus, Wan 2.2 T2I, Wan 2.1 (1.3B / 14B), and SkyReels V2 are 14B models that use 13.8 GB of memory. Vicuna 13B, HunyuanVideo, HunyuanVideo-Avatar, LTX-Video / LTX-2, and FramePack are 13B models that use 12.8 GB of memory. You can also run 12B models at Q8_0 using 15.3 GB of memory, including Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, and FLUX.1 Kontext dev. FLUX.1 dev fits at FP8 / optimized using 14.4 GB of memory.
Running these models at their limit leaves little room for context memory. The listed memory usage figures are calculated with a basic 4k context window. If you increase the context window to process longer documents, the model will require more memory and may overflow the 16 GB limit of your card.
When a model is too large for the graphics card, you can offload layers to your system RAM. This offload process requires a system with at least 32 GB of system RAM, but it reduces generation speed. Solar Pro and Codestral 22B are 22B models that need 16.1 GB at Q4_K_M and require 18.1 GB of system RAM. Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 are 24B models that need 17.6 GB at Q4_K_M and require 19.6 GB of system RAM. Aria is a 25B model that needs 18.3 GB at Q4_K_M and requires 20.3 GB of system RAM.
Even larger models can run using system RAM offloading. Gemma 4 26B-A4B and Gemma 4 (all sizes) are 26B models that need 19 GB at Q4_K_M and require 21 GB of system RAM. Gemma 3 27B and Gemma 3 4B/12B/27B (vision) are 27B models that need 19.8 GB at Q4_K_M and require 21.8 GB of system RAM.