Best local AI models for AMD RX 9060 XT 16GB
16 GB GDDR6. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| gpt-oss-20b | 21B | Q4_K_M | 15.4 GB |
| Reka Flash 3 | 21B | Q4_K_M | 15.4 GB |
| Qwen-Image | 20B | Q4_K_M | 14.6 GB |
| Qwen-Image-Edit | 20B | Q4_K_M | 14.6 GB |
| CogVLM2 | 19B | Q4_K_M | 13.9 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q5_K_M | 14.5 GB |
| Ling-Coder-Lite | 16.8B | Q5_K_M | 14.3 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q6_K | 15.7 GB |
| Kimi-VL A3B | 16B | Q6_K | 15.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q6_K | 14.8 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q6_K | 14.8 GB |
| Qwen2.5 14B | 14.7B | Q6_K | 15.3 GB |
| Phi-3 Medium | 14B | Q6_K | 13.8 GB |
| Phi-4 | 14B | Q6_K | 13.8 GB |
| Phi-4-reasoning / -plus | 14B | Q6_K | 13.8 GB |
| Wan 2.2 T2I | 14B | Q6_K | 13.8 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q6_K | 13.8 GB |
| SkyReels V2 | 14B | Q6_K | 13.8 GB |
| Vicuna 13B | 13B | Q6_K | 12.8 GB |
| HunyuanVideo | 13B | Q6_K | 12.8 GB |
| HunyuanVideo-Avatar | 13B | Q6_K | 12.8 GB |
| LTX-Video / LTX-2 | 13B | Q6_K | 12.8 GB |
| FramePack | 13B | Q6_K | 12.8 GB |
| FLUX.1 dev | 12B | FP8 / optimized | 14.4 GB |
| Gemma 3 12B | 12B | Q8_0 | 15.3 GB |
| Gemma 4 12B | 12B | Q8_0 | 15.3 GB |
| Mistral NeMo 12B | 12B | Q8_0 | 15.3 GB |
| Pixtral 12B | 12B | Q8_0 | 15.3 GB |
| FLUX.1 schnell | 12B | Q8_0 | 15.3 GB |
| FLUX.1 Kontext dev | 12B | Q8_0 | 15.3 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
| Mistral Small 3.2 | 24B | 17.6 GB needed | 19.6 GB |
| Magistral Small | 24B | 17.6 GB needed | 19.6 GB |
| Devstral Small 1.1 | 24B | 17.6 GB needed | 19.6 GB |
| Aria | 25B | 18.3 GB needed | 20.3 GB |
| Gemma 4 26B-A4B | 26B | 19 GB needed | 21 GB |
| Gemma 4 (all sizes) | 26B | 19 GB needed | 21 GB |
| Gemma 3 27B | 27B | 19.8 GB needed | 21.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | 19.8 GB needed | 21.8 GB |
How to read this
The AMD Radeon RX 9060 XT graphics card comes equipped with 16 GB of GDDR6 dedicated video memory. This physical memory limit determines the maximum size of the artificial intelligence models you can run locally on your system. To load a model entirely onto the graphics processor for fast execution, the total memory footprint of the model must remain under the 16 GB threshold. This allocation must also leave a small amount of headroom for your operating system and display outputs.
The quantization column indicates the compression level applied to each model. Raw models are often too large for consumer hardware, so they are quantized to smaller bit widths. For example, the 21B gpt-oss-20b and Reka Flash 3 models fit into 15.4 GB of video memory when using the Q4_K_M quantization. Other models like Qwen-Image and Qwen-Image-Edit use 14.6 GB at Q4_K_M. As model sizes decrease, you can use higher quality quantizations. The 15B StarCoder2 and Apriel-1.5-15B-Thinker models run at Q6_K using 14.8 GB, while the 12B Gemma 3 12B and Mistral NeMo 12B can run at Q8_0 using 15.3 GB.
If a model exceeds the 16 GB video memory limit, you can use CPU offloading. This process splits the model layers between your graphics card and your system RAM. Offloading allows you to run larger models, but it significantly reduces processing speed because system RAM is much slower than GDDR6 memory. For instance, running the 22B Solar Pro or Codestral 22B at Q4_K_M requires 16.1 GB of memory, which demands 18.1 GB of system RAM. Larger models like Gemma 3 27B require 19.8 GB at Q4_K_M, which needs 21.8 GB of system RAM to function.
When selecting your quantization level, you must also consider the context window. The memory figures listed for these models assume a standard 4k context window. If you increase the context length to process longer documents or extended conversations, the memory required for the key value cache will grow. This extra memory usage can push a model that fits at 4k context over the 16 GB limit, which will force your system to offload layers to system RAM and slow down generation speeds.
The RX 9060 XT supports a wide variety of model types within its 16 GB limit. You can run vision models like CogVLM2 at Q4_K_M using 13.9 GB, or image generation models like HunyuanImage 2.1 / 3.0 at Q5_K_M using 14.5 GB. Video generation models are also compatible, with HunyuanVideo and LTX-Video running at Q6_K using 12.8 GB. For text and coding tasks, the DeepSeek-Coder-V2 16B / 236B model fits at Q6_K using 15.7 GB, and Phi-4 runs at Q6_K using 13.8 GB.