Best local AI models for AMD RX 6900 XT
16 GB GDDR6. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| gpt-oss-20b | 21B | Q4_K_M | 15.4 GB |
| Reka Flash 3 | 21B | Q4_K_M | 15.4 GB |
| Qwen-Image | 20B | Q4_K_M | 14.6 GB |
| Qwen-Image-Edit | 20B | Q4_K_M | 14.6 GB |
| CogVLM2 | 19B | Q4_K_M | 13.9 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q5_K_M | 14.5 GB |
| Ling-Coder-Lite | 16.8B | Q5_K_M | 14.3 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q6_K | 15.7 GB |
| Kimi-VL A3B | 16B | Q6_K | 15.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q6_K | 14.8 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q6_K | 14.8 GB |
| Qwen2.5 14B | 14.7B | Q6_K | 15.3 GB |
| Phi-3 Medium | 14B | Q6_K | 13.8 GB |
| Phi-4 | 14B | Q6_K | 13.8 GB |
| Phi-4-reasoning / -plus | 14B | Q6_K | 13.8 GB |
| Wan 2.2 T2I | 14B | Q6_K | 13.8 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q6_K | 13.8 GB |
| SkyReels V2 | 14B | Q6_K | 13.8 GB |
| Vicuna 13B | 13B | Q6_K | 12.8 GB |
| HunyuanVideo | 13B | Q6_K | 12.8 GB |
| HunyuanVideo-Avatar | 13B | Q6_K | 12.8 GB |
| LTX-Video / LTX-2 | 13B | Q6_K | 12.8 GB |
| FramePack | 13B | Q6_K | 12.8 GB |
| FLUX.1 dev | 12B | FP8 / optimized | 14.4 GB |
| Gemma 3 12B | 12B | Q8_0 | 15.3 GB |
| Gemma 4 12B | 12B | Q8_0 | 15.3 GB |
| Mistral NeMo 12B | 12B | Q8_0 | 15.3 GB |
| Pixtral 12B | 12B | Q8_0 | 15.3 GB |
| FLUX.1 schnell | 12B | Q8_0 | 15.3 GB |
| FLUX.1 Kontext dev | 12B | Q8_0 | 15.3 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
| Mistral Small 3.2 | 24B | 17.6 GB needed | 19.6 GB |
| Magistral Small | 24B | 17.6 GB needed | 19.6 GB |
| Devstral Small 1.1 | 24B | 17.6 GB needed | 19.6 GB |
| Aria | 25B | 18.3 GB needed | 20.3 GB |
| Gemma 4 26B-A4B | 26B | 19 GB needed | 21 GB |
| Gemma 4 (all sizes) | 26B | 19 GB needed | 21 GB |
| Gemma 3 27B | 27B | 19.8 GB needed | 21.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | 19.8 GB needed | 21.8 GB |
How to read this
The AMD Radeon RX 6900 XT graphics card features 16 GB of GDDR6 VRAM. This memory capacity determines the size of the artificial intelligence models you can run entirely on the hardware. To run a model at maximum speed, the entire model must fit inside this 16 GB limit. If a model exceeds this limit, the system must transfer data between the graphics card and system memory, which slows down performance.
The quantization column shows the compression level applied to each model. Raw models are too large for consumer hardware, so they are compressed into smaller formats like Q4_K_M, Q5_K_M, Q6_K, or Q8_0. A higher quantization number like Q8_0 preserves more original model quality but requires more memory. A lower quantization like Q4_K_M uses less memory, allowing you to run larger models like the 21B gpt-oss-20b or Reka Flash 3 within 15.4 GB of VRAM.
For models that fit completely in VRAM, you can run the 20B Qwen-Image at Q4_K_M using 14.6 GB. The 19B CogVLM2 fits at Q4_K_M using 13.9 GB. You can run the 17B HunyuanImage 2.1 / 3.0 at Q5_K_M using 14.5 GB. The 16.8B Ling-Coder-Lite fits at Q5_K_M using 14.3 GB. The 16B DeepSeek-Coder-V2 16B / 236B and Kimi-VL A3B both fit at Q6_K using 15.7 GB. The 15B Apriel-1.5-15B-Thinker and StarCoder2 3B / 7B / 15B fit at Q6_K using 14.8 GB.
Several 14B models fit at Q6_K using 13.8 GB, including Qwen2.5 14B, Phi-3 Medium, Phi-4, Phi-4-reasoning / -plus, Wan 2.2 T2I, Wan 2.1 (1.3B / 14B), and SkyReels V2. The 13B Vicuna 13B, HunyuanVideo, HunyuanVideo-Avatar, LTX-Video / LTX-2, and FramePack fit at Q6_K using 12.8 GB. At Q8_0 quantization, you can run 12B models like Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, and FLUX.1 Kontext dev using 15.3 GB. FLUX.1 dev fits using 14.4 GB at FP8 / optimized.
When a model is too large for the 16 GB VRAM, you can offload the extra weight to your system RAM. This requires a system with at least 32 GB of system RAM. Offloading allows you to run larger models, but it reduces generation speed because the system RAM is much slower than GDDR6 VRAM. For example, Solar Pro and Codestral 22B need 16.1 GB at Q4_K_M and require 18.1 GB of system RAM. Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 need 17.6 GB at Q4_K_M and require 19.6 GB of system RAM.
Other offload options include Aria, which needs 18.3 GB at Q4_K_M and 20.3 GB of system RAM. Gemma 4 26B-A4B and Gemma 4 (all sizes) need 19 GB at Q4_K_M and 21 GB of system RAM. The largest offload options are Gemma 3 27B and Gemma 3 4B/12B/27B (vision), which need 19.8 GB at Q4_K_M and 21.8 GB of system RAM. Note that all memory calculations assume a standard 4k context window. If you increase the context window to process longer texts, the model will require significantly more VRAM.