Best local AI models for AMD RX 6950 XT
16 GB GDDR6. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| gpt-oss-20b | 21B | Q4_K_M | 15.4 GB |
| Reka Flash 3 | 21B | Q4_K_M | 15.4 GB |
| Qwen-Image | 20B | Q4_K_M | 14.6 GB |
| Qwen-Image-Edit | 20B | Q4_K_M | 14.6 GB |
| CogVLM2 | 19B | Q4_K_M | 13.9 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q5_K_M | 14.5 GB |
| Ling-Coder-Lite | 16.8B | Q5_K_M | 14.3 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q6_K | 15.7 GB |
| Kimi-VL A3B | 16B | Q6_K | 15.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q6_K | 14.8 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q6_K | 14.8 GB |
| Qwen2.5 14B | 14.7B | Q6_K | 15.3 GB |
| Phi-3 Medium | 14B | Q6_K | 13.8 GB |
| Phi-4 | 14B | Q6_K | 13.8 GB |
| Phi-4-reasoning / -plus | 14B | Q6_K | 13.8 GB |
| Wan 2.2 T2I | 14B | Q6_K | 13.8 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q6_K | 13.8 GB |
| SkyReels V2 | 14B | Q6_K | 13.8 GB |
| Vicuna 13B | 13B | Q6_K | 12.8 GB |
| HunyuanVideo | 13B | Q6_K | 12.8 GB |
| HunyuanVideo-Avatar | 13B | Q6_K | 12.8 GB |
| LTX-Video / LTX-2 | 13B | Q6_K | 12.8 GB |
| FramePack | 13B | Q6_K | 12.8 GB |
| FLUX.1 dev | 12B | FP8 / optimized | 14.4 GB |
| Gemma 3 12B | 12B | Q8_0 | 15.3 GB |
| Gemma 4 12B | 12B | Q8_0 | 15.3 GB |
| Mistral NeMo 12B | 12B | Q8_0 | 15.3 GB |
| Pixtral 12B | 12B | Q8_0 | 15.3 GB |
| FLUX.1 schnell | 12B | Q8_0 | 15.3 GB |
| FLUX.1 Kontext dev | 12B | Q8_0 | 15.3 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
| Mistral Small 3.2 | 24B | 17.6 GB needed | 19.6 GB |
| Magistral Small | 24B | 17.6 GB needed | 19.6 GB |
| Devstral Small 1.1 | 24B | 17.6 GB needed | 19.6 GB |
| Aria | 25B | 18.3 GB needed | 20.3 GB |
| Gemma 4 26B-A4B | 26B | 19 GB needed | 21 GB |
| Gemma 4 (all sizes) | 26B | 19 GB needed | 21 GB |
| Gemma 3 27B | 27B | 19.8 GB needed | 21.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | 19.8 GB needed | 21.8 GB |
How to read this
The AMD Radeon RX 6950 XT graphics card features 16 GB of GDDR6 memory. This onboard memory determines the maximum size of the artificial intelligence models you can run locally. To run a model entirely on your hardware at full speed, the entire model must fit inside this VRAM limit. If a model exceeds 16 GB, your system must use alternative execution methods.
Quantization is a compression technique that reduces model size. The quantization column shows the best quality level that fits within your hardware limit. For example, the 21B gpt-oss-20b and Reka Flash 3 models fit using the Q4_K_M quantization, which uses 15.4 GB of memory. Smaller models like the 15B StarCoder2 3B / 7B / 15B and Apriel-1.5-15B-Thinker can run at a higher Q6_K quantization, using 14.8 GB of memory. The 12B Gemma 3 12B, Gemma 4 12B, and Mistral NeMo 12B models can run at the high quality Q8_0 quantization, using 15.3 GB of memory.
Vision and image generation models also fit within this memory limit. The 20B Qwen-Image and Qwen-Image-Edit models use 14.6 GB of memory at Q4_K_M. The 12B FLUX.1 dev model fits using the FP8 / optimized quantization, which uses 14.4 GB of memory. Video generation models like HunyuanVideo, HunyuanVideo-Avatar, and LTX-Video / LTX-2 use 12.8 GB of memory at Q6_K.
When a model is too large for the 16 GB VRAM, you can use CPU offloading. This process splits the model between your graphics card and your system RAM. CPU offloading allows you to run larger models, but it significantly reduces processing speed. For this setup, we assume your computer has 32 GB of system RAM to handle the overflow.
With CPU offloading, you can run the 22B Solar Pro and Codestral 22B models, which need 16.1 GB of memory at Q4_K_M and require 18.1 GB of system RAM. You can also run the 24B Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 models, which need 17.6 GB at Q4_K_M and require 19.6 GB of system RAM. The largest offload options include the 27B Gemma 3 27B and Gemma 3 4B/12B/27B (vision) models, which need 19.8 GB at Q4_K_M and require 21.8 GB of system RAM.
All memory calculations are based on a standard 4k context window. The context window is the amount of text the model can read and write at one time. If you increase the context window beyond 4k, the model will require more memory. This extra memory requirement may force you to use a lower quantization level or switch to CPU offloading.