Best local AI models for AMD RX 6800 XT
16 GB GDDR6. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| gpt-oss-20b | 21B | Q4_K_M | 15.4 GB |
| Reka Flash 3 | 21B | Q4_K_M | 15.4 GB |
| Qwen-Image | 20B | Q4_K_M | 14.6 GB |
| Qwen-Image-Edit | 20B | Q4_K_M | 14.6 GB |
| CogVLM2 | 19B | Q4_K_M | 13.9 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q5_K_M | 14.5 GB |
| Ling-Coder-Lite | 16.8B | Q5_K_M | 14.3 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q6_K | 15.7 GB |
| Kimi-VL A3B | 16B | Q6_K | 15.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q6_K | 14.8 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q6_K | 14.8 GB |
| Qwen2.5 14B | 14.7B | Q6_K | 15.3 GB |
| Phi-3 Medium | 14B | Q6_K | 13.8 GB |
| Phi-4 | 14B | Q6_K | 13.8 GB |
| Phi-4-reasoning / -plus | 14B | Q6_K | 13.8 GB |
| Wan 2.2 T2I | 14B | Q6_K | 13.8 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q6_K | 13.8 GB |
| SkyReels V2 | 14B | Q6_K | 13.8 GB |
| Vicuna 13B | 13B | Q6_K | 12.8 GB |
| HunyuanVideo | 13B | Q6_K | 12.8 GB |
| HunyuanVideo-Avatar | 13B | Q6_K | 12.8 GB |
| LTX-Video / LTX-2 | 13B | Q6_K | 12.8 GB |
| FramePack | 13B | Q6_K | 12.8 GB |
| FLUX.1 dev | 12B | FP8 / optimized | 14.4 GB |
| Gemma 3 12B | 12B | Q8_0 | 15.3 GB |
| Gemma 4 12B | 12B | Q8_0 | 15.3 GB |
| Mistral NeMo 12B | 12B | Q8_0 | 15.3 GB |
| Pixtral 12B | 12B | Q8_0 | 15.3 GB |
| FLUX.1 schnell | 12B | Q8_0 | 15.3 GB |
| FLUX.1 Kontext dev | 12B | Q8_0 | 15.3 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
| Mistral Small 3.2 | 24B | 17.6 GB needed | 19.6 GB |
| Magistral Small | 24B | 17.6 GB needed | 19.6 GB |
| Devstral Small 1.1 | 24B | 17.6 GB needed | 19.6 GB |
| Aria | 25B | 18.3 GB needed | 20.3 GB |
| Gemma 4 26B-A4B | 26B | 19 GB needed | 21 GB |
| Gemma 4 (all sizes) | 26B | 19 GB needed | 21 GB |
| Gemma 3 27B | 27B | 19.8 GB needed | 21.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | 19.8 GB needed | 21.8 GB |
How to read this
The AMD Radeon RX 6800 XT graphics card features 16 GB of GDDR6 VRAM. This onboard memory determines the maximum size of the artificial intelligence models you can run locally. To run a model entirely on your graphics hardware, the model files and the active memory context must fit within this 16 GB limit. Running models locally on your GPU ensures the fastest generation speeds.
Quantization is a method that compresses model files to save space. In our listings, the best quant column shows the highest quality quantization level that fits within your VRAM. For example, the 21B gpt-oss-20b and Reka Flash 3 models fit at Q4_K_M quantization, using 15.4 GB of VRAM. The 20B Qwen-Image and Qwen-Image-Edit models fit at Q4_K_M quantization, using 14.6 GB of VRAM. CogVLM2 at 19B fits at Q4_K_M quantization, using 13.9 GB of VRAM.
As model sizes decrease, you can use higher quality quantizations. HunyuanImage 2.1 / 3.0 at 17B fits at Q5_K_M quantization, using 14.5 GB of VRAM. Ling-Coder-Lite at 16.8B fits at Q5_K_M quantization, using 14.3 GB of VRAM. Several 16B and 15B models can run at Q6_K quantization. These include DeepSeek-Coder-V2 16B / 236B and Kimi-VL A3B, which use 15.7 GB of VRAM. Apriel-1.5-15B-Thinker and StarCoder2 3B / 7B / 15B also run at Q6_K quantization, using 14.8 GB of VRAM.
Many popular 14B and 13B models run at Q6_K quantization. Qwen2.5 14B uses 15.3 GB of VRAM. Phi-3 Medium, Phi-4, Phi-4-reasoning / -plus, Wan 2.2 T2I, Wan 2.1 (1.3B / 14B), and SkyReels V2 all use 13.8 GB of VRAM. Vicuna 13B, HunyuanVideo, HunyuanVideo-Avatar, LTX-Video / LTX-2, and FramePack use 12.8 GB of VRAM. At the 12B size, you can run the highest Q8_0 quantization. Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, and FLUX.1 Kontext dev use 15.3 GB of VRAM at Q8_0. FLUX.1 dev uses 14.4 GB of VRAM with FP8 / optimized quantization.
If a model exceeds 16 GB of VRAM, you can offload parts of the model to your system RAM. This offloading process requires a system with at least 32 GB of system RAM. Offloading allows you to run larger models, but it significantly reduces processing speeds. Solar Pro and Codestral 22B need 16.1 GB at Q4_K_M quantization and require 18.1 GB of system RAM. Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 need 17.6 GB at Q4_K_M quantization and require 19.6 GB of system RAM.
Larger offload options include Aria, which needs 18.3 GB at Q4_K_M quantization and requires 20.3 GB of system RAM. Gemma 4 26B-A4B and Gemma 4 (all sizes) need 19 GB at Q4_K_M quantization and require 21 GB of system RAM. Gemma 3 27B and Gemma 3 4B/12B/27B (vision) need 19.8 GB at Q4_K_M quantization and require 21.8 GB of system RAM.
All VRAM calculations assume a standard 4k context window. The context window is the memory used for the history of your current conversation. If you increase the context window beyond 4k tokens, the model will require more VRAM. This extra memory usage might force you to use a lower quantization level or offload more layers to system RAM.