Best local AI models for AMD RX 6800
16 GB GDDR6. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| gpt-oss-20b | 21B | Q4_K_M | 15.4 GB |
| Reka Flash 3 | 21B | Q4_K_M | 15.4 GB |
| Qwen-Image | 20B | Q4_K_M | 14.6 GB |
| Qwen-Image-Edit | 20B | Q4_K_M | 14.6 GB |
| CogVLM2 | 19B | Q4_K_M | 13.9 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q5_K_M | 14.5 GB |
| Ling-Coder-Lite | 16.8B | Q5_K_M | 14.3 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q6_K | 15.7 GB |
| Kimi-VL A3B | 16B | Q6_K | 15.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q6_K | 14.8 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q6_K | 14.8 GB |
| Qwen2.5 14B | 14.7B | Q6_K | 15.3 GB |
| Phi-3 Medium | 14B | Q6_K | 13.8 GB |
| Phi-4 | 14B | Q6_K | 13.8 GB |
| Phi-4-reasoning / -plus | 14B | Q6_K | 13.8 GB |
| Wan 2.2 T2I | 14B | Q6_K | 13.8 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q6_K | 13.8 GB |
| SkyReels V2 | 14B | Q6_K | 13.8 GB |
| Vicuna 13B | 13B | Q6_K | 12.8 GB |
| HunyuanVideo | 13B | Q6_K | 12.8 GB |
| HunyuanVideo-Avatar | 13B | Q6_K | 12.8 GB |
| LTX-Video / LTX-2 | 13B | Q6_K | 12.8 GB |
| FramePack | 13B | Q6_K | 12.8 GB |
| FLUX.1 dev | 12B | FP8 / optimized | 14.4 GB |
| Gemma 3 12B | 12B | Q8_0 | 15.3 GB |
| Gemma 4 12B | 12B | Q8_0 | 15.3 GB |
| Mistral NeMo 12B | 12B | Q8_0 | 15.3 GB |
| Pixtral 12B | 12B | Q8_0 | 15.3 GB |
| FLUX.1 schnell | 12B | Q8_0 | 15.3 GB |
| FLUX.1 Kontext dev | 12B | Q8_0 | 15.3 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
| Mistral Small 3.2 | 24B | 17.6 GB needed | 19.6 GB |
| Magistral Small | 24B | 17.6 GB needed | 19.6 GB |
| Devstral Small 1.1 | 24B | 17.6 GB needed | 19.6 GB |
| Aria | 25B | 18.3 GB needed | 20.3 GB |
| Gemma 4 26B-A4B | 26B | 19 GB needed | 21 GB |
| Gemma 4 (all sizes) | 26B | 19 GB needed | 21 GB |
| Gemma 3 27B | 27B | 19.8 GB needed | 21.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | 19.8 GB needed | 21.8 GB |
How to read this
The AMD RX 6800 graphics card features 16 GB GDDR6 of dedicated video memory. This memory size determines which artificial intelligence models you can run entirely on your hardware. When a model fits completely inside this video memory, it processes tokens at maximum speed. If a model exceeds this limit, you must use alternative execution methods.
The quant column shows the specific quantization level recommended for each model. Quantization reduces the size of a model by compressing its numerical weights. For example, the 21B gpt-oss-20b and Reka Flash 3 models fit in 15.4 GB used with the Q4_K_M quant. The 12B Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, and FLUX.1 Kontext dev models use the Q8_0 quant, requiring 15.3 GB used.
Larger models require different quantization levels to fit within the 16 GB limit. The 17B HunyuanImage 2.1 / 3.0 and the 16.8B Ling-Coder-Lite models run at the Q5_K_M quant, using 14.5 GB and 14.3 GB respectively. Many models run at the Q6_K quant, such as the 16B DeepSeek-Coder-V2 16B / 236B and Kimi-VL A3B at 15.7 GB used, the 15B Apriel-1.5-15B-Thinker and StarCoder2 3B / 7B / 15B at 14.8 GB used, and the 14.7B Qwen2.5 14B at 15.3 GB used.
Other models also utilize the Q6_K quant to stay under the memory ceiling. The 14B Phi-3 Medium, Phi-4, Phi-4-reasoning / -plus, Wan 2.2 T2I, Wan 2.1 (1.3B / 14B), and SkyReels V2 models all use 13.8 GB. The 13B Vicuna 13B, HunyuanVideo, HunyuanVideo-Avatar, LTX-Video / LTX-2, and FramePack models use 12.8 GB. The 12B FLUX.1 dev model uses an FP8 / optimized quant requiring 14.4 GB used, while the 20B Qwen-Image and Qwen-Image-Edit models use Q4_K_M at 14.6 GB used, and CogVLM2 uses Q4_K_M at 13.9 GB used.
When a model is too large for the video memory, you can offload parts of it to your system RAM. This offload process allows you to run larger models but reduces processing speed significantly. For these cases, we assume your computer has 32 GB of system RAM. The 22B Solar Pro and Codestral 22B models need 16.1 GB at Q4_K_M, requiring 18.1 GB system RAM. The 24B Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 models need 17.6 GB at Q4_K_M, requiring 19.6 GB system RAM.
Even larger models can run using system RAM offloading. The 25B Aria model needs 18.3 GB at Q4_K_M, requiring 20.3 GB system RAM. The 26B Gemma 4 26B-A4B and Gemma 4 (all sizes) models need 19 GB at Q4_K_M, requiring 21 GB system RAM. The 27B Gemma 3 27B and Gemma 3 4B/12B/27B (vision) models need 19.8 GB at Q4_K_M, requiring 21.8 GB system RAM.
All memory calculations in this guide assume a standard 4k context window. If you increase the context window to process longer texts, the model will require more video memory. Running close to the 16 GB limit of your AMD RX 6800 means that large context windows might cause the system to run out of memory or force slow CPU offloading.