Best local AI models for AMD RX 6850M XT
12 GB GDDR6. At a 4k context, 147 of the 233 models in our catalog with verified parameter counts fit fully, up to DeepSeek-Coder-V2 16B / 236B at 16B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 147 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| DeepSeek-Coder-V2 16B / 236B | 16B | Q4_K_M | 11.7 GB |
| Kimi-VL A3B | 16B | Q4_K_M | 11.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q4_K_M | 11 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q4_K_M | 11 GB |
| Qwen2.5 14B | 14.7B | Q4_K_M | 11.6 GB |
| Phi-3 Medium | 14B | Q5_K_M | 11.9 GB |
| Phi-4 | 14B | Q5_K_M | 11.9 GB |
| Phi-4-reasoning / -plus | 14B | Q5_K_M | 11.9 GB |
| Wan 2.2 T2I | 14B | Q5_K_M | 11.9 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q5_K_M | 11.9 GB |
| SkyReels V2 | 14B | Q5_K_M | 11.9 GB |
| Vicuna 13B | 13B | Q5_K_M | 11.1 GB |
| HunyuanVideo | 13B | Q5_K_M | 11.1 GB |
| HunyuanVideo-Avatar | 13B | Q5_K_M | 11.1 GB |
| LTX-Video / LTX-2 | 13B | Q5_K_M | 11.1 GB |
| FramePack | 13B | Q5_K_M | 11.1 GB |
| Gemma 3 12B | 12B | Q6_K | 11.8 GB |
| Gemma 4 12B | 12B | Q6_K | 11.8 GB |
| Mistral NeMo 12B | 12B | Q6_K | 11.8 GB |
| Pixtral 12B | 12B | Q6_K | 11.8 GB |
| FLUX.1 schnell | 12B | Q6_K | 11.8 GB |
| FLUX.1 Kontext dev | 12B | Q6_K | 11.8 GB |
| FLUX.1 Krea dev | 12B | Q6_K | 11.8 GB |
| Open-Sora 2.0 | 11B | Q6_K | 10.8 GB |
| Mochi 1 | 10B | Q6_K | 9.8 GB |
| Gemma 2 9B | 9B | Q6_K | 10.3 GB |
| Nemotron Nano 4B / 9B | 9B | Q8_0 | 11.4 GB |
| GLM-4 9B / GLM-4.5-Air | 9B | Q8_0 | 11.4 GB |
| Yi-Coder 1.5B / 9B | 9B | Q8_0 | 11.4 GB |
| GLM-4-9B-Chat / CodeGeeX4 | 9B | Q8_0 | 11.4 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| FLUX.1 dev | 12B | 14.4 GB needed | 16.4 GB |
| Ling-Coder-Lite | 16.8B | 12.3 GB needed | 14.3 GB |
| HunyuanImage 2.1 / 3.0 | 17B | 12.4 GB needed | 14.4 GB |
| CogVLM2 | 19B | 13.9 GB needed | 15.9 GB |
| Qwen-Image | 20B | 14.6 GB needed | 16.6 GB |
| Qwen-Image-Edit | 20B | 14.6 GB needed | 16.6 GB |
| gpt-oss-20b | 21B | 15.4 GB needed | 17.4 GB |
| Reka Flash 3 | 21B | 15.4 GB needed | 17.4 GB |
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
How to read this
The AMD Radeon RX 6850M XT is a mobile graphics processor equipped with 12 GB of GDDR6 dedicated video memory. This memory capacity determines the maximum size of the artificial intelligence models you can run entirely on your hardware. To execute local models efficiently, the model weights and the active context window must fit within this 12 GB limit. Running out of video memory causes the system to slow down significantly.
Quantization is a method that compresses model files to save space. The quant column indicates the optimal compression level for each model on this hardware. For example, a Q4_K_M quant represents a four bit medium quantization, while Q6_K and Q8_0 represent six bit and eight bit options. Higher quants preserve more original model accuracy but require more video memory. Lower quants like Q4_K_M allow larger models to fit inside your 12 GB limit.
Several high performance models fit completely within the video memory of the RX 6850M XT. The DeepSeek-Coder-V2 16B and Kimi-VL A3B models fit at Q4_K_M quant, using 11.7 GB of memory. The Phi-4 and Phi-4-reasoning models fit at Q5_K_M quant, using 11.9 GB of memory. For image and video tasks, the FLUX.1 schnell and Mistral NeMo 12B models run at Q6_K quant, using 11.8 GB of memory. The GLM-4 9B model can run at a high quality Q8_0 quant, using 11.4 GB of memory.
When a model is too large for the 12 GB video memory, you can use CPU offloading. This process splits the model weights between your video memory and your system RAM. For instance, the FLUX.1 dev model requires 14.4 GB at FP8 and needs 16.4 GB of system RAM. The Codestral 22B model requires 16.1 GB at Q4_K_M and needs 18.1 GB of system RAM. While offloading allows you to run larger models like CogVLM2 or Solar Pro, it reduces processing speed because system RAM is slower than GDDR6 memory.
You must also consider the memory required for the context window. The memory figures listed are calculated using a baseline 4k context window. If you increase the context window to process longer documents or chat histories, the model will require more memory. To prevent crashes or slow performance on your RX 6850M XT, you may need to choose a smaller model or a lower quantization level when working with large context sizes.