Best local AI models for AMD RX 6850M XT

12 GB GDDR6. At a 4k context, 147 of the 233 models in our catalog with verified parameter counts fit fully, up to DeepSeek-Coder-V2 16B / 236B at 16B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 147 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
DeepSeek-Coder-V2 16B / 236B16BQ4_K_M11.7 GB
Kimi-VL A3B16BQ4_K_M11.7 GB
Apriel-1.5-15B-Thinker15BQ4_K_M11 GB
StarCoder2 3B / 7B / 15B15BQ4_K_M11 GB
Qwen2.5 14B14.7BQ4_K_M11.6 GB
Phi-3 Medium14BQ5_K_M11.9 GB
Phi-414BQ5_K_M11.9 GB
Phi-4-reasoning / -plus14BQ5_K_M11.9 GB
Wan 2.2 T2I14BQ5_K_M11.9 GB
Wan 2.1 (1.3B / 14B)14BQ5_K_M11.9 GB
SkyReels V214BQ5_K_M11.9 GB
Vicuna 13B13BQ5_K_M11.1 GB
HunyuanVideo13BQ5_K_M11.1 GB
HunyuanVideo-Avatar13BQ5_K_M11.1 GB
LTX-Video / LTX-213BQ5_K_M11.1 GB
FramePack13BQ5_K_M11.1 GB
Gemma 3 12B12BQ6_K11.8 GB
Gemma 4 12B12BQ6_K11.8 GB
Mistral NeMo 12B12BQ6_K11.8 GB
Pixtral 12B12BQ6_K11.8 GB
FLUX.1 schnell12BQ6_K11.8 GB
FLUX.1 Kontext dev12BQ6_K11.8 GB
FLUX.1 Krea dev12BQ6_K11.8 GB
Open-Sora 2.011BQ6_K10.8 GB
Mochi 110BQ6_K9.8 GB
Gemma 2 9B9BQ6_K10.3 GB
Nemotron Nano 4B / 9B9BQ8_011.4 GB
GLM-4 9B / GLM-4.5-Air9BQ8_011.4 GB
Yi-Coder 1.5B / 9B9BQ8_011.4 GB
GLM-4-9B-Chat / CodeGeeX49BQ8_011.4 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
FLUX.1 dev12B14.4 GB needed16.4 GB
Ling-Coder-Lite16.8B12.3 GB needed14.3 GB
HunyuanImage 2.1 / 3.017B12.4 GB needed14.4 GB
CogVLM219B13.9 GB needed15.9 GB
Qwen-Image20B14.6 GB needed16.6 GB
Qwen-Image-Edit20B14.6 GB needed16.6 GB
gpt-oss-20b21B15.4 GB needed17.4 GB
Reka Flash 321B15.4 GB needed17.4 GB
Solar Pro22B16.1 GB needed18.1 GB
Codestral 22B22B16.1 GB needed18.1 GB

How to read this

The AMD Radeon RX 6850M XT is a mobile graphics processor equipped with 12 GB of GDDR6 dedicated video memory. This memory capacity determines the maximum size of the artificial intelligence models you can run entirely on your hardware. To execute local models efficiently, the model weights and the active context window must fit within this 12 GB limit. Running out of video memory causes the system to slow down significantly.

Quantization is a method that compresses model files to save space. The quant column indicates the optimal compression level for each model on this hardware. For example, a Q4_K_M quant represents a four bit medium quantization, while Q6_K and Q8_0 represent six bit and eight bit options. Higher quants preserve more original model accuracy but require more video memory. Lower quants like Q4_K_M allow larger models to fit inside your 12 GB limit.

Several high performance models fit completely within the video memory of the RX 6850M XT. The DeepSeek-Coder-V2 16B and Kimi-VL A3B models fit at Q4_K_M quant, using 11.7 GB of memory. The Phi-4 and Phi-4-reasoning models fit at Q5_K_M quant, using 11.9 GB of memory. For image and video tasks, the FLUX.1 schnell and Mistral NeMo 12B models run at Q6_K quant, using 11.8 GB of memory. The GLM-4 9B model can run at a high quality Q8_0 quant, using 11.4 GB of memory.

When a model is too large for the 12 GB video memory, you can use CPU offloading. This process splits the model weights between your video memory and your system RAM. For instance, the FLUX.1 dev model requires 14.4 GB at FP8 and needs 16.4 GB of system RAM. The Codestral 22B model requires 16.1 GB at Q4_K_M and needs 18.1 GB of system RAM. While offloading allows you to run larger models like CogVLM2 or Solar Pro, it reduces processing speed because system RAM is slower than GDDR6 memory.

You must also consider the memory required for the context window. The memory figures listed are calculated using a baseline 4k context window. If you increase the context window to process longer documents or chat histories, the model will require more memory. To prevent crashes or slow performance on your RX 6850M XT, you may need to choose a smaller model or a lower quantization level when working with large context sizes.