Best local AI models for AMD RX 6800M

12 GB GDDR6. At a 4k context, 147 of the 233 models in our catalog with verified parameter counts fit fully, up to DeepSeek-Coder-V2 16B / 236B at 16B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 147 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
DeepSeek-Coder-V2 16B / 236B16BQ4_K_M11.7 GB
Kimi-VL A3B16BQ4_K_M11.7 GB
Apriel-1.5-15B-Thinker15BQ4_K_M11 GB
StarCoder2 3B / 7B / 15B15BQ4_K_M11 GB
Qwen2.5 14B14.7BQ4_K_M11.6 GB
Phi-3 Medium14BQ5_K_M11.9 GB
Phi-414BQ5_K_M11.9 GB
Phi-4-reasoning / -plus14BQ5_K_M11.9 GB
Wan 2.2 T2I14BQ5_K_M11.9 GB
Wan 2.1 (1.3B / 14B)14BQ5_K_M11.9 GB
SkyReels V214BQ5_K_M11.9 GB
Vicuna 13B13BQ5_K_M11.1 GB
HunyuanVideo13BQ5_K_M11.1 GB
HunyuanVideo-Avatar13BQ5_K_M11.1 GB
LTX-Video / LTX-213BQ5_K_M11.1 GB
FramePack13BQ5_K_M11.1 GB
Gemma 3 12B12BQ6_K11.8 GB
Gemma 4 12B12BQ6_K11.8 GB
Mistral NeMo 12B12BQ6_K11.8 GB
Pixtral 12B12BQ6_K11.8 GB
FLUX.1 schnell12BQ6_K11.8 GB
FLUX.1 Kontext dev12BQ6_K11.8 GB
FLUX.1 Krea dev12BQ6_K11.8 GB
Open-Sora 2.011BQ6_K10.8 GB
Mochi 110BQ6_K9.8 GB
Gemma 2 9B9BQ6_K10.3 GB
Nemotron Nano 4B / 9B9BQ8_011.4 GB
GLM-4 9B / GLM-4.5-Air9BQ8_011.4 GB
Yi-Coder 1.5B / 9B9BQ8_011.4 GB
GLM-4-9B-Chat / CodeGeeX49BQ8_011.4 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
FLUX.1 dev12B14.4 GB needed16.4 GB
Ling-Coder-Lite16.8B12.3 GB needed14.3 GB
HunyuanImage 2.1 / 3.017B12.4 GB needed14.4 GB
CogVLM219B13.9 GB needed15.9 GB
Qwen-Image20B14.6 GB needed16.6 GB
Qwen-Image-Edit20B14.6 GB needed16.6 GB
gpt-oss-20b21B15.4 GB needed17.4 GB
Reka Flash 321B15.4 GB needed17.4 GB
Solar Pro22B16.1 GB needed18.1 GB
Codestral 22B22B16.1 GB needed18.1 GB

How to read this

The AMD Radeon RX 6800M is a mobile graphics processor equipped with 12 GB of GDDR6 memory. This physical memory limit determines which artificial intelligence models can run entirely on your hardware. For local execution, the model weights must fit inside this video memory space to ensure fast processing speeds. If a model exceeds this limit, the system must transfer data between your graphics card and system memory.

The quantization column indicates the compression level applied to each model. Quantization reduces the precision of model weights to save space. A Q4_K_M quant represents a four bit quantization level which offers a balanced mix of size and accuracy. Higher quants like Q5_K_M, Q6_K, and Q8_0 provide better output quality but require more memory. For example, the 14B parameter Phi-4 model fits within 11.9 GB of memory when using a Q5_K_M quant.

Models like DeepSeek-Coder-V2 16B and Kimi-VL A3B utilize 11.7 GB of memory at Q4_K_M. This leaves very little remaining space in your 12 GB video memory. When running these large models, you must consider the context window. Standard calculations assume a base 4k context window. Expanding your context window beyond this limit requires additional memory for the key value cache, which can cause out of memory errors on a 12 GB card.

When a model is too large for the 12 GB video memory, you can use CPU offloading. This technique splits the workload between your graphics card and your system RAM. For CPU offloading, we assume your computer has 32 GB of system RAM. Offloading allows you to run larger models, but it significantly reduces generation speed because system RAM is much slower than GDDR6 video memory.

Using CPU offloading, you can run the 12B parameter FLUX.1 dev model which needs 14.4 GB of space at FP8 and uses 16.4 GB of system RAM. You can also run the 22B parameter Codestral 22B model. This model requires 16.1 GB of space at Q4_K_M and utilizes 18.1 GB of system RAM. Other offload options include the 20B parameter Qwen-Image model and the 21B parameter Reka Flash 3 model.