Best local AI models for AMD Pro W6800X Duo

32 GB GDDR6. At a 4k context, 183 of the 233 models in our catalog with verified parameter counts fit fully, up to Seed-OSS 36B at 36B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 183 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Seed-OSS 36B36BQ5_K_M30.7 GB
Qwen3.6-35B-A3B35BQ5_K_M29.8 GB
Command R (35B)35BQ5_K_M29.8 GB
Yi 1.5 9B / 34B34BQ5_K_M29 GB
Granite Code 3B to 34B34BQ5_K_M29 GB
LLaVA 1.5 / 1.6 (7B to 34B)34BQ5_K_M29 GB
Ovis 234BQ5_K_M29 GB
DeepSeek-Coder 1.3B / 6.7B / 33B33BQ5_K_M28.1 GB
WizardCoder 33B33BQ5_K_M28.1 GB
OTel 2.0 LLM 31B IT32.1BQ5_K_M31.4 GB
Qwen3 8B / 14B / 32B32BQ6_K31.5 GB
Qwen3.5 (dense variants)32BQ6_K31.5 GB
Aya Expanse 8B / 32B32BQ6_K31.5 GB
Granite 4.0 Small/Tiny32BQ6_K31.5 GB
Qwen2.5-Coder 0.5B to 32B32BQ6_K31.5 GB
Qwen3-30B-A3B30BQ6_K29.5 GB
Qwen3-Coder 30B-A3B30BQ6_K29.5 GB
Gemma 3 27B27BQ6_K26.6 GB
Gemma 3 4B/12B/27B (vision)27BQ6_K26.6 GB
Wan 2.2 / 2.527BQ6_K26.6 GB
Gemma 4 26B-A4B26BQ6_K25.6 GB
Gemma 4 (all sizes)26BQ6_K25.6 GB
Aria25BQ8_031.8 GB
Mistral Small 3.224BQ8_030.5 GB
Magistral Small24BQ8_030.5 GB
Devstral Small 1.124BQ8_030.5 GB
Solar Pro22BQ8_028 GB
Codestral 22B22BQ8_028 GB
gpt-oss-20b21BQ8_026.7 GB
Reka Flash 321BQ8_026.7 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at Q4_K_MSystem RAM at 4k
Mixtral 8x7B47B34.4 GB needed36.4 GB
Llama 3.1 Nemotron 51B51B37.3 GB needed39.3 GB
Jamba 1.5 Mini / Large52B38.1 GB needed40.1 GB

How to read this

The AMD Pro W6800X Duo graphics card features 32 GB GDDR6 memory. This memory capacity determines the size of the artificial intelligence models you can run locally. To run a model entirely on the graphics hardware, the model files and its working memory must fit within this 32 GB limit.

The quantization column indicates the compression level applied to each model. Quantization reduces the size of a model so it uses less memory. A Q5_K_M quant represents a high quality five bit quantization. A Q6_K quant represents a six bit quantization. A Q8_0 quant represents an eight bit quantization which preserves more original model accuracy but requires more memory.

For maximum performance, you can run models like Seed-OSS 36B at Q5_K_M which uses 30.7 GB of memory. You can also run Qwen3.6-35B-A3B or Command R (35B) at Q5_K_M using 29.8 GB of memory. Other options include Yi 1.5 9B / 34B, Granite Code 3B to 34B, LLaVA 1.5 / 1.6 (7B to 34B), and Ovis 2 at Q5_K_M using 29 GB of memory.

Models like Qwen3 8B / 14B / 32B, Qwen3.5 (dense variants), Aya Expanse 8B / 32B, Granite 4.0 Small/Tiny, and Qwen2.5-Coder 0.5B to 32B run at Q6_K using 31.5 GB of memory. You can also run Gemma 3 27B, Gemma 3 4B/12B/27B (vision), and Wan 2.2 / 2.5 at Q6_K using 26.6 GB of memory. Aria runs at Q8_0 using 31.8 GB of memory.

When a model exceeds the 32 GB graphics memory, you must offload parts of the model to your system RAM and CPU. CPU offload allows you to run larger models but it reduces processing speed. For example, Mixtral 8x7B needs 34.4 GB at Q4_K_M and requires 36.4 GB system RAM. Llama 3.1 Nemotron 51B needs 37.3 GB at Q4_K_M and requires 39.3 GB system RAM. Jamba 1.5 Mini / Large needs 38.1 GB at Q4_K_M and requires 40.1 GB system RAM.

All memory calculations in this guide assume a standard 4k context window. If you increase the context window to process longer documents, the model will require more memory. This extra memory usage might require you to choose a lower quantization level or use CPU offload.