Best local AI models for AMD Pro W6800X
32 GB GDDR6. At a 4k context, 183 of the 233 models in our catalog with verified parameter counts fit fully, up to Seed-OSS 36B at 36B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 183 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Seed-OSS 36B | 36B | Q5_K_M | 30.7 GB |
| Qwen3.6-35B-A3B | 35B | Q5_K_M | 29.8 GB |
| Command R (35B) | 35B | Q5_K_M | 29.8 GB |
| Yi 1.5 9B / 34B | 34B | Q5_K_M | 29 GB |
| Granite Code 3B to 34B | 34B | Q5_K_M | 29 GB |
| LLaVA 1.5 / 1.6 (7B to 34B) | 34B | Q5_K_M | 29 GB |
| Ovis 2 | 34B | Q5_K_M | 29 GB |
| DeepSeek-Coder 1.3B / 6.7B / 33B | 33B | Q5_K_M | 28.1 GB |
| WizardCoder 33B | 33B | Q5_K_M | 28.1 GB |
| OTel 2.0 LLM 31B IT | 32.1B | Q5_K_M | 31.4 GB |
| Qwen3 8B / 14B / 32B | 32B | Q6_K | 31.5 GB |
| Qwen3.5 (dense variants) | 32B | Q6_K | 31.5 GB |
| Aya Expanse 8B / 32B | 32B | Q6_K | 31.5 GB |
| Granite 4.0 Small/Tiny | 32B | Q6_K | 31.5 GB |
| Qwen2.5-Coder 0.5B to 32B | 32B | Q6_K | 31.5 GB |
| Qwen3-30B-A3B | 30B | Q6_K | 29.5 GB |
| Qwen3-Coder 30B-A3B | 30B | Q6_K | 29.5 GB |
| Gemma 3 27B | 27B | Q6_K | 26.6 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | Q6_K | 26.6 GB |
| Wan 2.2 / 2.5 | 27B | Q6_K | 26.6 GB |
| Gemma 4 26B-A4B | 26B | Q6_K | 25.6 GB |
| Gemma 4 (all sizes) | 26B | Q6_K | 25.6 GB |
| Aria | 25B | Q8_0 | 31.8 GB |
| Mistral Small 3.2 | 24B | Q8_0 | 30.5 GB |
| Magistral Small | 24B | Q8_0 | 30.5 GB |
| Devstral Small 1.1 | 24B | Q8_0 | 30.5 GB |
| Solar Pro | 22B | Q8_0 | 28 GB |
| Codestral 22B | 22B | Q8_0 | 28 GB |
| gpt-oss-20b | 21B | Q8_0 | 26.7 GB |
| Reka Flash 3 | 21B | Q8_0 | 26.7 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Mixtral 8x7B | 47B | 34.4 GB needed | 36.4 GB |
| Llama 3.1 Nemotron 51B | 51B | 37.3 GB needed | 39.3 GB |
| Jamba 1.5 Mini / Large | 52B | 38.1 GB needed | 40.1 GB |
How to read this
The AMD Radeon Pro W6800X is a workstation graphics card equipped with 32 GB of GDDR6 memory. This dedicated video memory determines the maximum size of the artificial intelligence models you can run locally. To achieve optimal processing speeds, the entire model must fit directly inside this onboard memory. If a model exceeds this capacity, execution slows down significantly.
Model quantization is a method that compresses neural networks to save memory. In our catalog, the best quantization column shows the highest quality format that fits within your hardware limits. For this card, models in the 33B to 36B parameter range run best using the Q5_K_M quantization level. Examples include Seed-OSS 36B using 30.7 GB, Qwen3.6-35B-A3B using 29.8 GB, and Command R (35B) using 29.8 GB. You can also run Yi 1.5 34B, Granite Code 34B, LLaVA 1.5 / 1.6 34B, and Ovis 2 at Q5_K_M, with each using 29 GB of memory. DeepSeek-Coder 33B and WizardCoder 33B fit at Q5_K_M using 28.1 GB.
Models in the 25B to 32B range can run at higher precision levels. You can run OTel 2.0 LLM 31B IT at Q5_K_M using 31.4 GB. The Q6_K quantization level is available for Qwen3 32B, Qwen3.5 dense variants, Aya Expanse 32B, Granite 4.0 Small/Tiny, and Qwen2.5-Coder 32B, which all use 31.5 GB. Qwen3-30B-A3B and Qwen3-Coder 30B-A3B use 29.5 GB at Q6_K. Gemma 3 27B, Gemma 3 vision 27B, and Wan 2.2 / 2.5 use 26.6 GB at Q6_K. Gemma 4 26B-A4B and Gemma 4 all sizes use 25.6 GB at Q6_K.
For maximum output quality, you can run models up to 25B parameters using the uncompromised Q8_0 quantization level. Aria fits this category using 31.8 GB. Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 use 30.5 GB of memory at Q8_0. Solar Pro and Codestral 22B use 28 GB at Q8_0. The gpt-oss-20b model and Reka Flash 3 use 26.7 GB at Q8_0.
When a model is too large for the 32 GB of onboard graphics memory, you must offload parts of it to your system CPU and system RAM. This offloading process allows you to run larger models but reduces processing speed. For example, Mixtral 8x7B requires 34.4 GB at Q4_K_M and needs 36.4 GB of system RAM. Llama 3.1 Nemotron 51B requires 37.3 GB at Q4_K_M and needs 39.3 GB of system RAM. Jamba 1.5 Mini / Large requires 38.1 GB at Q4_K_M and needs 40.1 GB of system RAM.
All memory calculations are based on a standard 4k context window. Generating longer responses or processing larger documents increases memory consumption. If you increase the context window beyond 4k tokens, you will need to select smaller models or lower quantization levels to prevent running out of memory.