Best local AI models for AMD PRO W6800

32 GB GDDR6. At a 4k context, 183 of the 233 models in our catalog with verified parameter counts fit fully, up to Seed-OSS 36B at 36B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 183 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Seed-OSS 36B36BQ5_K_M30.7 GB
Qwen3.6-35B-A3B35BQ5_K_M29.8 GB
Command R (35B)35BQ5_K_M29.8 GB
Yi 1.5 9B / 34B34BQ5_K_M29 GB
Granite Code 3B to 34B34BQ5_K_M29 GB
LLaVA 1.5 / 1.6 (7B to 34B)34BQ5_K_M29 GB
Ovis 234BQ5_K_M29 GB
DeepSeek-Coder 1.3B / 6.7B / 33B33BQ5_K_M28.1 GB
WizardCoder 33B33BQ5_K_M28.1 GB
OTel 2.0 LLM 31B IT32.1BQ5_K_M31.4 GB
Qwen3 8B / 14B / 32B32BQ6_K31.5 GB
Qwen3.5 (dense variants)32BQ6_K31.5 GB
Aya Expanse 8B / 32B32BQ6_K31.5 GB
Granite 4.0 Small/Tiny32BQ6_K31.5 GB
Qwen2.5-Coder 0.5B to 32B32BQ6_K31.5 GB
Qwen3-30B-A3B30BQ6_K29.5 GB
Qwen3-Coder 30B-A3B30BQ6_K29.5 GB
Gemma 3 27B27BQ6_K26.6 GB
Gemma 3 4B/12B/27B (vision)27BQ6_K26.6 GB
Wan 2.2 / 2.527BQ6_K26.6 GB
Gemma 4 26B-A4B26BQ6_K25.6 GB
Gemma 4 (all sizes)26BQ6_K25.6 GB
Aria25BQ8_031.8 GB
Mistral Small 3.224BQ8_030.5 GB
Magistral Small24BQ8_030.5 GB
Devstral Small 1.124BQ8_030.5 GB
Solar Pro22BQ8_028 GB
Codestral 22B22BQ8_028 GB
gpt-oss-20b21BQ8_026.7 GB
Reka Flash 321BQ8_026.7 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at Q4_K_MSystem RAM at 4k
Mixtral 8x7B47B34.4 GB needed36.4 GB
Llama 3.1 Nemotron 51B51B37.3 GB needed39.3 GB
Jamba 1.5 Mini / Large52B38.1 GB needed40.1 GB

How to read this

The AMD Radeon PRO W6800 workstation graphics card features 32 GB of GDDR6 memory. This dedicated onboard memory determines the maximum size of the artificial intelligence models you can run locally. To achieve optimal processing speeds, the entire model must reside directly within this graphics memory. If a model exceeds this capacity, performance drops significantly.

Quantization is a method that compresses model files to save space. The quant column indicates the specific level of compression applied to each model. For this hardware, the best quant for models like Seed-OSS 36B, Qwen3.6-35B-A3B, and Command R (35B) is Q5_K_M. This compression level uses 30.7 GB of memory for Seed-OSS 36B and 29.8 GB for the others, which fits comfortably inside your limits.

Other models can run at higher precision levels. The Q6_K quant is ideal for Qwen3 32B, Qwen3.5 dense variants, Aya Expanse 32B, Granite 4.0 Small, and Qwen2.5-Coder 32B. These configurations require 31.5 GB of memory. Gemma 3 27B and Wan 2.5 also use the Q6_K quant, consuming 26.6 GB of memory. For models like Aria, Mistral Small 3.2, and Codestral 22B, you can run the uncompressed Q8_0 quant which uses up to 31.8 GB.

When a model is too large for the 32 GB of graphics memory, you must use CPU offloading. This process splits the workload between your graphics card and your system RAM. For example, Mixtral 8x7B requires 34.4 GB of memory at the Q4_K_M quant, which demands at least 36.4 GB of system RAM. Llama 3.1 Nemotron 51B requires 37.3 GB at Q4_K_M, needing 39.3 GB of system RAM. Offloading allows you to run these larger models, but it reduces generation speeds.

All memory calculations are based on a standard 4k context window. The context window is the amount of text the model can remember during a conversation. If you increase this context window beyond 4k tokens, the model will require more memory. This extra memory usage might force you to use a lower quant or rely on slower CPU offloading to prevent out of memory errors.