Best local AI models for AMD PRO W7800

32 GB GDDR6. At a 4k context, 183 of the 233 models in our catalog with verified parameter counts fit fully, up to Seed-OSS 36B at 36B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 183 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Seed-OSS 36B36BQ5_K_M30.7 GB
Qwen3.6-35B-A3B35BQ5_K_M29.8 GB
Command R (35B)35BQ5_K_M29.8 GB
Yi 1.5 9B / 34B34BQ5_K_M29 GB
Granite Code 3B to 34B34BQ5_K_M29 GB
LLaVA 1.5 / 1.6 (7B to 34B)34BQ5_K_M29 GB
Ovis 234BQ5_K_M29 GB
DeepSeek-Coder 1.3B / 6.7B / 33B33BQ5_K_M28.1 GB
WizardCoder 33B33BQ5_K_M28.1 GB
OTel 2.0 LLM 31B IT32.1BQ5_K_M31.4 GB
Qwen3 8B / 14B / 32B32BQ6_K31.5 GB
Qwen3.5 (dense variants)32BQ6_K31.5 GB
Aya Expanse 8B / 32B32BQ6_K31.5 GB
Granite 4.0 Small/Tiny32BQ6_K31.5 GB
Qwen2.5-Coder 0.5B to 32B32BQ6_K31.5 GB
Qwen3-30B-A3B30BQ6_K29.5 GB
Qwen3-Coder 30B-A3B30BQ6_K29.5 GB
Gemma 3 27B27BQ6_K26.6 GB
Gemma 3 4B/12B/27B (vision)27BQ6_K26.6 GB
Wan 2.2 / 2.527BQ6_K26.6 GB
Gemma 4 26B-A4B26BQ6_K25.6 GB
Gemma 4 (all sizes)26BQ6_K25.6 GB
Aria25BQ8_031.8 GB
Mistral Small 3.224BQ8_030.5 GB
Magistral Small24BQ8_030.5 GB
Devstral Small 1.124BQ8_030.5 GB
Solar Pro22BQ8_028 GB
Codestral 22B22BQ8_028 GB
gpt-oss-20b21BQ8_026.7 GB
Reka Flash 321BQ8_026.7 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at Q4_K_MSystem RAM at 4k
Mixtral 8x7B47B34.4 GB needed36.4 GB
Llama 3.1 Nemotron 51B51B37.3 GB needed39.3 GB
Jamba 1.5 Mini / Large52B38.1 GB needed40.1 GB

How to read this

The AMD Radeon PRO W7800 workstation graphics card comes equipped with 32 GB of GDDR6 dedicated video memory. This onboard VRAM determines the maximum size of the artificial intelligence models you can run entirely on the hardware. Keeping the model files within this 32 GB boundary ensures fast processing speeds because the GPU can access the model weights directly without waiting on slower system memory.

The quantization column indicates the compression level applied to each model. Quantization reduces the precision of the model weights to make the files smaller. For this hardware, the best quant for models like Seed-OSS 36B, Qwen3.6-35B-A3B, and Command R (35B) is Q5_K_M, which uses 30.7 GB, 29.8 GB, and 29.8 GB of VRAM respectively. Other models like Yi 1.5 9B / 34B, Granite Code 3B to 34B, LLaVA 1.5 / 1.6 (7B to 34B), and Ovis 2 also run at Q5_K_M using 29 GB of memory.

Slightly smaller models can run at higher precision levels. DeepSeek-Coder 1.3B / 6.7B / 33B and WizardCoder 33B fit at Q5_K_M using 28.1 GB. The OTel 2.0 LLM 31B IT fits at Q5_K_M using 31.4 GB. Models such as Qwen3 8B / 14B / 32B, Qwen3.5 (dense variants), Aya Expanse 8B / 32B, Granite 4.0 Small/Tiny, and Qwen2.5-Coder 0.5B to 32B can run at the Q6_K quantization level using 31.5 GB of VRAM.

Other models running at Q6_K include Qwen3-30B-A3B and Qwen3-Coder 30B-A3B using 29.5 GB. Gemma 3 27B, Gemma 3 4B/12B/27B (vision), and Wan 2.2 / 2.5 use 26.6 GB. Gemma 4 26B-A4B and Gemma 4 (all sizes) use 25.6 GB. For maximum precision, Aria uses 31.8 GB at Q8_0. Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 use 30.5 GB at Q8_0. Solar Pro and Codestral 22B use 28 GB at Q8_0, while gpt-oss-20b and Reka Flash 3 use 26.7 GB at Q8_0.

When a model is too large for the 32 GB VRAM, you must offload parts of it to your system CPU. This offloading process allows you to run larger models but reduces the processing speed significantly. For example, Mixtral 8x7B requires 34.4 GB of memory at Q4_K_M and needs 36.4 GB of system RAM. Llama 3.1 Nemotron 51B requires 37.3 GB at Q4_K_M and needs 39.3 GB of system RAM. Jamba 1.5 Mini / Large requires 38.1 GB at Q4_K_M and needs 40.1 GB of system RAM.

All VRAM calculations assume a standard 4k context window. If you increase the context window to process longer documents or chat histories, the memory usage will increase. This extra memory demand might force you to choose a lower quantization level or offload layers to the CPU to avoid running out of video memory.