Best local AI models for AMD Pro WX 9100
16 GB HBM2. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| gpt-oss-20b | 21B | Q4_K_M | 15.4 GB |
| Reka Flash 3 | 21B | Q4_K_M | 15.4 GB |
| Qwen-Image | 20B | Q4_K_M | 14.6 GB |
| Qwen-Image-Edit | 20B | Q4_K_M | 14.6 GB |
| CogVLM2 | 19B | Q4_K_M | 13.9 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q5_K_M | 14.5 GB |
| Ling-Coder-Lite | 16.8B | Q5_K_M | 14.3 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q6_K | 15.7 GB |
| Kimi-VL A3B | 16B | Q6_K | 15.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q6_K | 14.8 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q6_K | 14.8 GB |
| Qwen2.5 14B | 14.7B | Q6_K | 15.3 GB |
| Phi-3 Medium | 14B | Q6_K | 13.8 GB |
| Phi-4 | 14B | Q6_K | 13.8 GB |
| Phi-4-reasoning / -plus | 14B | Q6_K | 13.8 GB |
| Wan 2.2 T2I | 14B | Q6_K | 13.8 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q6_K | 13.8 GB |
| SkyReels V2 | 14B | Q6_K | 13.8 GB |
| Vicuna 13B | 13B | Q6_K | 12.8 GB |
| HunyuanVideo | 13B | Q6_K | 12.8 GB |
| HunyuanVideo-Avatar | 13B | Q6_K | 12.8 GB |
| LTX-Video / LTX-2 | 13B | Q6_K | 12.8 GB |
| FramePack | 13B | Q6_K | 12.8 GB |
| FLUX.1 dev | 12B | FP8 / optimized | 14.4 GB |
| Gemma 3 12B | 12B | Q8_0 | 15.3 GB |
| Gemma 4 12B | 12B | Q8_0 | 15.3 GB |
| Mistral NeMo 12B | 12B | Q8_0 | 15.3 GB |
| Pixtral 12B | 12B | Q8_0 | 15.3 GB |
| FLUX.1 schnell | 12B | Q8_0 | 15.3 GB |
| FLUX.1 Kontext dev | 12B | Q8_0 | 15.3 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
| Mistral Small 3.2 | 24B | 17.6 GB needed | 19.6 GB |
| Magistral Small | 24B | 17.6 GB needed | 19.6 GB |
| Devstral Small 1.1 | 24B | 17.6 GB needed | 19.6 GB |
| Aria | 25B | 18.3 GB needed | 20.3 GB |
| Gemma 4 26B-A4B | 26B | 19 GB needed | 21 GB |
| Gemma 4 (all sizes) | 26B | 19 GB needed | 21 GB |
| Gemma 3 27B | 27B | 19.8 GB needed | 21.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | 19.8 GB needed | 21.8 GB |
How to read this
The AMD Radeon Pro WX 9100 workstation graphics card features 16 GB of high bandwidth HBM2 memory. This dedicated video memory determines the maximum size of the artificial intelligence models you can run locally. To load a model entirely on the graphics hardware for fast processing, the model files and the active context must fit within this 16 GB limit.
The quantization column indicates the compression level applied to each model. Raw models are often too large for local hardware, so they are quantized to reduce their footprint. For example, the 21B models gpt-oss-20b and Reka Flash 3 fit within 15.4 GB of video memory when using the Q4_K_M quantization. Other models like the 20B Qwen-Image and Qwen-Image-Edit require 14.6 GB of memory at the same Q4_K_M level.
As model sizes decrease, you can use higher quality quantizations. The 17B HunyuanImage 2.1 / 3.0 fits in 14.5 GB of memory using a Q5_K_M quantization, while the 16.8B Ling-Coder-Lite uses 14.3 GB. For even better precision, you can run the 16B DeepSeek-Coder-V2 16B / 236B or Kimi-VL A3B at Q6_K quantization, which uses 15.7 GB of video memory. The 15B Apriel-1.5-15B-Thinker and StarCoder2 3B / 7B / 15B also run at Q6_K quantization using 14.8 GB.
Several 14B models run efficiently at Q6_K quantization using 13.8 GB of memory. These include Phi-3 Medium, Phi-4, Phi-4-reasoning / -plus, Wan 2.2 T2I, Wan 2.1 (1.3B / 14B), and SkyReels V2. The 13B models like Vicuna 13B, HunyuanVideo, HunyuanVideo-Avatar, LTX-Video / LTX-2, and FramePack use 12.8 GB at Q6_K. Highly precise Q8_0 quantizations are available for 12B models like Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, and FLUX.1 Kontext dev, which require 15.3 GB of memory.
When a model exceeds the 16 GB video memory limit, you can offload some layers to your system RAM. This requires a system with at least 32 GB of system memory. For example, Solar Pro 22B and Codestral 22B need 16.1 GB of memory at Q4_K_M quantization and require 18.1 GB of system RAM. Larger models like Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 need 17.6 GB at Q4_K_M and require 19.6 GB of system RAM.
Other offload options include Aria 25B, which needs 18.3 GB at Q4_K_M and uses 20.3 GB of system RAM. The 26B models Gemma 4 26B-A4B and Gemma 4 (all sizes) need 19 GB at Q4_K_M and use 21 GB of system RAM. The largest offload options are Gemma 3 27B and Gemma 3 4B/12B/27B (vision), which need 19.8 GB at Q4_K_M and use 21.8 GB of system RAM. Offloading layers to system memory allows you to run these larger models, but it reduces processing speed.
All memory calculations are based on a standard 4k context window. If you increase the context window to process longer documents or larger chat histories, the memory usage will increase. You must leave enough free video memory on your AMD Pro WX 9100 to accommodate this extra context data.