Best local AI models for AMD FirePro W9100
16 GB GDDR5. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| gpt-oss-20b | 21B | Q4_K_M | 15.4 GB |
| Reka Flash 3 | 21B | Q4_K_M | 15.4 GB |
| Qwen-Image | 20B | Q4_K_M | 14.6 GB |
| Qwen-Image-Edit | 20B | Q4_K_M | 14.6 GB |
| CogVLM2 | 19B | Q4_K_M | 13.9 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q5_K_M | 14.5 GB |
| Ling-Coder-Lite | 16.8B | Q5_K_M | 14.3 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q6_K | 15.7 GB |
| Kimi-VL A3B | 16B | Q6_K | 15.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q6_K | 14.8 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q6_K | 14.8 GB |
| Qwen2.5 14B | 14.7B | Q6_K | 15.3 GB |
| Phi-3 Medium | 14B | Q6_K | 13.8 GB |
| Phi-4 | 14B | Q6_K | 13.8 GB |
| Phi-4-reasoning / -plus | 14B | Q6_K | 13.8 GB |
| Wan 2.2 T2I | 14B | Q6_K | 13.8 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q6_K | 13.8 GB |
| SkyReels V2 | 14B | Q6_K | 13.8 GB |
| Vicuna 13B | 13B | Q6_K | 12.8 GB |
| HunyuanVideo | 13B | Q6_K | 12.8 GB |
| HunyuanVideo-Avatar | 13B | Q6_K | 12.8 GB |
| LTX-Video / LTX-2 | 13B | Q6_K | 12.8 GB |
| FramePack | 13B | Q6_K | 12.8 GB |
| FLUX.1 dev | 12B | FP8 / optimized | 14.4 GB |
| Gemma 3 12B | 12B | Q8_0 | 15.3 GB |
| Gemma 4 12B | 12B | Q8_0 | 15.3 GB |
| Mistral NeMo 12B | 12B | Q8_0 | 15.3 GB |
| Pixtral 12B | 12B | Q8_0 | 15.3 GB |
| FLUX.1 schnell | 12B | Q8_0 | 15.3 GB |
| FLUX.1 Kontext dev | 12B | Q8_0 | 15.3 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
| Mistral Small 3.2 | 24B | 17.6 GB needed | 19.6 GB |
| Magistral Small | 24B | 17.6 GB needed | 19.6 GB |
| Devstral Small 1.1 | 24B | 17.6 GB needed | 19.6 GB |
| Aria | 25B | 18.3 GB needed | 20.3 GB |
| Gemma 4 26B-A4B | 26B | 19 GB needed | 21 GB |
| Gemma 4 (all sizes) | 26B | 19 GB needed | 21 GB |
| Gemma 3 27B | 27B | 19.8 GB needed | 21.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | 19.8 GB needed | 21.8 GB |
How to read this
The AMD FirePro W9100 is equipped with 16 GB of GDDR5 memory. This onboard memory determines the maximum size of the artificial intelligence models you can run locally. To fit a model entirely within this hardware limit, the total memory footprint of the model must remain under 16 GB. Running models fully inside the graphics memory ensures the fastest possible processing speeds.
Quantization is a method that compresses model files to save space. The quant column shows the best compression level that fits your hardware. For example, the 21B models gpt-oss-20b and Reka Flash 3 both fit at a Q4_K_M quantization which uses 15.4 GB of memory. Vision models like Qwen-Image and Qwen-Image-Edit fit at Q4_K_M using 14.6 GB of memory. CogVLM2 fits at Q4_K_M using 13.9 GB of memory.
Higher quality quantizations are available for slightly smaller models. HunyuanImage 2.1 / 3.0 fits at Q5_K_M using 14.5 GB of memory. Ling-Coder-Lite fits at Q5_K_M using 14.3 GB of memory. You can run DeepSeek-Coder-V2 16B / 236B, Kimi-VL A3B, Apriel-1.5-15B-Thinker, and StarCoder2 3B / 7B / 15B at Q6_K quantization. These models use between 14.8 GB and 15.7 GB of memory.
Other models like Qwen2.5 14B, Phi-3 Medium, Phi-4, Phi-4-reasoning / -plus, Wan 2.2 T2I, Wan 2.1 (1.3B / 14B), SkyReels V2, Vicuna 13B, HunyuanVideo, HunyuanVideo-Avatar, LTX-Video / LTX-2, and FramePack also run at Q6_K quantization. For maximum precision, you can run Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, and FLUX.1 Kontext dev at Q8_0 quantization using 15.3 GB of memory. FLUX.1 dev fits at FP8 / optimized using 14.4 GB of memory.
When a model is too large for the 16 GB graphics card, you can use CPU offload if you have 32 GB of system RAM. This process splits the model between your graphics card and your system memory. Offloading allows you to run larger models but it reduces your processing speed. For example, Solar Pro and Codestral 22B need 16.1 GB at Q4_K_M and require 18.1 GB of system RAM. Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 need 17.6 GB at Q4_K_M and require 19.6 GB of system RAM.
Larger offload options include Aria which needs 18.3 GB at Q4_K_M and requires 20.3 GB of system RAM. Gemma 4 26B-A4B and Gemma 4 (all sizes) need 19 GB at Q4_K_M and require 21 GB of system RAM. Gemma 3 27B and Gemma 3 4B/12B/27B (vision) need 19.8 GB at Q4_K_M and require 21.8 GB of system RAM. Note that all memory calculations are based on a standard 4k context window. Increasing the context window size will require more memory and may force you to use lower quantization levels.