Best local AI models for AMD FirePro W9000
6 GB GDDR5. At a 4k context, 114 of the 233 models in our catalog with verified parameter counts fit fully, up to Granite 3.3 2B / 8B at 8B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 114 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Granite 3.3 2B / 8B | 8B | Q4_K_M | 5.9 GB |
| Ministral 3B / 8B | 8B | Q4_K_M | 5.9 GB |
| InternLM 3 8B | 8B | Q4_K_M | 5.9 GB |
| OpenCoder 1.5B / 8B | 8B | Q4_K_M | 5.9 GB |
| Seed-Coder 8B | 8B | Q4_K_M | 5.9 GB |
| MiniCPM-V 2.6 / MiniCPM-o 2.6 | 8B | Q4_K_M | 5.9 GB |
| Idefics 3 8B | 8B | Q4_K_M | 5.9 GB |
| Fuyu-8B | 8B | Q4_K_M | 5.9 GB |
| Emu3 | 8B | Q4_K_M | 5.9 GB |
| Stable Diffusion 3.5 Large / Turbo | 8B | Q4_K_M | 5.9 GB |
| EXAONE 3.5 2.4B / 7.8B | 7.8B | Q4_K_M | 5.7 GB |
| Mistral 7B | 7B | Q4_K_M | 5.7 GB |
| Qwen2.5 0.5B / 1.5B / 3B / 7B | 7B | Q5_K_M | 6 GB |
| OLMo 2 1B / 7B | 7B | Q5_K_M | 6 GB |
| Falcon 3 1B / 3B / 7B | 7B | Q5_K_M | 6 GB |
| Command R7B | 7B | Q5_K_M | 6 GB |
| OpenHermes 2.5 | 7B | Q5_K_M | 6 GB |
| Zephyr 7B Beta | 7B | Q5_K_M | 6 GB |
| OpenChat 3.5 | 7B | Q5_K_M | 6 GB |
| Starling LM 7B | 7B | Q5_K_M | 6 GB |
| Codestral Mamba 7B | 7B | Q5_K_M | 6 GB |
| CodeGemma 2B / 7B | 7B | Q5_K_M | 6 GB |
| aiXcoder-7B | 7B | Q5_K_M | 6 GB |
| Nxcode / CodeQwen 1.5 7B | 7B | Q5_K_M | 6 GB |
| Janus-Pro 1B / 7B | 7B | Q5_K_M | 6 GB |
| Ruyi-Mini-7B | 7B | Q5_K_M | 6 GB |
| Qwen2-Audio 7B | 7B | Q5_K_M | 6 GB |
| Qwen2.5-Omni 3B / 7B | 7B | Q5_K_M | 6 GB |
| YuE | 7B | Q5_K_M | 6 GB |
| Magicoder-S-DS 6.7B | 6.7B | Q5_K_M | 5.7 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Llama 3.1 8B | 8B | 6.4 GB needed | 8.4 GB |
| Chroma | 8.9B | 6.5 GB needed | 8.5 GB |
| Gemma 2 9B | 9B | 8 GB needed | 10 GB |
| Nemotron Nano 4B / 9B | 9B | 6.6 GB needed | 8.6 GB |
| GLM-4 9B / GLM-4.5-Air | 9B | 6.6 GB needed | 8.6 GB |
| Yi-Coder 1.5B / 9B | 9B | 6.6 GB needed | 8.6 GB |
| GLM-4-9B-Chat / CodeGeeX4 | 9B | 6.6 GB needed | 8.6 GB |
| GLM-4V-9B / GLM-4.1V-Thinking | 9B | 6.6 GB needed | 8.6 GB |
| Mochi 1 | 10B | 7.3 GB needed | 9.3 GB |
| Open-Sora 2.0 | 11B | 8.1 GB needed | 10.1 GB |
How to read this
The AMD FirePro W9000 is an enterprise workstation graphics card equipped with 6 GB of GDDR5 memory. This onboard memory capacity determines the maximum size of the artificial intelligence models you can run entirely on the hardware. When a model fits completely within this 6 GB limit, the graphics processor handles all calculations. This local execution ensures the fastest possible generation speeds without relying on external cloud servers.
To fit modern models into this memory footprint, you must use quantized versions. Quantization is a compression technique that reduces the precision of model weights. The quant column indicates the specific level of compression applied to the model. For example, the Q4_K_M quant represents a medium four bit quantization, while the Q5_K_M quant represents a five bit quantization. Higher quantization levels preserve more model intelligence but require more memory.
Several high quality models fit entirely within the 6 GB memory of the card. You can run the 8B versions of Granite 3.3, Ministral, InternLM 3, OpenCoder, Seed-Coder, MiniCPM-V 2.6, Idefics 3, Fuyu-8B, Emu3, and Stable Diffusion 3.5 at the Q4_K_M quant, which uses 5.9 GB of memory. The 7.8B version of EXAONE 3.5 and the 7B version of Mistral also fit at Q4_K_M, consuming 5.7 GB of memory. Magicoder-S-DS 6.7B fits at Q5_K_M while using 5.7 GB of memory.
Other 7B models can run at the Q5_K_M quant by utilizing exactly 6 GB of memory. These models include Qwen2.5, OLMo 2, Falcon 3, Command R7B, OpenHermes 2.5, Zephyr 7B Beta, OpenChat 3.5, Starling LM 7B, Codestral Mamba 7B, CodeGemma, aiXcoder-7B, Nxcode, Janus-Pro, Ruyi-Mini-7B, Qwen2-Audio, Qwen2.5-Omni, and YuE. Running these models at 6 GB leaves no spare room in the onboard memory.
If a model exceeds the 6 GB onboard memory, you can offload the remaining parts to your system RAM. This setup assumes you have 32 GB of system RAM available. For example, Llama 3.1 8B requires 6.4 GB at Q4_K_M, which uses 8.4 GB of system RAM for offloading. Chroma requires 6.5 GB at Q4_K_M and uses 8.5 GB of system RAM. Gemma 2 9B requires 8 GB at Q4_K_M and uses 10 GB of system RAM. Offloading allows you to run larger models like Mochi 1 or Open-Sora 2.0, but it reduces processing speed because system RAM is slower than GDDR5 memory.
Other offload options include Nemotron Nano 9B, GLM-4 9B, Yi-Coder 9B, GLM-4-9B-Chat, and GLM-4V-9B. These models require 6.6 GB at Q4_K_M and use 8.6 GB of system RAM. All memory calculations for these models are based on a standard 4k context window. If you increase the context window to process longer documents, the memory requirements will rise, which may force you to use lower quants or offload more data to system RAM.