Best local AI models for NVIDIA RTX PRO 4000 Blackwell
24 GB GDDR7. At a 4k context, 173 of the 233 models in our catalog with verified parameter counts fit fully, up to Qwen3 8B / 14B / 32B at 32B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 173 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Qwen3 8B / 14B / 32B | 32B | Q4_K_M | 23.4 GB |
| Qwen3.5 (dense variants) | 32B | Q4_K_M | 23.4 GB |
| Aya Expanse 8B / 32B | 32B | Q4_K_M | 23.4 GB |
| Granite 4.0 Small/Tiny | 32B | Q4_K_M | 23.4 GB |
| Qwen2.5-Coder 0.5B to 32B | 32B | Q4_K_M | 23.4 GB |
| Qwen3-30B-A3B | 30B | Q4_K_M | 22 GB |
| Qwen3-Coder 30B-A3B | 30B | Q4_K_M | 22 GB |
| Gemma 3 27B | 27B | Q5_K_M | 23 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | Q5_K_M | 23 GB |
| Wan 2.2 / 2.5 | 27B | Q5_K_M | 23 GB |
| Gemma 4 26B-A4B | 26B | Q5_K_M | 22.2 GB |
| Gemma 4 (all sizes) | 26B | Q5_K_M | 22.2 GB |
| Aria | 25B | Q5_K_M | 21.3 GB |
| Mistral Small 3.2 | 24B | Q6_K | 23.6 GB |
| Magistral Small | 24B | Q6_K | 23.6 GB |
| Devstral Small 1.1 | 24B | Q6_K | 23.6 GB |
| Solar Pro | 22B | Q6_K | 21.6 GB |
| Codestral 22B | 22B | Q6_K | 21.6 GB |
| gpt-oss-20b | 21B | Q6_K | 20.7 GB |
| Reka Flash 3 | 21B | Q6_K | 20.7 GB |
| Qwen-Image | 20B | Q6_K | 19.7 GB |
| Qwen-Image-Edit | 20B | Q6_K | 19.7 GB |
| CogVLM2 | 19B | Q6_K | 18.7 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q8_0 | 21.6 GB |
| Ling-Coder-Lite | 16.8B | Q8_0 | 21.4 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q8_0 | 20.4 GB |
| Kimi-VL A3B | 16B | Q8_0 | 20.4 GB |
| Apriel-1.5-15B-Thinker | 15B | Q8_0 | 19.1 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q8_0 | 19.1 GB |
| Qwen2.5 14B | 14.7B | Q8_0 | 19.5 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| OTel 2.0 LLM 31B IT | 32.1B | 27.5 GB needed | 29.5 GB |
| DeepSeek-Coder 1.3B / 6.7B / 33B | 33B | 24.2 GB needed | 26.2 GB |
| WizardCoder 33B | 33B | 24.2 GB needed | 26.2 GB |
| Yi 1.5 9B / 34B | 34B | 24.9 GB needed | 26.9 GB |
| Granite Code 3B to 34B | 34B | 24.9 GB needed | 26.9 GB |
| LLaVA 1.5 / 1.6 (7B to 34B) | 34B | 24.9 GB needed | 26.9 GB |
| Ovis 2 | 34B | 24.9 GB needed | 26.9 GB |
| Qwen3.6-35B-A3B | 35B | 25.6 GB needed | 27.6 GB |
| Command R (35B) | 35B | 25.6 GB needed | 27.6 GB |
| Seed-OSS 36B | 36B | 26.4 GB needed | 28.4 GB |
How to read this
The NVIDIA RTX PRO 4000 Blackwell workstation graphics card features 24 GB of GDDR7 memory. This dedicated memory capacity determines which artificial intelligence models can run entirely on the hardware. When a model fits completely within this VRAM limit, it achieves the fastest possible processing speeds. If a model exceeds this limit, some data must spill over to your system memory.
To fit larger models into the 24 GB memory space, developers use quantization. The quant column indicates the compression level applied to the model weights. For example, the Q4_K_M quant represents a four bit quantization level. This compression reduces the memory footprint significantly while preserving most of the original model accuracy. Higher quants like Q5_K_M, Q6_K, and Q8_0 offer better precision but require more memory.
Several high performance models fit entirely within the VRAM of this card. The Qwen3 32B, Qwen3.5 dense 32B, Aya Expanse 32B, Granite 4.0 Small or Tiny 32B, and Qwen2.5-Coder 32B models all run locally at Q4_K_M quantization using 23.4 GB of memory. The Qwen3-30B-A3B and Qwen3-Coder 30B-A3B models fit at Q4_K_M using 22 GB. For slightly smaller models, you can run Gemma 3 27B, Gemma 3 vision 27B, and Wan 2.2 or 2.5 27B at Q5_K_M quantization using 23 GB. Gemma 4 26B-A4B and Gemma 4 all sizes 26B run at Q5_K_M using 22.2 GB, while Aria 25B fits at Q5_K_M using 21.3 GB.
You can also run highly precise models at Q6_K and Q8_0 quantization. Mistral Small 3.2 24B, Magistral Small 24B, and Devstral Small 1.1 24B fit at Q6_K using 23.6 GB. Solar Pro 22B and Codestral 22B run at Q6_K using 21.6 GB. The gpt-oss-20b 21B and Reka Flash 3 21B models use 20.7 GB at Q6_K. Qwen-Image 20B and Qwen-Image-Edit 20B use 19.7 GB at Q6_K. CogVLM2 19B uses 18.7 GB at Q6_K. At the Q8_0 level, HunyuanImage 2.1 or 3.0 17B uses 21.6 GB, Ling-Coder-Lite 16.8B uses 21.4 GB, DeepSeek-Coder-V2 16B uses 20.4 GB, Kimi-VL A3B 16B uses 20.4 GB, Apriel-1.5-15B-Thinker 15B uses 19.1 GB, StarCoder2 15B uses 19.1 GB, and Qwen2.5 14B uses 19.5 GB.
When a model is too large for the 24 GB VRAM, you can offload parts of it to your system RAM. This offload process allows you to run larger models but reduces generation speeds. Assuming a 32 GB system RAM setup, OTel 2.0 LLM 31B IT needs 27.5 GB at Q4_K_M and uses 29.5 GB of system RAM. DeepSeek-Coder 33B and WizardCoder 33B need 24.2 GB at Q4_K_M and use 26.2 GB of system RAM. Yi 1.5 34B, Granite Code 34B, LLaVA 1.5 or 1.6 34B, and Ovis 2 34B need 24.9 GB at Q4_K_M and use 26.9 GB of system RAM. Qwen3.6-35B-A3B and Command R 35B need 25.6 GB at Q4_K_M and use 27.6 GB of system RAM. Seed-OSS 36B needs 26.4 GB at Q4_K_M and uses 28.4 GB of system RAM.
All memory calculations listed here assume a standard 4k context window. If you increase the context window to handle longer documents, the memory requirements will grow. This extra context data can push a model that normally fits in VRAM into system RAM offloading.