Best local AI models for NVIDIA RTX PRO 5000 72 GB Blackwell
72 GB GDDR7. At a 4k context, 186 of the 233 models in our catalog with verified parameter counts fit fully, up to Jamba 1.5 Mini / Large at 52B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 186 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Jamba 1.5 Mini / Large | 52B | Q8_0 | 66.1 GB |
| Llama 3.1 Nemotron 51B | 51B | Q8_0 | 64.9 GB |
| Mixtral 8x7B | 47B | Q8_0 | 59.8 GB |
| Seed-OSS 36B | 36B | Q8_0 | 45.8 GB |
| Qwen3.6-35B-A3B | 35B | Q8_0 | 44.5 GB |
| Command R (35B) | 35B | Q8_0 | 44.5 GB |
| Yi 1.5 9B / 34B | 34B | Q8_0 | 43.2 GB |
| Granite Code 3B to 34B | 34B | Q8_0 | 43.2 GB |
| LLaVA 1.5 / 1.6 (7B to 34B) | 34B | Q8_0 | 43.2 GB |
| Ovis 2 | 34B | Q8_0 | 43.2 GB |
| DeepSeek-Coder 1.3B / 6.7B / 33B | 33B | Q8_0 | 42 GB |
| WizardCoder 33B | 33B | Q8_0 | 42 GB |
| OTel 2.0 LLM 31B IT | 32.1B | Q8_0 | 44.9 GB |
| Qwen3 8B / 14B / 32B | 32B | Q8_0 | 40.7 GB |
| Qwen3.5 (dense variants) | 32B | Q8_0 | 40.7 GB |
| Aya Expanse 8B / 32B | 32B | Q8_0 | 40.7 GB |
| Granite 4.0 Small/Tiny | 32B | Q8_0 | 40.7 GB |
| Qwen2.5-Coder 0.5B to 32B | 32B | Q8_0 | 40.7 GB |
| Qwen3-30B-A3B | 30B | FP16 | 72 GB |
| Qwen3-Coder 30B-A3B | 30B | FP16 | 72 GB |
| Gemma 3 27B | 27B | FP16 | 64.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | FP16 | 64.8 GB |
| Wan 2.2 / 2.5 | 27B | FP16 | 64.8 GB |
| Gemma 4 26B-A4B | 26B | FP16 | 62.4 GB |
| Gemma 4 (all sizes) | 26B | FP16 | 62.4 GB |
| Aria | 25B | FP16 | 60 GB |
| Mistral Small 3.2 | 24B | FP16 | 57.6 GB |
| Magistral Small | 24B | FP16 | 57.6 GB |
| Devstral Small 1.1 | 24B | FP16 | 57.6 GB |
| Solar Pro | 22B | FP16 | 52.8 GB |
How to read this
The NVIDIA RTX PRO 5000 Blackwell graphics card features 72 GB of GDDR7 memory. This dedicated memory determines the maximum size of the artificial intelligence model you can run locally. When a model fits entirely within this onboard memory, the system achieves the fastest possible processing speeds.
The quantization column indicates the compression level applied to each model. For example, the Jamba 1.5 Mini or Large 52B model runs best at the Q8_0 quantization level, which uses 66.1 GB of memory. Similarly, the Llama 3.1 Nemotron 51B model fits at Q8_0 quantization and uses 64.9 GB of memory. Other models like the Mixtral 8x7B 47B model use 59.8 GB at Q8_0 quantization.
Several models can run at full precision without compression. The Qwen3-30B-A3B and Qwen3-Coder 30B-A3B models run at the FP16 quantization level, utilizing the full 72 GB of memory. The Gemma 3 27B model, Gemma 3 27B vision model, and Wan 2.2 or 2.5 models run at FP16 quantization using 64.8 GB of memory. Gemma 4 26B-A4B and other Gemma 4 variants use 62.4 GB at FP16 quantization.
Other notable options include the Seed-OSS 36B model using 45.8 GB at Q8_0 quantization. The Qwen3.6-35B-A3B and Command R 35B models both use 44.5 GB at Q8_0 quantization. The Yi 1.5 34B, Granite Code 34B, LLaVA 1.5 or 1.6 34B, and Ovis 2 models all require 43.2 GB at Q8_0 quantization. The DeepSeek-Coder 33B and WizardCoder 33B models use 42 GB at Q8_0 quantization.
For smaller models, the OTel 2.0 LLM 31B IT uses 44.9 GB at Q8_0 quantization. The Qwen3 32B, Qwen3.5 dense variants, Aya Expanse 32B, Granite 4.0 Small or Tiny 32B, and Qwen2.5-Coder 32B models all use 40.7 GB at Q8_0 quantization. The Aria 25B model uses 60 GB at FP16 quantization. Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 use 57.6 GB at FP16 quantization, while Solar Pro uses 52.8 GB at FP16 quantization.
Offloading parts of a model to system RAM is not required for these configurations because they fit within the graphics memory. Offloading to system RAM always introduces a speed penalty because system memory is slower than GDDR7 memory. All memory calculations are based on a standard 4k context window, and increasing this context window will require more memory.