Best local AI models for NVIDIA RTX PRO 5000 72 GB Blackwell

72 GB GDDR7. At a 4k context, 186 of the 233 models in our catalog with verified parameter counts fit fully, up to Jamba 1.5 Mini / Large at 52B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 186 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Jamba 1.5 Mini / Large52BQ8_066.1 GB
Llama 3.1 Nemotron 51B51BQ8_064.9 GB
Mixtral 8x7B47BQ8_059.8 GB
Seed-OSS 36B36BQ8_045.8 GB
Qwen3.6-35B-A3B35BQ8_044.5 GB
Command R (35B)35BQ8_044.5 GB
Yi 1.5 9B / 34B34BQ8_043.2 GB
Granite Code 3B to 34B34BQ8_043.2 GB
LLaVA 1.5 / 1.6 (7B to 34B)34BQ8_043.2 GB
Ovis 234BQ8_043.2 GB
DeepSeek-Coder 1.3B / 6.7B / 33B33BQ8_042 GB
WizardCoder 33B33BQ8_042 GB
OTel 2.0 LLM 31B IT32.1BQ8_044.9 GB
Qwen3 8B / 14B / 32B32BQ8_040.7 GB
Qwen3.5 (dense variants)32BQ8_040.7 GB
Aya Expanse 8B / 32B32BQ8_040.7 GB
Granite 4.0 Small/Tiny32BQ8_040.7 GB
Qwen2.5-Coder 0.5B to 32B32BQ8_040.7 GB
Qwen3-30B-A3B30BFP1672 GB
Qwen3-Coder 30B-A3B30BFP1672 GB
Gemma 3 27B27BFP1664.8 GB
Gemma 3 4B/12B/27B (vision)27BFP1664.8 GB
Wan 2.2 / 2.527BFP1664.8 GB
Gemma 4 26B-A4B26BFP1662.4 GB
Gemma 4 (all sizes)26BFP1662.4 GB
Aria25BFP1660 GB
Mistral Small 3.224BFP1657.6 GB
Magistral Small24BFP1657.6 GB
Devstral Small 1.124BFP1657.6 GB
Solar Pro22BFP1652.8 GB

How to read this

The NVIDIA RTX PRO 5000 Blackwell graphics card features 72 GB of GDDR7 memory. This dedicated memory determines the maximum size of the artificial intelligence model you can run locally. When a model fits entirely within this onboard memory, the system achieves the fastest possible processing speeds.

The quantization column indicates the compression level applied to each model. For example, the Jamba 1.5 Mini or Large 52B model runs best at the Q8_0 quantization level, which uses 66.1 GB of memory. Similarly, the Llama 3.1 Nemotron 51B model fits at Q8_0 quantization and uses 64.9 GB of memory. Other models like the Mixtral 8x7B 47B model use 59.8 GB at Q8_0 quantization.

Several models can run at full precision without compression. The Qwen3-30B-A3B and Qwen3-Coder 30B-A3B models run at the FP16 quantization level, utilizing the full 72 GB of memory. The Gemma 3 27B model, Gemma 3 27B vision model, and Wan 2.2 or 2.5 models run at FP16 quantization using 64.8 GB of memory. Gemma 4 26B-A4B and other Gemma 4 variants use 62.4 GB at FP16 quantization.

Other notable options include the Seed-OSS 36B model using 45.8 GB at Q8_0 quantization. The Qwen3.6-35B-A3B and Command R 35B models both use 44.5 GB at Q8_0 quantization. The Yi 1.5 34B, Granite Code 34B, LLaVA 1.5 or 1.6 34B, and Ovis 2 models all require 43.2 GB at Q8_0 quantization. The DeepSeek-Coder 33B and WizardCoder 33B models use 42 GB at Q8_0 quantization.

For smaller models, the OTel 2.0 LLM 31B IT uses 44.9 GB at Q8_0 quantization. The Qwen3 32B, Qwen3.5 dense variants, Aya Expanse 32B, Granite 4.0 Small or Tiny 32B, and Qwen2.5-Coder 32B models all use 40.7 GB at Q8_0 quantization. The Aria 25B model uses 60 GB at FP16 quantization. Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 use 57.6 GB at FP16 quantization, while Solar Pro uses 52.8 GB at FP16 quantization.

Offloading parts of a model to system RAM is not required for these configurations because they fit within the graphics memory. Offloading to system RAM always introduces a speed penalty because system memory is slower than GDDR7 memory. All memory calculations are based on a standard 4k context window, and increasing this context window will require more memory.