Best local AI models for NVIDIA RTX PRO 5000 Blackwell

48 GB GDDR7. At a 4k context, 186 of the 233 models in our catalog with verified parameter counts fit fully, up to Jamba 1.5 Mini / Large at 52B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 186 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Jamba 1.5 Mini / Large52BQ5_K_M44.3 GB
Llama 3.1 Nemotron 51B51BQ5_K_M43.5 GB
Mixtral 8x7B47BQ6_K46.2 GB
Seed-OSS 36B36BQ8_045.8 GB
Qwen3.6-35B-A3B35BQ8_044.5 GB
Command R (35B)35BQ8_044.5 GB
Yi 1.5 9B / 34B34BQ8_043.2 GB
Granite Code 3B to 34B34BQ8_043.2 GB
LLaVA 1.5 / 1.6 (7B to 34B)34BQ8_043.2 GB
Ovis 234BQ8_043.2 GB
DeepSeek-Coder 1.3B / 6.7B / 33B33BQ8_042 GB
WizardCoder 33B33BQ8_042 GB
OTel 2.0 LLM 31B IT32.1BQ8_044.9 GB
Qwen3 8B / 14B / 32B32BQ8_040.7 GB
Qwen3.5 (dense variants)32BQ8_040.7 GB
Aya Expanse 8B / 32B32BQ8_040.7 GB
Granite 4.0 Small/Tiny32BQ8_040.7 GB
Qwen2.5-Coder 0.5B to 32B32BQ8_040.7 GB
Qwen3-30B-A3B30BQ8_038.2 GB
Qwen3-Coder 30B-A3B30BQ8_038.2 GB
Gemma 3 27B27BQ8_034.3 GB
Gemma 3 4B/12B/27B (vision)27BQ8_034.3 GB
Wan 2.2 / 2.527BQ8_034.3 GB
Gemma 4 26B-A4B26BQ8_033.1 GB
Gemma 4 (all sizes)26BQ8_033.1 GB
Aria25BQ8_031.8 GB
Mistral Small 3.224BQ8_030.5 GB
Magistral Small24BQ8_030.5 GB
Devstral Small 1.124BQ8_030.5 GB
Solar Pro22BQ8_028 GB

How to read this

The NVIDIA RTX PRO 5000 Blackwell workstation graphics card features 48 GB of GDDR7 memory. This dedicated memory determines the size of the artificial intelligence models you can run locally. To run a model at native speed, the entire model must fit inside this 48 GB space. If a model exceeds this limit, it cannot run entirely on the fast graphics hardware.

The quantization column indicates the compression level applied to each model. Quantization reduces the size of model weights to save memory. A Q8_0 quant represents an eight bit quantization which preserves excellent output quality. For larger models like Jamba 1.5 Mini / Large 52B or Llama 3.1 Nemotron 51B, a Q5_K_M quant is used to fit the model into 44.3 GB and 43.5 GB of memory respectively.

Running models within the 48 GB limit avoids CPU offloading entirely. There are no CPU offload cases listed for this hardware configuration with 32 GB of system RAM. Keeping the model fully inside the GDDR7 memory ensures you get the maximum generation speed. Offloading parts of a model to system RAM would slow down the processing speed significantly.

The memory usage figures are calculated using a standard 4k context window. If you increase the context window to process longer documents, the memory usage will rise. You must leave some of the 48 GB memory free to handle this extra context data. For example, running Mixtral 8x7B at Q6_K uses 46.2 GB, which leaves very little room for context expansion.

Many high performance models fit comfortably on this hardware. You can run Qwen3.6-35B-A3B or Command R (35B) at Q8_0 quant using 44.5 GB of memory. Other options include Yi 1.5 34B, Granite Code 34B, LLaVA 1.5 / 1.6 34B, and Ovis 2, which all require 43.2 GB of memory at Q8_0. DeepSeek-Coder 33B and WizardCoder 33B both fit at Q8_0 using 42 GB.

For users needing more memory headroom for long conversations, slightly smaller models are ideal. Qwen3 32B, Qwen3.5 dense variants, Aya Expanse 32B, Granite 4.0 Small/Tiny 32B, and Qwen2.5-Coder 32B all use 40.7 GB of memory at Q8_0. Gemma 3 27B and Wan 2.2 / 2.5 27B use 34.3 GB at Q8_0, leaving ample space for large context demands.