Best local AI models for NVIDIA RTX PRO 2000 Blackwell
16 GB GDDR7. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| gpt-oss-20b | 21B | Q4_K_M | 15.4 GB |
| Reka Flash 3 | 21B | Q4_K_M | 15.4 GB |
| Qwen-Image | 20B | Q4_K_M | 14.6 GB |
| Qwen-Image-Edit | 20B | Q4_K_M | 14.6 GB |
| CogVLM2 | 19B | Q4_K_M | 13.9 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q5_K_M | 14.5 GB |
| Ling-Coder-Lite | 16.8B | Q5_K_M | 14.3 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q6_K | 15.7 GB |
| Kimi-VL A3B | 16B | Q6_K | 15.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q6_K | 14.8 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q6_K | 14.8 GB |
| Qwen2.5 14B | 14.7B | Q6_K | 15.3 GB |
| Phi-3 Medium | 14B | Q6_K | 13.8 GB |
| Phi-4 | 14B | Q6_K | 13.8 GB |
| Phi-4-reasoning / -plus | 14B | Q6_K | 13.8 GB |
| Wan 2.2 T2I | 14B | Q6_K | 13.8 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q6_K | 13.8 GB |
| SkyReels V2 | 14B | Q6_K | 13.8 GB |
| Vicuna 13B | 13B | Q6_K | 12.8 GB |
| HunyuanVideo | 13B | Q6_K | 12.8 GB |
| HunyuanVideo-Avatar | 13B | Q6_K | 12.8 GB |
| LTX-Video / LTX-2 | 13B | Q6_K | 12.8 GB |
| FramePack | 13B | Q6_K | 12.8 GB |
| FLUX.1 dev | 12B | FP8 / optimized | 14.4 GB |
| Gemma 3 12B | 12B | Q8_0 | 15.3 GB |
| Gemma 4 12B | 12B | Q8_0 | 15.3 GB |
| Mistral NeMo 12B | 12B | Q8_0 | 15.3 GB |
| Pixtral 12B | 12B | Q8_0 | 15.3 GB |
| FLUX.1 schnell | 12B | Q8_0 | 15.3 GB |
| FLUX.1 Kontext dev | 12B | Q8_0 | 15.3 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
| Mistral Small 3.2 | 24B | 17.6 GB needed | 19.6 GB |
| Magistral Small | 24B | 17.6 GB needed | 19.6 GB |
| Devstral Small 1.1 | 24B | 17.6 GB needed | 19.6 GB |
| Aria | 25B | 18.3 GB needed | 20.3 GB |
| Gemma 4 26B-A4B | 26B | 19 GB needed | 21 GB |
| Gemma 4 (all sizes) | 26B | 19 GB needed | 21 GB |
| Gemma 3 27B | 27B | 19.8 GB needed | 21.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | 19.8 GB needed | 21.8 GB |
How to read this
The NVIDIA RTX PRO 2000 Blackwell workstation graphics card features 16 GB of GDDR7 memory. This dedicated memory determines the maximum size of the artificial intelligence models you can run locally. To load a model entirely on the graphics card, the model files and the active memory space must fit within this 16 GB limit. Running models fully on the graphics card ensures the fastest processing speeds.
The quantization column indicates the compression level applied to each model. Raw models are often too large for local hardware. Quantization reduces the precision of the model weights to save space. For example, the Q4_K_M quant represents a medium four bit quantization. The Q6_K and Q8_0 quants offer higher precision but require more memory. Choosing the best quantization level balances model intelligence with available hardware memory.
Several large models fit completely within the 16 GB memory limit of this card. You can run the 21B gpt-oss-20b and Reka Flash 3 models at Q4_K_M quantization using 15.4 GB of memory. Vision models like Qwen-Image and Qwen-Image-Edit at 20B fit at Q4_K_M using 14.6 GB. The 19B CogVLM2 model fits at Q4_K_M using 13.9 GB. For higher precision, you can run the 12B Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, and FLUX.1 Kontext dev models at Q8_0 quantization using 15.3 GB of memory.
When a model exceeds the 16 GB graphics memory, you can offload parts of the model to your system RAM. This process requires at least 32 GB of system RAM. For example, the 22B Solar Pro and Codestral 22B models need 16.1 GB at Q4_K_M quantization and require 18.1 GB of system RAM. The 24B Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 models need 17.6 GB at Q4_K_M quantization and require 19.6 GB of system RAM. Offloading allows you to run larger models like the Gemma 3 27B which needs 19.8 GB at Q4_K_M and requires 21.8 GB of system RAM.
Offloading models to system RAM comes with a performance cost. System RAM is much slower than the GDDR7 memory on the graphics card. When parts of the model run on the system processor and RAM, the generation speed drops significantly. For the best user experience, keep your primary models fully loaded within the local graphics memory.
All listed memory figures are calculated using a standard 4k context window. The context window is the amount of text the model can read and write at one time. If you increase the context window beyond 4k tokens, the model will require more memory. This extra memory usage might force you to use a lower quantization level or offload more data to your system RAM.