Best local AI models for NVIDIA Quadro P5000

16 GB GDDR5X. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
gpt-oss-20b21BQ4_K_M15.4 GB
Reka Flash 321BQ4_K_M15.4 GB
Qwen-Image20BQ4_K_M14.6 GB
Qwen-Image-Edit20BQ4_K_M14.6 GB
CogVLM219BQ4_K_M13.9 GB
HunyuanImage 2.1 / 3.017BQ5_K_M14.5 GB
Ling-Coder-Lite16.8BQ5_K_M14.3 GB
DeepSeek-Coder-V2 16B / 236B16BQ6_K15.7 GB
Kimi-VL A3B16BQ6_K15.7 GB
Apriel-1.5-15B-Thinker15BQ6_K14.8 GB
StarCoder2 3B / 7B / 15B15BQ6_K14.8 GB
Qwen2.5 14B14.7BQ6_K15.3 GB
Phi-3 Medium14BQ6_K13.8 GB
Phi-414BQ6_K13.8 GB
Phi-4-reasoning / -plus14BQ6_K13.8 GB
Wan 2.2 T2I14BQ6_K13.8 GB
Wan 2.1 (1.3B / 14B)14BQ6_K13.8 GB
SkyReels V214BQ6_K13.8 GB
Vicuna 13B13BQ6_K12.8 GB
HunyuanVideo13BQ6_K12.8 GB
HunyuanVideo-Avatar13BQ6_K12.8 GB
LTX-Video / LTX-213BQ6_K12.8 GB
FramePack13BQ6_K12.8 GB
FLUX.1 dev12BFP8 / optimized14.4 GB
Gemma 3 12B12BQ8_015.3 GB
Gemma 4 12B12BQ8_015.3 GB
Mistral NeMo 12B12BQ8_015.3 GB
Pixtral 12B12BQ8_015.3 GB
FLUX.1 schnell12BQ8_015.3 GB
FLUX.1 Kontext dev12BQ8_015.3 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at Q4_K_MSystem RAM at 4k
Solar Pro22B16.1 GB needed18.1 GB
Codestral 22B22B16.1 GB needed18.1 GB
Mistral Small 3.224B17.6 GB needed19.6 GB
Magistral Small24B17.6 GB needed19.6 GB
Devstral Small 1.124B17.6 GB needed19.6 GB
Aria25B18.3 GB needed20.3 GB
Gemma 4 26B-A4B26B19 GB needed21 GB
Gemma 4 (all sizes)26B19 GB needed21 GB
Gemma 3 27B27B19.8 GB needed21.8 GB
Gemma 3 4B/12B/27B (vision)27B19.8 GB needed21.8 GB

How to read this

The NVIDIA Quadro P5000 is equipped with 16 GB of GDDR5X frame buffer. This dedicated memory capacity determines the size of the artificial intelligence models you can run entirely on the graphics hardware. Keeping the model parameters inside this video memory is critical for maintaining acceptable generation speeds.

When selecting a model, the quantization level indicates how much the original model weights are compressed. A higher quantization level like Q8_0 or Q6_K preserves more of the original model accuracy. Lower quantization levels like Q5_K_M or Q4_K_M compress the weights further to fit larger models into the 16 GB limit.

For maximum performance inside the local video memory, you can run models up to 21B parameters. The gpt-oss-20b and Reka Flash 3 models fit at Q4_K_M quantization using 15.4 GB of video memory. You can also run Qwen-Image at Q4_K_M using 14.6 GB or HunyuanImage 2.1 / 3.0 at Q5_K_M using 14.5 GB. Highly accurate Q8_0 models like Gemma 3 12B, Gemma 4 12B, and Mistral NeMo 12B fit comfortably using 15.3 GB of video memory.

If you want to run larger models, you must offload some layers to your system memory. This process requires at least 32 GB of system RAM. For example, running Solar Pro or Codestral 22B at Q4_K_M requires 16.1 GB of memory, which uses 18.1 GB of system RAM. Running Gemma 3 27B at Q4_K_M requires 19.8 GB of memory, which uses 21.8 GB of system RAM. Offloading layers to system RAM prevents out of memory errors but significantly reduces processing speeds.

You must also account for the context window when calculating memory usage. The listed memory figures represent the model weights at a standard 4k context limit. Increasing the context window to process longer documents or chat histories will require additional video memory. If you run out of video memory due to a large context window, the system will slow down or fail to generate text.