Best local AI models for NVIDIA Quadro RTX 5000 MAX-Q

16 GB GDDR6. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
gpt-oss-20b21BQ4_K_M15.4 GB
Reka Flash 321BQ4_K_M15.4 GB
Qwen-Image20BQ4_K_M14.6 GB
Qwen-Image-Edit20BQ4_K_M14.6 GB
CogVLM219BQ4_K_M13.9 GB
HunyuanImage 2.1 / 3.017BQ5_K_M14.5 GB
Ling-Coder-Lite16.8BQ5_K_M14.3 GB
DeepSeek-Coder-V2 16B / 236B16BQ6_K15.7 GB
Kimi-VL A3B16BQ6_K15.7 GB
Apriel-1.5-15B-Thinker15BQ6_K14.8 GB
StarCoder2 3B / 7B / 15B15BQ6_K14.8 GB
Qwen2.5 14B14.7BQ6_K15.3 GB
Phi-3 Medium14BQ6_K13.8 GB
Phi-414BQ6_K13.8 GB
Phi-4-reasoning / -plus14BQ6_K13.8 GB
Wan 2.2 T2I14BQ6_K13.8 GB
Wan 2.1 (1.3B / 14B)14BQ6_K13.8 GB
SkyReels V214BQ6_K13.8 GB
Vicuna 13B13BQ6_K12.8 GB
HunyuanVideo13BQ6_K12.8 GB
HunyuanVideo-Avatar13BQ6_K12.8 GB
LTX-Video / LTX-213BQ6_K12.8 GB
FramePack13BQ6_K12.8 GB
FLUX.1 dev12BFP8 / optimized14.4 GB
Gemma 3 12B12BQ8_015.3 GB
Gemma 4 12B12BQ8_015.3 GB
Mistral NeMo 12B12BQ8_015.3 GB
Pixtral 12B12BQ8_015.3 GB
FLUX.1 schnell12BQ8_015.3 GB
FLUX.1 Kontext dev12BQ8_015.3 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at Q4_K_MSystem RAM at 4k
Solar Pro22B16.1 GB needed18.1 GB
Codestral 22B22B16.1 GB needed18.1 GB
Mistral Small 3.224B17.6 GB needed19.6 GB
Magistral Small24B17.6 GB needed19.6 GB
Devstral Small 1.124B17.6 GB needed19.6 GB
Aria25B18.3 GB needed20.3 GB
Gemma 4 26B-A4B26B19 GB needed21 GB
Gemma 4 (all sizes)26B19 GB needed21 GB
Gemma 3 27B27B19.8 GB needed21.8 GB
Gemma 3 4B/12B/27B (vision)27B19.8 GB needed21.8 GB

How to read this

The NVIDIA Quadro RTX 5000 MAX-Q is a mobile workstation graphics card equipped with 16 GB of GDDR6 memory. This dedicated video memory determines the maximum size of the artificial intelligence models you can run locally. To load and run a model entirely on the graphics hardware, the model files and the active workspace must fit within this 16 GB limit.

The quantization column indicates the compression level applied to each model. Raw models are often too large for consumer hardware, so they are quantized to smaller bit widths. For example, the 21B models gpt-oss-20b and Reka Flash 3 fit within 15.4 GB of memory when compressed to the Q4_K_M quantization. Larger models like Gemma 3 12B and Mistral NeMo 12B can run at a higher quality Q8_0 quantization because they use 15.3 GB of memory.

When a model exceeds the 16 GB video memory limit, you can use CPU offloading if your computer has at least 32 GB of system RAM. This process splits the model between your graphics card and your system memory. Offloading allows you to run larger models like Solar Pro or Codestral 22B, which require 16.1 GB of memory at Q4_K_M quantization and 18.1 GB of system RAM. You can even run Mistral Small 3.2 or Devstral Small 1.1, which require 17.6 GB of memory at Q4_K_M and 19.6 GB of system RAM.

CPU offloading comes with a significant performance cost. Processing data across the system memory bus is much slower than processing it directly on the GDDR6 memory of the graphics card. While offloading makes it possible to run massive models like Gemma 3 27B or Gemma 4 26B-A4B, the generation speed will drop noticeably compared to running smaller models entirely on the graphics hardware.

You must also consider the memory required for context. The memory figures listed for these models, such as 13.8 GB for Phi-4 or 12.8 GB for Vicuna 13B, are measured with a standard 4k context window. If you increase the context window to process longer documents or chat histories, the memory usage will rise. Keeping a safety margin below 16 GB prevents your system from running out of memory during long conversations.