Best local AI models for NVIDIA Quadro GV100

32 GB HBM2. At a 4k context, 183 of the 233 models in our catalog with verified parameter counts fit fully, up to Seed-OSS 36B at 36B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 183 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Seed-OSS 36B36BQ5_K_M30.7 GB
Qwen3.6-35B-A3B35BQ5_K_M29.8 GB
Command R (35B)35BQ5_K_M29.8 GB
Yi 1.5 9B / 34B34BQ5_K_M29 GB
Granite Code 3B to 34B34BQ5_K_M29 GB
LLaVA 1.5 / 1.6 (7B to 34B)34BQ5_K_M29 GB
Ovis 234BQ5_K_M29 GB
DeepSeek-Coder 1.3B / 6.7B / 33B33BQ5_K_M28.1 GB
WizardCoder 33B33BQ5_K_M28.1 GB
OTel 2.0 LLM 31B IT32.1BQ5_K_M31.4 GB
Qwen3 8B / 14B / 32B32BQ6_K31.5 GB
Qwen3.5 (dense variants)32BQ6_K31.5 GB
Aya Expanse 8B / 32B32BQ6_K31.5 GB
Granite 4.0 Small/Tiny32BQ6_K31.5 GB
Qwen2.5-Coder 0.5B to 32B32BQ6_K31.5 GB
Qwen3-30B-A3B30BQ6_K29.5 GB
Qwen3-Coder 30B-A3B30BQ6_K29.5 GB
Gemma 3 27B27BQ6_K26.6 GB
Gemma 3 4B/12B/27B (vision)27BQ6_K26.6 GB
Wan 2.2 / 2.527BQ6_K26.6 GB
Gemma 4 26B-A4B26BQ6_K25.6 GB
Gemma 4 (all sizes)26BQ6_K25.6 GB
Aria25BQ8_031.8 GB
Mistral Small 3.224BQ8_030.5 GB
Magistral Small24BQ8_030.5 GB
Devstral Small 1.124BQ8_030.5 GB
Solar Pro22BQ8_028 GB
Codestral 22B22BQ8_028 GB
gpt-oss-20b21BQ8_026.7 GB
Reka Flash 321BQ8_026.7 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at Q4_K_MSystem RAM at 4k
Mixtral 8x7B47B34.4 GB needed36.4 GB
Llama 3.1 Nemotron 51B51B37.3 GB needed39.3 GB
Jamba 1.5 Mini / Large52B38.1 GB needed40.1 GB

How to read this

The NVIDIA Quadro GV100 is equipped with 32 GB of high speed HBM2 memory. This onboard memory determines the maximum size of the artificial intelligence models you can run locally. To run a model entirely on the graphics hardware, the model files and the active memory space must fit within this 32 GB limit. Keeping the entire model on the card ensures the fastest processing speeds.

Quantization is a method that compresses model files to save space. In our lists, the best quant column shows the highest quality compression level that still fits within your hardware limits. For example, the Seed-OSS 36B model fits at the Q5_K_M quantization level which uses 30.7 GB of memory. Other models like the Qwen3.6-35B-A3B and Command R (35B) also run at the Q5_K_M level using 29.8 GB of memory.

Several high performance models fit well within the 32 GB limit of your card. The Yi 1.5 9B / 34B, Granite Code 3B to 34B, LLaVA 1.5 / 1.6 (7B to 34B), and Ovis 2 models all fit at the Q5_K_M level using 29 GB of memory. You can also run DeepSeek-Coder 1.3B / 6.7B / 33B and WizardCoder 33B at Q5_K_M using 28.1 GB of memory. The OTel 2.0 LLM 31B IT model fits at Q5_K_M using 31.4 GB of memory.

If you want higher quality quantization levels, you can select slightly smaller models. The Qwen3 8B / 14B / 32B, Qwen3.5 (dense variants), Aya Expanse 8B / 32B, Granite 4.0 Small/Tiny, and Qwen2.5-Coder 0.5B to 32B models all run at the Q6_K level using 31.5 GB of memory. The Qwen3-30B-A3B and Qwen3-Coder 30B-A3B models use 29.5 GB of memory at Q6_K. Gemma 3 27B, Gemma 3 4B/12B/27B (vision), and Wan 2.2 / 2.5 use 26.6 GB at Q6_K. Gemma 4 26B-A4B and Gemma 4 (all sizes) use 25.6 GB at Q6_K.

For maximum precision, some models can run at the Q8_0 level. Aria uses 31.8 GB of memory at Q8_0. Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 use 30.5 GB of memory at Q8_0. Solar Pro and Codestral 22B use 28 GB of memory at Q8_0. The gpt-oss-20b and Reka Flash 3 models use 26.7 GB of memory at Q8_0.

When a model is too large for the 32 GB card, you can offload parts of it to your system CPU and system RAM. This offloading allows you to run larger models but it reduces your processing speed. For example, Mixtral 8x7B needs 34.4 GB at Q4_K_M and requires 36.4 GB of system RAM. Llama 3.1 Nemotron 51B needs 37.3 GB at Q4_K_M and requires 39.3 GB of system RAM. Jamba 1.5 Mini / Large needs 38.1 GB at Q4_K_M and requires 40.1 GB of system RAM.

All memory calculations are based on a standard 4k context window. The context window is the amount of text the model can read and write at one time. If you increase this context window, the model will require more memory. This extra memory usage might force you to use a lower quantization level or offload data to your system RAM.