Best local AI models for NVIDIA Quadro GP100
16 GB HBM2. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| gpt-oss-20b | 21B | Q4_K_M | 15.4 GB |
| Reka Flash 3 | 21B | Q4_K_M | 15.4 GB |
| Qwen-Image | 20B | Q4_K_M | 14.6 GB |
| Qwen-Image-Edit | 20B | Q4_K_M | 14.6 GB |
| CogVLM2 | 19B | Q4_K_M | 13.9 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q5_K_M | 14.5 GB |
| Ling-Coder-Lite | 16.8B | Q5_K_M | 14.3 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q6_K | 15.7 GB |
| Kimi-VL A3B | 16B | Q6_K | 15.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q6_K | 14.8 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q6_K | 14.8 GB |
| Qwen2.5 14B | 14.7B | Q6_K | 15.3 GB |
| Phi-3 Medium | 14B | Q6_K | 13.8 GB |
| Phi-4 | 14B | Q6_K | 13.8 GB |
| Phi-4-reasoning / -plus | 14B | Q6_K | 13.8 GB |
| Wan 2.2 T2I | 14B | Q6_K | 13.8 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q6_K | 13.8 GB |
| SkyReels V2 | 14B | Q6_K | 13.8 GB |
| Vicuna 13B | 13B | Q6_K | 12.8 GB |
| HunyuanVideo | 13B | Q6_K | 12.8 GB |
| HunyuanVideo-Avatar | 13B | Q6_K | 12.8 GB |
| LTX-Video / LTX-2 | 13B | Q6_K | 12.8 GB |
| FramePack | 13B | Q6_K | 12.8 GB |
| FLUX.1 dev | 12B | FP8 / optimized | 14.4 GB |
| Gemma 3 12B | 12B | Q8_0 | 15.3 GB |
| Gemma 4 12B | 12B | Q8_0 | 15.3 GB |
| Mistral NeMo 12B | 12B | Q8_0 | 15.3 GB |
| Pixtral 12B | 12B | Q8_0 | 15.3 GB |
| FLUX.1 schnell | 12B | Q8_0 | 15.3 GB |
| FLUX.1 Kontext dev | 12B | Q8_0 | 15.3 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
| Mistral Small 3.2 | 24B | 17.6 GB needed | 19.6 GB |
| Magistral Small | 24B | 17.6 GB needed | 19.6 GB |
| Devstral Small 1.1 | 24B | 17.6 GB needed | 19.6 GB |
| Aria | 25B | 18.3 GB needed | 20.3 GB |
| Gemma 4 26B-A4B | 26B | 19 GB needed | 21 GB |
| Gemma 4 (all sizes) | 26B | 19 GB needed | 21 GB |
| Gemma 3 27B | 27B | 19.8 GB needed | 21.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | 19.8 GB needed | 21.8 GB |
How to read this
The NVIDIA Quadro GP100 graphics card features 16 GB of HBM2 frame buffer memory. This high bandwidth memory determines the maximum size of the artificial intelligence models you can run locally. To fit a model entirely within this hardware limit, the model parameters must undergo quantization. Quantization is a compression process that reduces the precision of model weights to save space. The quant column indicates the highest quality quantization level that fits within your hardware limits without causing out of memory errors.
For fully local execution on the graphics card, the largest fitting models include gpt-oss-20b and Reka Flash 3. Both are 21B parameter models that run at the Q4_K_M quantization level using 15.4 GB of memory. You can also run the 20B Qwen-Image and Qwen-Image-Edit models at Q4_K_M using 14.6 GB of memory. The 19B CogVLM2 model fits at Q4_K_M using 13.9 GB of memory. HunyuanImage 2.1 and HunyuanImage 3.0 are 17B models that run at Q5_K_M using 14.5 GB of memory.
Several high quality models fit at the Q6_K quantization level. The DeepSeek-Coder-V2 16B / 236B and Kimi-VL A3B models use 15.7 GB of memory. Apriel-1.5-15B-Thinker and StarCoder2 3B / 7B / 15B use 14.8 GB of memory. The Qwen2.5 14B model uses 15.3 GB of memory. You can run Phi-3 Medium, Phi-4, Phi-4-reasoning / -plus, Wan 2.2 T2I, Wan 2.1 (1.3B / 14B), and SkyReels V2 using 13.8 GB of memory. Vicuna 13B, HunyuanVideo, HunyuanVideo-Avatar, LTX-Video / LTX-2, and FramePack use 12.8 GB of memory.
If you want to use the highest precision Q8_0 quantization level, you can run 12B models. Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, and FLUX.1 Kontext dev all use 15.3 GB of memory at Q8_0. The FLUX.1 dev model uses 14.4 GB of memory at its optimized FP8 quantization level. These configurations ensure that all computations remain on the fast HBM2 memory of your graphics card.
When a model is too large for the 16 GB frame buffer, you can offload parts of it to your system RAM. This offload process allows you to run larger models but reduces processing speed. Assuming a system with 32 GB of system RAM, you can run Solar Pro or Codestral 22B at Q4_K_M, which requires 16.1 GB of memory and 18.1 GB of system RAM. Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 require 17.6 GB of memory and 19.6 GB of system RAM. Aria requires 18.3 GB of memory and 20.3 GB of system RAM.
Even larger models can run with system RAM offloading. Gemma 4 26B-A4B and Gemma 4 (all sizes) require 19 GB of memory and 21 GB of system RAM at Q4_K_M. Gemma 3 27B and Gemma 3 4B/12B/27B (vision) require 19.8 GB of memory and 21.8 GB of system RAM at Q4_K_M. Be aware that all memory calculations assume a standard 4k context window. Increasing the context window size requires more memory for active processing, which will reduce the maximum model size you can load.