Best local AI models for NVIDIA Quadro RTX 6000
24 GB GDDR6. At a 4k context, 173 of the 233 models in our catalog with verified parameter counts fit fully, up to Qwen3 8B / 14B / 32B at 32B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 173 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Qwen3 8B / 14B / 32B | 32B | Q4_K_M | 23.4 GB |
| Qwen3.5 (dense variants) | 32B | Q4_K_M | 23.4 GB |
| Aya Expanse 8B / 32B | 32B | Q4_K_M | 23.4 GB |
| Granite 4.0 Small/Tiny | 32B | Q4_K_M | 23.4 GB |
| Qwen2.5-Coder 0.5B to 32B | 32B | Q4_K_M | 23.4 GB |
| Qwen3-30B-A3B | 30B | Q4_K_M | 22 GB |
| Qwen3-Coder 30B-A3B | 30B | Q4_K_M | 22 GB |
| Gemma 3 27B | 27B | Q5_K_M | 23 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | Q5_K_M | 23 GB |
| Wan 2.2 / 2.5 | 27B | Q5_K_M | 23 GB |
| Gemma 4 26B-A4B | 26B | Q5_K_M | 22.2 GB |
| Gemma 4 (all sizes) | 26B | Q5_K_M | 22.2 GB |
| Aria | 25B | Q5_K_M | 21.3 GB |
| Mistral Small 3.2 | 24B | Q6_K | 23.6 GB |
| Magistral Small | 24B | Q6_K | 23.6 GB |
| Devstral Small 1.1 | 24B | Q6_K | 23.6 GB |
| Solar Pro | 22B | Q6_K | 21.6 GB |
| Codestral 22B | 22B | Q6_K | 21.6 GB |
| gpt-oss-20b | 21B | Q6_K | 20.7 GB |
| Reka Flash 3 | 21B | Q6_K | 20.7 GB |
| Qwen-Image | 20B | Q6_K | 19.7 GB |
| Qwen-Image-Edit | 20B | Q6_K | 19.7 GB |
| CogVLM2 | 19B | Q6_K | 18.7 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q8_0 | 21.6 GB |
| Ling-Coder-Lite | 16.8B | Q8_0 | 21.4 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q8_0 | 20.4 GB |
| Kimi-VL A3B | 16B | Q8_0 | 20.4 GB |
| Apriel-1.5-15B-Thinker | 15B | Q8_0 | 19.1 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q8_0 | 19.1 GB |
| Qwen2.5 14B | 14.7B | Q8_0 | 19.5 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| OTel 2.0 LLM 31B IT | 32.1B | 27.5 GB needed | 29.5 GB |
| DeepSeek-Coder 1.3B / 6.7B / 33B | 33B | 24.2 GB needed | 26.2 GB |
| WizardCoder 33B | 33B | 24.2 GB needed | 26.2 GB |
| Yi 1.5 9B / 34B | 34B | 24.9 GB needed | 26.9 GB |
| Granite Code 3B to 34B | 34B | 24.9 GB needed | 26.9 GB |
| LLaVA 1.5 / 1.6 (7B to 34B) | 34B | 24.9 GB needed | 26.9 GB |
| Ovis 2 | 34B | 24.9 GB needed | 26.9 GB |
| Qwen3.6-35B-A3B | 35B | 25.6 GB needed | 27.6 GB |
| Command R (35B) | 35B | 25.6 GB needed | 27.6 GB |
| Seed-OSS 36B | 36B | 26.4 GB needed | 28.4 GB |
How to read this
The NVIDIA Quadro RTX 6000 graphics card features 24 GB of GDDR6 memory. This dedicated memory determines the size of the artificial intelligence models you can run locally. To fit a model entirely on the hardware, the total memory footprint must remain under this 24 GB limit. Running models fully inside this fast graphics memory ensures the highest speed and lowest latency during text generation.
The quantization column indicates the compression level applied to each model. Quantization reduces the precision of model weights to save space. For example, a Q4_K_M quantization represents a medium four bit format. This allows larger models like Qwen3 32B, Qwen3.5 32B, Aya Expanse 32B, Granite 4.0 32B, and Qwen2.5-Coder 32B to fit into 23.4 GB of graphics memory. Other models like Qwen3-30B-A3B and Qwen3-Coder 30B-A3B fit into 22 GB using the same Q4_K_M quantization.
Higher precision quantizations require more memory but preserve more original model quality. You can run Gemma 3 27B, Gemma 3 27B vision, and Wan 2.2 / 2.5 at Q5_K_M quantization using 23 GB of memory. Gemma 4 26B-A4B and Gemma 4 all sizes fit at Q5_K_M using 22.2 GB. Aria fits at Q5_K_M using 21.3 GB. Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 can run at Q6_K quantization using 23.6 GB. Solar Pro and Codestral 22B fit at Q6_K using 21.6 GB. You can also run gpt-oss-20b and Reka Flash 3 at Q6_K using 20.7 GB.
For even higher precision, the Q8_0 quantization uses eight bits. This format works well for slightly smaller models. HunyuanImage 2.1 / 3.0 fits at Q8_0 using 21.6 GB. Ling-Coder-Lite fits at Q8_0 using 21.4 GB. DeepSeek-Coder-V2 16B and Kimi-VL A3B fit at Q8_0 using 20.4 GB. Apriel-1.5-15B-Thinker and StarCoder2 15B fit at Q8_0 using 19.1 GB. Qwen2.5 14B fits at Q8_0 using 19.5 GB. Other models like Qwen-Image and Qwen-Image-Edit fit at Q6_K using 19.7 GB, while CogVLM2 fits at Q6_K using 18.7 GB.
When a model exceeds the 24 GB graphics memory, you must offload parts of it to your system RAM. This offloading process slows down generation speeds significantly. Assuming you have a 32 GB system RAM, you can run OTel 2.0 LLM 31B IT at Q4_K_M which needs 27.5 GB of memory and 29.5 GB of system RAM. DeepSeek-Coder 33B and WizardCoder 33B need 24.2 GB at Q4_K_M and 26.2 GB of system RAM. Yi 1.5 34B, Granite Code 34B, LLaVA 1.5 / 1.6 34B, and Ovis 2 need 24.9 GB at Q4_K_M and 26.9 GB of system RAM.
Other offload options include Qwen3.6-35B-A3B and Command R 35B which need 25.6 GB at Q4_K_M and 27.6 GB of system RAM. Seed-OSS 36B needs 26.4 GB at Q4_K_M and 28.4 GB of system RAM. Please note that all memory calculations assume a standard context window of four thousand tokens. Increasing the context window size will require more memory and may force you to use lower quantizations or more CPU offloading.