Best local AI models for NVIDIA Quadro RTX 8000
48 GB GDDR6. At a 4k context, 186 of the 233 models in our catalog with verified parameter counts fit fully, up to Jamba 1.5 Mini / Large at 52B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 186 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Jamba 1.5 Mini / Large | 52B | Q5_K_M | 44.3 GB |
| Llama 3.1 Nemotron 51B | 51B | Q5_K_M | 43.5 GB |
| Mixtral 8x7B | 47B | Q6_K | 46.2 GB |
| Seed-OSS 36B | 36B | Q8_0 | 45.8 GB |
| Qwen3.6-35B-A3B | 35B | Q8_0 | 44.5 GB |
| Command R (35B) | 35B | Q8_0 | 44.5 GB |
| Yi 1.5 9B / 34B | 34B | Q8_0 | 43.2 GB |
| Granite Code 3B to 34B | 34B | Q8_0 | 43.2 GB |
| LLaVA 1.5 / 1.6 (7B to 34B) | 34B | Q8_0 | 43.2 GB |
| Ovis 2 | 34B | Q8_0 | 43.2 GB |
| DeepSeek-Coder 1.3B / 6.7B / 33B | 33B | Q8_0 | 42 GB |
| WizardCoder 33B | 33B | Q8_0 | 42 GB |
| OTel 2.0 LLM 31B IT | 32.1B | Q8_0 | 44.9 GB |
| Qwen3 8B / 14B / 32B | 32B | Q8_0 | 40.7 GB |
| Qwen3.5 (dense variants) | 32B | Q8_0 | 40.7 GB |
| Aya Expanse 8B / 32B | 32B | Q8_0 | 40.7 GB |
| Granite 4.0 Small/Tiny | 32B | Q8_0 | 40.7 GB |
| Qwen2.5-Coder 0.5B to 32B | 32B | Q8_0 | 40.7 GB |
| Qwen3-30B-A3B | 30B | Q8_0 | 38.2 GB |
| Qwen3-Coder 30B-A3B | 30B | Q8_0 | 38.2 GB |
| Gemma 3 27B | 27B | Q8_0 | 34.3 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | Q8_0 | 34.3 GB |
| Wan 2.2 / 2.5 | 27B | Q8_0 | 34.3 GB |
| Gemma 4 26B-A4B | 26B | Q8_0 | 33.1 GB |
| Gemma 4 (all sizes) | 26B | Q8_0 | 33.1 GB |
| Aria | 25B | Q8_0 | 31.8 GB |
| Mistral Small 3.2 | 24B | Q8_0 | 30.5 GB |
| Magistral Small | 24B | Q8_0 | 30.5 GB |
| Devstral Small 1.1 | 24B | Q8_0 | 30.5 GB |
| Solar Pro | 22B | Q8_0 | 28 GB |
How to read this
The NVIDIA Quadro RTX 8000 graphics card features 48 GB GDDR6 of onboard video memory. This large frame buffer determines the size of the artificial intelligence models you can run locally. For optimal performance, the entire model must fit within this physical memory limit. If a model exceeds the available space, it cannot run entirely on the graphics hardware.
The quantization column indicates the level of compression applied to the model weights. Running models at higher precision requires more memory. Using quantized formats like Q8_0 or Q5_K_M reduces the memory footprint while preserving most of the model intelligence. This allows you to run larger parameter architectures on a single workstation.
With 48 GB of video memory, you can run several large models at high precision. The Jamba 1.5 Mini or Large 52B model fits at a Q5_K_M quantization using 44.3 GB of memory. The Llama 3.1 Nemotron 51B model also runs at the Q5_K_M quantization level using 43.5 GB of memory. Mixtral 8x7B fits at a Q6_K quantization using 46.2 GB of memory.
Many powerful models run at the maximum Q8_0 quantization level. The Seed-OSS 36B model fits using 45.8 GB of memory. The Qwen3.6-35B-A3B and Command R 35B models both use 44.5 GB of memory. Yi 1.5 34B, Granite Code 34B, LLaVA 1.5 or 1.6 34B, and Ovis 2 use 43.2 GB of memory. DeepSeek-Coder 33B and WizardCoder 33B use 42 GB of memory. The OTel 2.0 LLM 31B IT model fits using 44.9 GB of memory.
Other notable models fit comfortably within the memory limit at Q8_0 quantization. Qwen3 32B, Qwen3.5 dense variants, Aya Expanse 32B, Granite 4.0 Small, and Qwen2.5-Coder 32B all use 40.7 GB of memory. Qwen3-30B-A3B and Qwen3-Coder 30B-A3B use 38.2 GB of memory. Gemma 3 27B, Gemma 3 vision 27B, and Wan 2.2 or 2.5 use 34.3 GB of memory. Gemma 4 26B-A4B and Gemma 4 all sizes use 33.1 GB of memory. Aria uses 31.8 GB of memory. Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 use 30.5 GB of memory. Solar Pro uses 28 GB of memory.
When selecting a model, you must account for the context window. The memory figures listed are calculated using a standard 4k context window. If you increase the context length to process longer documents, the system will require additional video memory. This extra memory requirement may force you to choose a smaller model or a lower quantization level to avoid running out of memory.
System RAM offloading is not required for these configurations. Offloading parts of a model to system RAM slows down processing speeds significantly. Keeping the entire model within the 48 GB GDDR6 memory of the NVIDIA Quadro RTX 8000 ensures maximum generation speed and efficiency.