Best local AI models for NVIDIA Quadro RTX 8000

48 GB GDDR6. At a 4k context, 186 of the 233 models in our catalog with verified parameter counts fit fully, up to Jamba 1.5 Mini / Large at 52B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 186 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Jamba 1.5 Mini / Large52BQ5_K_M44.3 GB
Llama 3.1 Nemotron 51B51BQ5_K_M43.5 GB
Mixtral 8x7B47BQ6_K46.2 GB
Seed-OSS 36B36BQ8_045.8 GB
Qwen3.6-35B-A3B35BQ8_044.5 GB
Command R (35B)35BQ8_044.5 GB
Yi 1.5 9B / 34B34BQ8_043.2 GB
Granite Code 3B to 34B34BQ8_043.2 GB
LLaVA 1.5 / 1.6 (7B to 34B)34BQ8_043.2 GB
Ovis 234BQ8_043.2 GB
DeepSeek-Coder 1.3B / 6.7B / 33B33BQ8_042 GB
WizardCoder 33B33BQ8_042 GB
OTel 2.0 LLM 31B IT32.1BQ8_044.9 GB
Qwen3 8B / 14B / 32B32BQ8_040.7 GB
Qwen3.5 (dense variants)32BQ8_040.7 GB
Aya Expanse 8B / 32B32BQ8_040.7 GB
Granite 4.0 Small/Tiny32BQ8_040.7 GB
Qwen2.5-Coder 0.5B to 32B32BQ8_040.7 GB
Qwen3-30B-A3B30BQ8_038.2 GB
Qwen3-Coder 30B-A3B30BQ8_038.2 GB
Gemma 3 27B27BQ8_034.3 GB
Gemma 3 4B/12B/27B (vision)27BQ8_034.3 GB
Wan 2.2 / 2.527BQ8_034.3 GB
Gemma 4 26B-A4B26BQ8_033.1 GB
Gemma 4 (all sizes)26BQ8_033.1 GB
Aria25BQ8_031.8 GB
Mistral Small 3.224BQ8_030.5 GB
Magistral Small24BQ8_030.5 GB
Devstral Small 1.124BQ8_030.5 GB
Solar Pro22BQ8_028 GB

How to read this

The NVIDIA Quadro RTX 8000 graphics card features 48 GB GDDR6 of onboard video memory. This large frame buffer determines the size of the artificial intelligence models you can run locally. For optimal performance, the entire model must fit within this physical memory limit. If a model exceeds the available space, it cannot run entirely on the graphics hardware.

The quantization column indicates the level of compression applied to the model weights. Running models at higher precision requires more memory. Using quantized formats like Q8_0 or Q5_K_M reduces the memory footprint while preserving most of the model intelligence. This allows you to run larger parameter architectures on a single workstation.

With 48 GB of video memory, you can run several large models at high precision. The Jamba 1.5 Mini or Large 52B model fits at a Q5_K_M quantization using 44.3 GB of memory. The Llama 3.1 Nemotron 51B model also runs at the Q5_K_M quantization level using 43.5 GB of memory. Mixtral 8x7B fits at a Q6_K quantization using 46.2 GB of memory.

Many powerful models run at the maximum Q8_0 quantization level. The Seed-OSS 36B model fits using 45.8 GB of memory. The Qwen3.6-35B-A3B and Command R 35B models both use 44.5 GB of memory. Yi 1.5 34B, Granite Code 34B, LLaVA 1.5 or 1.6 34B, and Ovis 2 use 43.2 GB of memory. DeepSeek-Coder 33B and WizardCoder 33B use 42 GB of memory. The OTel 2.0 LLM 31B IT model fits using 44.9 GB of memory.

Other notable models fit comfortably within the memory limit at Q8_0 quantization. Qwen3 32B, Qwen3.5 dense variants, Aya Expanse 32B, Granite 4.0 Small, and Qwen2.5-Coder 32B all use 40.7 GB of memory. Qwen3-30B-A3B and Qwen3-Coder 30B-A3B use 38.2 GB of memory. Gemma 3 27B, Gemma 3 vision 27B, and Wan 2.2 or 2.5 use 34.3 GB of memory. Gemma 4 26B-A4B and Gemma 4 all sizes use 33.1 GB of memory. Aria uses 31.8 GB of memory. Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 use 30.5 GB of memory. Solar Pro uses 28 GB of memory.

When selecting a model, you must account for the context window. The memory figures listed are calculated using a standard 4k context window. If you increase the context length to process longer documents, the system will require additional video memory. This extra memory requirement may force you to choose a smaller model or a lower quantization level to avoid running out of memory.

System RAM offloading is not required for these configurations. Offloading parts of a model to system RAM slows down processing speeds significantly. Keeping the entire model within the 48 GB GDDR6 memory of the NVIDIA Quadro RTX 8000 ensures maximum generation speed and efficiency.