Best local AI models for NVIDIA RTX 6000 Ada Generation

48 GB GDDR6. At a 4k context, 186 of the 233 models in our catalog with verified parameter counts fit fully, up to Jamba 1.5 Mini / Large at 52B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 186 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Jamba 1.5 Mini / Large52BQ5_K_M44.3 GB
Llama 3.1 Nemotron 51B51BQ5_K_M43.5 GB
Mixtral 8x7B47BQ6_K46.2 GB
Seed-OSS 36B36BQ8_045.8 GB
Qwen3.6-35B-A3B35BQ8_044.5 GB
Command R (35B)35BQ8_044.5 GB
Yi 1.5 9B / 34B34BQ8_043.2 GB
Granite Code 3B to 34B34BQ8_043.2 GB
LLaVA 1.5 / 1.6 (7B to 34B)34BQ8_043.2 GB
Ovis 234BQ8_043.2 GB
DeepSeek-Coder 1.3B / 6.7B / 33B33BQ8_042 GB
WizardCoder 33B33BQ8_042 GB
OTel 2.0 LLM 31B IT32.1BQ8_044.9 GB
Qwen3 8B / 14B / 32B32BQ8_040.7 GB
Qwen3.5 (dense variants)32BQ8_040.7 GB
Aya Expanse 8B / 32B32BQ8_040.7 GB
Granite 4.0 Small/Tiny32BQ8_040.7 GB
Qwen2.5-Coder 0.5B to 32B32BQ8_040.7 GB
Qwen3-30B-A3B30BQ8_038.2 GB
Qwen3-Coder 30B-A3B30BQ8_038.2 GB
Gemma 3 27B27BQ8_034.3 GB
Gemma 3 4B/12B/27B (vision)27BQ8_034.3 GB
Wan 2.2 / 2.527BQ8_034.3 GB
Gemma 4 26B-A4B26BQ8_033.1 GB
Gemma 4 (all sizes)26BQ8_033.1 GB
Aria25BQ8_031.8 GB
Mistral Small 3.224BQ8_030.5 GB
Magistral Small24BQ8_030.5 GB
Devstral Small 1.124BQ8_030.5 GB
Solar Pro22BQ8_028 GB

How to read this

The NVIDIA RTX 6000 Ada Generation workstation graphics card features 48 GB GDDR6 of dedicated video memory. This large frame buffer allows you to run highly capable artificial intelligence models completely on the local hardware. When you run models locally, your data remains private and you do not rely on external cloud APIs.

To fit larger models into the 48 GB memory space, developers use quantization. Quantization reduces the precision of model weights to save space. The best quant column shows the highest quality format that fits comfortably within your hardware limits. For example, the 52B Jamba 1.5 Mini or Large model fits at a Q5_K_M quantization level which uses 44.3 GB of video memory. The 51B Llama 3.1 Nemotron 51B model fits at Q5_K_M using 43.5 GB of memory.

Many other large models can run at the maximum Q8_0 quantization level which preserves excellent output quality. The Mixtral 8x7B model fits at Q6_K using 46.2 GB of memory. Models like Seed-OSS 36B at Q8_0 use 45.8 GB of memory. The Qwen3.6-35B-A3B and Command R (35B) models both run at Q8_0 using 44.5 GB of memory. You can also run Yi 1.5 9B / 34B, Granite Code 3B to 34B, LLaVA 1.5 / 1.6 (7B to 34B), and Ovis 2 at Q8_0 using 43.2 GB of memory.

Medium sized models fit with even more room to spare. DeepSeek-Coder 1.3B / 6.7B / 33B and WizardCoder 33B run at Q8_0 using 42 GB of memory. The OTel 2.0 LLM 31B IT model runs at Q8_0 using 44.9 GB of memory. Several 32B models like Qwen3 8B / 14B / 32B, Qwen3.5 (dense variants), Aya Expanse 8B / 32B, Granite 4.0 Small/Tiny, and Qwen2.5-Coder 0.5B to 32B run at Q8_0 using 40.7 GB of memory. Qwen3-30B-A3B and Qwen3-Coder 30B-A3B run at Q8_0 using 38.2 GB of memory.

Smaller models run with complete ease on this hardware. Gemma 3 27B, Gemma 3 4B/12B/27B (vision), and Wan 2.2 / 2.5 run at Q8_0 using 34.3 GB of memory. Gemma 4 26B-A4B and Gemma 4 (all sizes) run at Q8_0 using 33.1 GB of memory. Aria runs at Q8_0 using 31.8 GB of memory. Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 run at Q8_0 using 30.5 GB of memory. Solar Pro runs at Q8_0 using 28 GB of memory.

All memory calculations are based on a standard 4k context window. If you increase the context window to process longer documents, the system will require significantly more memory. Running out of video memory forces the system to offload processing to system RAM. Offloading to system RAM prevents crashes but it slows down generation speeds dramatically because system RAM is much slower than GDDR6 video memory.