Best local AI models for NVIDIA RTX 6000 Ada Generation
48 GB GDDR6. At a 4k context, 186 of the 233 models in our catalog with verified parameter counts fit fully, up to Jamba 1.5 Mini / Large at 52B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 186 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Jamba 1.5 Mini / Large | 52B | Q5_K_M | 44.3 GB |
| Llama 3.1 Nemotron 51B | 51B | Q5_K_M | 43.5 GB |
| Mixtral 8x7B | 47B | Q6_K | 46.2 GB |
| Seed-OSS 36B | 36B | Q8_0 | 45.8 GB |
| Qwen3.6-35B-A3B | 35B | Q8_0 | 44.5 GB |
| Command R (35B) | 35B | Q8_0 | 44.5 GB |
| Yi 1.5 9B / 34B | 34B | Q8_0 | 43.2 GB |
| Granite Code 3B to 34B | 34B | Q8_0 | 43.2 GB |
| LLaVA 1.5 / 1.6 (7B to 34B) | 34B | Q8_0 | 43.2 GB |
| Ovis 2 | 34B | Q8_0 | 43.2 GB |
| DeepSeek-Coder 1.3B / 6.7B / 33B | 33B | Q8_0 | 42 GB |
| WizardCoder 33B | 33B | Q8_0 | 42 GB |
| OTel 2.0 LLM 31B IT | 32.1B | Q8_0 | 44.9 GB |
| Qwen3 8B / 14B / 32B | 32B | Q8_0 | 40.7 GB |
| Qwen3.5 (dense variants) | 32B | Q8_0 | 40.7 GB |
| Aya Expanse 8B / 32B | 32B | Q8_0 | 40.7 GB |
| Granite 4.0 Small/Tiny | 32B | Q8_0 | 40.7 GB |
| Qwen2.5-Coder 0.5B to 32B | 32B | Q8_0 | 40.7 GB |
| Qwen3-30B-A3B | 30B | Q8_0 | 38.2 GB |
| Qwen3-Coder 30B-A3B | 30B | Q8_0 | 38.2 GB |
| Gemma 3 27B | 27B | Q8_0 | 34.3 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | Q8_0 | 34.3 GB |
| Wan 2.2 / 2.5 | 27B | Q8_0 | 34.3 GB |
| Gemma 4 26B-A4B | 26B | Q8_0 | 33.1 GB |
| Gemma 4 (all sizes) | 26B | Q8_0 | 33.1 GB |
| Aria | 25B | Q8_0 | 31.8 GB |
| Mistral Small 3.2 | 24B | Q8_0 | 30.5 GB |
| Magistral Small | 24B | Q8_0 | 30.5 GB |
| Devstral Small 1.1 | 24B | Q8_0 | 30.5 GB |
| Solar Pro | 22B | Q8_0 | 28 GB |
How to read this
The NVIDIA RTX 6000 Ada Generation workstation graphics card features 48 GB GDDR6 of dedicated video memory. This large frame buffer allows you to run highly capable artificial intelligence models completely on the local hardware. When you run models locally, your data remains private and you do not rely on external cloud APIs.
To fit larger models into the 48 GB memory space, developers use quantization. Quantization reduces the precision of model weights to save space. The best quant column shows the highest quality format that fits comfortably within your hardware limits. For example, the 52B Jamba 1.5 Mini or Large model fits at a Q5_K_M quantization level which uses 44.3 GB of video memory. The 51B Llama 3.1 Nemotron 51B model fits at Q5_K_M using 43.5 GB of memory.
Many other large models can run at the maximum Q8_0 quantization level which preserves excellent output quality. The Mixtral 8x7B model fits at Q6_K using 46.2 GB of memory. Models like Seed-OSS 36B at Q8_0 use 45.8 GB of memory. The Qwen3.6-35B-A3B and Command R (35B) models both run at Q8_0 using 44.5 GB of memory. You can also run Yi 1.5 9B / 34B, Granite Code 3B to 34B, LLaVA 1.5 / 1.6 (7B to 34B), and Ovis 2 at Q8_0 using 43.2 GB of memory.
Medium sized models fit with even more room to spare. DeepSeek-Coder 1.3B / 6.7B / 33B and WizardCoder 33B run at Q8_0 using 42 GB of memory. The OTel 2.0 LLM 31B IT model runs at Q8_0 using 44.9 GB of memory. Several 32B models like Qwen3 8B / 14B / 32B, Qwen3.5 (dense variants), Aya Expanse 8B / 32B, Granite 4.0 Small/Tiny, and Qwen2.5-Coder 0.5B to 32B run at Q8_0 using 40.7 GB of memory. Qwen3-30B-A3B and Qwen3-Coder 30B-A3B run at Q8_0 using 38.2 GB of memory.
Smaller models run with complete ease on this hardware. Gemma 3 27B, Gemma 3 4B/12B/27B (vision), and Wan 2.2 / 2.5 run at Q8_0 using 34.3 GB of memory. Gemma 4 26B-A4B and Gemma 4 (all sizes) run at Q8_0 using 33.1 GB of memory. Aria runs at Q8_0 using 31.8 GB of memory. Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 run at Q8_0 using 30.5 GB of memory. Solar Pro runs at Q8_0 using 28 GB of memory.
All memory calculations are based on a standard 4k context window. If you increase the context window to process longer documents, the system will require significantly more memory. Running out of video memory forces the system to offload processing to system RAM. Offloading to system RAM prevents crashes but it slows down generation speeds dramatically because system RAM is much slower than GDDR6 video memory.