Best local AI models for NVIDIA RTX 5090 D
32 GB GDDR7. At a 4k context, 183 of the 233 models in our catalog with verified parameter counts fit fully, up to Seed-OSS 36B at 36B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 183 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Seed-OSS 36B | 36B | Q5_K_M | 30.7 GB |
| Qwen3.6-35B-A3B | 35B | Q5_K_M | 29.8 GB |
| Command R (35B) | 35B | Q5_K_M | 29.8 GB |
| Yi 1.5 9B / 34B | 34B | Q5_K_M | 29 GB |
| Granite Code 3B to 34B | 34B | Q5_K_M | 29 GB |
| LLaVA 1.5 / 1.6 (7B to 34B) | 34B | Q5_K_M | 29 GB |
| Ovis 2 | 34B | Q5_K_M | 29 GB |
| DeepSeek-Coder 1.3B / 6.7B / 33B | 33B | Q5_K_M | 28.1 GB |
| WizardCoder 33B | 33B | Q5_K_M | 28.1 GB |
| OTel 2.0 LLM 31B IT | 32.1B | Q5_K_M | 31.4 GB |
| Qwen3 8B / 14B / 32B | 32B | Q6_K | 31.5 GB |
| Qwen3.5 (dense variants) | 32B | Q6_K | 31.5 GB |
| Aya Expanse 8B / 32B | 32B | Q6_K | 31.5 GB |
| Granite 4.0 Small/Tiny | 32B | Q6_K | 31.5 GB |
| Qwen2.5-Coder 0.5B to 32B | 32B | Q6_K | 31.5 GB |
| Qwen3-30B-A3B | 30B | Q6_K | 29.5 GB |
| Qwen3-Coder 30B-A3B | 30B | Q6_K | 29.5 GB |
| Gemma 3 27B | 27B | Q6_K | 26.6 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | Q6_K | 26.6 GB |
| Wan 2.2 / 2.5 | 27B | Q6_K | 26.6 GB |
| Gemma 4 26B-A4B | 26B | Q6_K | 25.6 GB |
| Gemma 4 (all sizes) | 26B | Q6_K | 25.6 GB |
| Aria | 25B | Q8_0 | 31.8 GB |
| Mistral Small 3.2 | 24B | Q8_0 | 30.5 GB |
| Magistral Small | 24B | Q8_0 | 30.5 GB |
| Devstral Small 1.1 | 24B | Q8_0 | 30.5 GB |
| Solar Pro | 22B | Q8_0 | 28 GB |
| Codestral 22B | 22B | Q8_0 | 28 GB |
| gpt-oss-20b | 21B | Q8_0 | 26.7 GB |
| Reka Flash 3 | 21B | Q8_0 | 26.7 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Mixtral 8x7B | 47B | 34.4 GB needed | 36.4 GB |
| Llama 3.1 Nemotron 51B | 51B | 37.3 GB needed | 39.3 GB |
| Jamba 1.5 Mini / Large | 52B | 38.1 GB needed | 40.1 GB |
How to read this
The NVIDIA RTX 5090 D graphics card features 32 GB of GDDR7 onboard memory. This memory capacity determines which artificial intelligence models can run entirely on the graphics hardware. Running a model completely in the graphics memory ensures the fastest possible processing speeds for text generation and reasoning tasks.
To fit larger models into the 32 GB limit you must use quantized versions. Quantization reduces the precision of model weights to save space. The best quantization level represents the highest quality format that still fits inside the available graphics memory. For example Seed-OSS 36B fits at the Q5_K_M quantization level using 30.7 GB of memory. Similarly Qwen3.6-35B-A3B and Command R (35B) both run at Q5_K_M using 29.8 GB of memory.
Several high performance models fit comfortably within this memory envelope at the Q5_K_M quantization level. These include Yi 1.5 9B / 34B Granite Code 3B to 34B LLaVA 1.5 / 1.6 (7B to 34B) and Ovis 2 which all consume 29 GB of memory. DeepSeek-Coder 1.3B / 6.7B / 33B and WizardCoder 33B require 28.1 GB of memory at Q5_K_M. OTel 2.0 LLM 31B IT fits at Q5_K_M using 31.4 GB of memory.
You can run highly accurate Q6_K quantized models on this hardware. Qwen3 8B / 14B / 32B Qwen3.5 (dense variants) Aya Expanse 8B / 32B Granite 4.0 Small/Tiny and Qwen2.5-Coder 0.5B to 32B all run at Q6_K using 31.5 GB of memory. Qwen3-30B-A3B and Qwen3-Coder 30B-A3B use 29.5 GB of memory at Q6_K. Gemma 3 27B Gemma 3 4B/12B/27B (vision) and Wan 2.2 / 2.5 use 26.6 GB of memory at Q6_K. Gemma 4 26B-A4B and Gemma 4 (all sizes) use 25.6 GB of memory at Q6_K.
For maximum precision you can run Q8_0 quantized models. Aria fits at Q8_0 using 31.8 GB of memory. Mistral Small 3.2 Magistral Small and Devstral Small 1.1 all run at Q8_0 using 30.5 GB of memory. Solar Pro and Codestral 22B use 28 GB of memory at Q8_0. Both gpt-oss-20b and Reka Flash 3 run at Q8_0 using 26.7 GB of memory.
If you want to run models that exceed 32 GB you must offload parts of the model to your system CPU and system RAM. This offloading process slows down processing speeds significantly. For example Mixtral 8x7B needs 34.4 GB at Q4_K_M and requires 36.4 GB of system RAM. Llama 3.1 Nemotron 51B needs 37.3 GB at Q4_K_M and requires 39.3 GB of system RAM. Jamba 1.5 Mini / Large needs 38.1 GB at Q4_K_M and requires 40.1 GB of system RAM.
All memory calculations assume a standard 4k context window. Generating longer responses or processing larger documents increases memory usage. If you increase the context window beyond 4k tokens the model may exceed the 32 GB limit and require system RAM offloading.