Best local AI models for NVIDIA RTX 5090 D V2
24 GB GDDR7. At a 4k context, 173 of the 233 models in our catalog with verified parameter counts fit fully, up to Qwen3 8B / 14B / 32B at 32B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 173 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Qwen3 8B / 14B / 32B | 32B | Q4_K_M | 23.4 GB |
| Qwen3.5 (dense variants) | 32B | Q4_K_M | 23.4 GB |
| Aya Expanse 8B / 32B | 32B | Q4_K_M | 23.4 GB |
| Granite 4.0 Small/Tiny | 32B | Q4_K_M | 23.4 GB |
| Qwen2.5-Coder 0.5B to 32B | 32B | Q4_K_M | 23.4 GB |
| Qwen3-30B-A3B | 30B | Q4_K_M | 22 GB |
| Qwen3-Coder 30B-A3B | 30B | Q4_K_M | 22 GB |
| Gemma 3 27B | 27B | Q5_K_M | 23 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | Q5_K_M | 23 GB |
| Wan 2.2 / 2.5 | 27B | Q5_K_M | 23 GB |
| Gemma 4 26B-A4B | 26B | Q5_K_M | 22.2 GB |
| Gemma 4 (all sizes) | 26B | Q5_K_M | 22.2 GB |
| Aria | 25B | Q5_K_M | 21.3 GB |
| Mistral Small 3.2 | 24B | Q6_K | 23.6 GB |
| Magistral Small | 24B | Q6_K | 23.6 GB |
| Devstral Small 1.1 | 24B | Q6_K | 23.6 GB |
| Solar Pro | 22B | Q6_K | 21.6 GB |
| Codestral 22B | 22B | Q6_K | 21.6 GB |
| gpt-oss-20b | 21B | Q6_K | 20.7 GB |
| Reka Flash 3 | 21B | Q6_K | 20.7 GB |
| Qwen-Image | 20B | Q6_K | 19.7 GB |
| Qwen-Image-Edit | 20B | Q6_K | 19.7 GB |
| CogVLM2 | 19B | Q6_K | 18.7 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q8_0 | 21.6 GB |
| Ling-Coder-Lite | 16.8B | Q8_0 | 21.4 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q8_0 | 20.4 GB |
| Kimi-VL A3B | 16B | Q8_0 | 20.4 GB |
| Apriel-1.5-15B-Thinker | 15B | Q8_0 | 19.1 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q8_0 | 19.1 GB |
| Qwen2.5 14B | 14.7B | Q8_0 | 19.5 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| OTel 2.0 LLM 31B IT | 32.1B | 27.5 GB needed | 29.5 GB |
| DeepSeek-Coder 1.3B / 6.7B / 33B | 33B | 24.2 GB needed | 26.2 GB |
| WizardCoder 33B | 33B | 24.2 GB needed | 26.2 GB |
| Yi 1.5 9B / 34B | 34B | 24.9 GB needed | 26.9 GB |
| Granite Code 3B to 34B | 34B | 24.9 GB needed | 26.9 GB |
| LLaVA 1.5 / 1.6 (7B to 34B) | 34B | 24.9 GB needed | 26.9 GB |
| Ovis 2 | 34B | 24.9 GB needed | 26.9 GB |
| Qwen3.6-35B-A3B | 35B | 25.6 GB needed | 27.6 GB |
| Command R (35B) | 35B | 25.6 GB needed | 27.6 GB |
| Seed-OSS 36B | 36B | 26.4 GB needed | 28.4 GB |
How to read this
The NVIDIA RTX 5090 D V2 graphics card comes equipped with 24 GB of GDDR7 memory. This dedicated video memory determines the maximum size of the artificial intelligence model you can run locally. To fit a model entirely on this hardware, the total memory footprint of the model and its operational workspace must remain under this 24 GB limit.
The quantization column indicates the compression level applied to each model. Quantization reduces the precision of model weights to save space. For example, the Qwen3 32B, Qwen3.5 32B, Aya Expanse 32B, Granite 4.0 Small/Tiny 32B, and Qwen2.5-Coder 32B models fit within 23.4 GB using the Q4_K_M quantization. Similarly, Gemma 3 27B, Wan 2.2 / 2.5 27B, and Gemma 4 26B-A4B utilize the Q5_K_M quantization to fit within 23 GB and 22.2 GB respectively.
Higher precision quantizations offer better output quality but require more memory. You can run Mistral Small 3.2 24B, Magistral Small 24B, and Devstral Small 1.1 24B at the Q6_K quantization level using 23.6 GB of memory. For maximum precision, models like DeepSeek-Coder-V2 16B, Kimi-VL A3B 16B, and Qwen2.5 14B run at the Q8_0 quantization level, consuming up to 21.4 GB of video memory.
When a model exceeds the 24 GB onboard memory, you must offload parts of it to your system RAM. This offloading process allows you to run larger models but slows down generation speeds significantly. For instance, running the Yi 1.5 34B or Granite Code 34B requires 24.9 GB of memory at Q4_K_M quantization, which demands 26.9 GB of system RAM on a standard 32 GB system.
Other offloading examples include the Qwen3.6-35B-A3B and Command R 35B models. These models require 25.6 GB at Q4_K_M quantization and need 27.6 GB of system RAM to function. The Seed-OSS 36B model requires 26.4 GB at Q4_K_M quantization and needs 28.4 GB of system RAM.
Memory consumption calculations assume a standard 4k context window. If you increase the context length to process longer documents or chat histories, the memory requirements will rise. This extra context overhead can push a model past the 24 GB physical limit of your card and trigger automatic system RAM offloading.