Best local AI models for NVIDIA RTX 5090
32 GB GDDR7. At a 4k context, 183 of the 233 models in our catalog with verified parameter counts fit fully, up to Seed-OSS 36B at 36B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 183 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Seed-OSS 36B | 36B | Q5_K_M | 30.7 GB |
| Qwen3.6-35B-A3B | 35B | Q5_K_M | 29.8 GB |
| Command R (35B) | 35B | Q5_K_M | 29.8 GB |
| Yi 1.5 9B / 34B | 34B | Q5_K_M | 29 GB |
| Granite Code 3B to 34B | 34B | Q5_K_M | 29 GB |
| LLaVA 1.5 / 1.6 (7B to 34B) | 34B | Q5_K_M | 29 GB |
| Ovis 2 | 34B | Q5_K_M | 29 GB |
| DeepSeek-Coder 1.3B / 6.7B / 33B | 33B | Q5_K_M | 28.1 GB |
| WizardCoder 33B | 33B | Q5_K_M | 28.1 GB |
| OTel 2.0 LLM 31B IT | 32.1B | Q5_K_M | 31.4 GB |
| Qwen3 8B / 14B / 32B | 32B | Q6_K | 31.5 GB |
| Qwen3.5 (dense variants) | 32B | Q6_K | 31.5 GB |
| Aya Expanse 8B / 32B | 32B | Q6_K | 31.5 GB |
| Granite 4.0 Small/Tiny | 32B | Q6_K | 31.5 GB |
| Qwen2.5-Coder 0.5B to 32B | 32B | Q6_K | 31.5 GB |
| Qwen3-30B-A3B | 30B | Q6_K | 29.5 GB |
| Qwen3-Coder 30B-A3B | 30B | Q6_K | 29.5 GB |
| Gemma 3 27B | 27B | Q6_K | 26.6 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | Q6_K | 26.6 GB |
| Wan 2.2 / 2.5 | 27B | Q6_K | 26.6 GB |
| Gemma 4 26B-A4B | 26B | Q6_K | 25.6 GB |
| Gemma 4 (all sizes) | 26B | Q6_K | 25.6 GB |
| Aria | 25B | Q8_0 | 31.8 GB |
| Mistral Small 3.2 | 24B | Q8_0 | 30.5 GB |
| Magistral Small | 24B | Q8_0 | 30.5 GB |
| Devstral Small 1.1 | 24B | Q8_0 | 30.5 GB |
| Solar Pro | 22B | Q8_0 | 28 GB |
| Codestral 22B | 22B | Q8_0 | 28 GB |
| gpt-oss-20b | 21B | Q8_0 | 26.7 GB |
| Reka Flash 3 | 21B | Q8_0 | 26.7 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Mixtral 8x7B | 47B | 34.4 GB needed | 36.4 GB |
| Llama 3.1 Nemotron 51B | 51B | 37.3 GB needed | 39.3 GB |
| Jamba 1.5 Mini / Large | 52B | 38.1 GB needed | 40.1 GB |
How to read this
The NVIDIA RTX 5090 features 32 GB of GDDR7 memory. This memory size determines which local AI models can run entirely on your graphics hardware. When a model fits completely within this 32 GB frame, you get the fastest possible generation speeds. If a model exceeds this limit, you must use CPU offload to system RAM, which slows down performance.
To fit larger models into the 32 GB memory space, we use quantized versions. The quant column shows the specific quantization level that maximizes model smartness while keeping the memory footprint under the limit. For example, the Seed-OSS 36B model fits at the Q5_K_M quant, using 30.7 GB of memory. Similarly, the Qwen3.6-35B-A3B and Command R (35B) models both run at the Q5_K_M quant, using 29.8 GB of memory.
Other high performance models fit comfortably within this hardware profile. The Yi 1.5 9B / 34B, Granite Code 3B to 34B, LLaVA 1.5 / 1.6 (7B to 34B), and Ovis 2 models all use 29 GB of memory at the Q5_K_M quant. The DeepSeek-Coder 1.3B / 6.7B / 33B and WizardCoder 33B models require 28.1 GB of memory at the Q5_K_M quant. The OTel 2.0 LLM 31B IT model utilizes 31.4 GB at the Q5_K_M quant.
You can run several models at the higher quality Q6_K quant. The Qwen3 8B / 14B / 32B, Qwen3.5 (dense variants), Aya Expanse 8B / 32B, Granite 4.0 Small/Tiny, and Qwen2.5-Coder 0.5B to 32B models all use 31.5 GB of memory at Q6_K. The Qwen3-30B-A3B and Qwen3-Coder 30B-A3B models use 29.5 GB at Q6_K. The Gemma 3 27B, Gemma 3 4B/12B/27B (vision), and Wan 2.2 / 2.5 models use 26.6 GB at Q6_K. The Gemma 4 26B-A4B and Gemma 4 (all sizes) models use 25.6 GB at Q6_K.
For maximum precision, some models run at the Q8_0 quant. The Aria model uses 31.8 GB of memory at Q8_0. The Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 models use 30.5 GB at Q8_0. The Solar Pro and Codestral 22B models use 28 GB at Q8_0. The gpt-oss-20b and Reka Flash 3 models use 26.7 GB at Q8_0.
When you want to run even larger models, you must offload layers to your system RAM. This offload process requires a system with at least 32 GB of system RAM, but larger models require more. For instance, Mixtral 8x7B needs 34.4 GB at Q4_K_M, which requires 36.4 GB of system RAM. Llama 3.1 Nemotron 51B needs 37.3 GB at Q4_K_M, requiring 39.3 GB of system RAM. Jamba 1.5 Mini / Large needs 38.1 GB at Q4_K_M, requiring 40.1 GB of system RAM.
All memory calculations in this guide assume a standard 4k context window. If you increase the context window to process longer documents, the active memory usage will grow. This extra memory demand might require you to use a lower quant or offload some layers to system RAM to prevent out of memory errors.