Best local AI models for NVIDIA RTX 3090 Ti
24 GB GDDR6X. At a 4k context, 173 of the 233 models in our catalog with verified parameter counts fit fully, up to Qwen3 8B / 14B / 32B at 32B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 173 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Qwen3 8B / 14B / 32B | 32B | Q4_K_M | 23.4 GB |
| Qwen3.5 (dense variants) | 32B | Q4_K_M | 23.4 GB |
| Aya Expanse 8B / 32B | 32B | Q4_K_M | 23.4 GB |
| Granite 4.0 Small/Tiny | 32B | Q4_K_M | 23.4 GB |
| Qwen2.5-Coder 0.5B to 32B | 32B | Q4_K_M | 23.4 GB |
| Qwen3-30B-A3B | 30B | Q4_K_M | 22 GB |
| Qwen3-Coder 30B-A3B | 30B | Q4_K_M | 22 GB |
| Gemma 3 27B | 27B | Q5_K_M | 23 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | Q5_K_M | 23 GB |
| Wan 2.2 / 2.5 | 27B | Q5_K_M | 23 GB |
| Gemma 4 26B-A4B | 26B | Q5_K_M | 22.2 GB |
| Gemma 4 (all sizes) | 26B | Q5_K_M | 22.2 GB |
| Aria | 25B | Q5_K_M | 21.3 GB |
| Mistral Small 3.2 | 24B | Q6_K | 23.6 GB |
| Magistral Small | 24B | Q6_K | 23.6 GB |
| Devstral Small 1.1 | 24B | Q6_K | 23.6 GB |
| Solar Pro | 22B | Q6_K | 21.6 GB |
| Codestral 22B | 22B | Q6_K | 21.6 GB |
| gpt-oss-20b | 21B | Q6_K | 20.7 GB |
| Reka Flash 3 | 21B | Q6_K | 20.7 GB |
| Qwen-Image | 20B | Q6_K | 19.7 GB |
| Qwen-Image-Edit | 20B | Q6_K | 19.7 GB |
| CogVLM2 | 19B | Q6_K | 18.7 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q8_0 | 21.6 GB |
| Ling-Coder-Lite | 16.8B | Q8_0 | 21.4 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q8_0 | 20.4 GB |
| Kimi-VL A3B | 16B | Q8_0 | 20.4 GB |
| Apriel-1.5-15B-Thinker | 15B | Q8_0 | 19.1 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q8_0 | 19.1 GB |
| Qwen2.5 14B | 14.7B | Q8_0 | 19.5 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| OTel 2.0 LLM 31B IT | 32.1B | 27.5 GB needed | 29.5 GB |
| DeepSeek-Coder 1.3B / 6.7B / 33B | 33B | 24.2 GB needed | 26.2 GB |
| WizardCoder 33B | 33B | 24.2 GB needed | 26.2 GB |
| Yi 1.5 9B / 34B | 34B | 24.9 GB needed | 26.9 GB |
| Granite Code 3B to 34B | 34B | 24.9 GB needed | 26.9 GB |
| LLaVA 1.5 / 1.6 (7B to 34B) | 34B | 24.9 GB needed | 26.9 GB |
| Ovis 2 | 34B | 24.9 GB needed | 26.9 GB |
| Qwen3.6-35B-A3B | 35B | 25.6 GB needed | 27.6 GB |
| Command R (35B) | 35B | 25.6 GB needed | 27.6 GB |
| Seed-OSS 36B | 36B | 26.4 GB needed | 28.4 GB |
How to read this
The NVIDIA RTX 3090 Ti features 24 GB of GDDR6X onboard memory. This dedicated video memory determines the size of the artificial intelligence models you can run entirely on your graphics card. When a model fits completely within this limit, you get the fastest possible processing speeds. If a model exceeds this capacity, you must use CPU offloading to system memory, which slows down performance.
The quantization column shows the compression level applied to each model. Quantization reduces the precision of model weights to save memory. For example, the Q4_K_M quant represents a four bit compression level. This allows larger models like Qwen3 32B, Qwen3.5 32B, Aya Expanse 32B, Granite 4.0 32B, and Qwen2.5-Coder 32B to fit into 23.4 GB of video memory. Similarly, Qwen3-30B-A3B and Qwen3-Coder 30B-A3B fit into 22 GB using Q4_K_M.
Higher precision quants require more memory but preserve more original model quality. The Q5_K_M quant is used for Gemma 3 27B, Gemma 3 27B vision, and Wan 2.2 / 2.5, which all consume 23 GB. Gemma 4 26B-A4B and Gemma 4 all sizes use Q5_K_M to fit in 22.2 GB, while Aria uses it to fit in 21.3 GB. The Q6_K quant allows Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 to run at 23.6 GB. Solar Pro and Codestral 22B use Q6_K at 21.6 GB, while gpt-oss-20b and Reka Flash 3 use it at 20.7 GB. Qwen-Image and Qwen-Image-Edit fit in 19.7 GB, and CogVLM2 fits in 18.7 GB using Q6_K.
The highest precision Q8_0 quant is ideal for smaller models. HunyuanImage 2.1 / 3.0 uses Q8_0 at 21.6 GB. Ling-Coder-Lite fits in 21.4 GB, while DeepSeek-Coder-V2 16B and Kimi-VL A3B use Q8_0 at 20.4 GB. Apriel-1.5-15B-Thinker and StarCoder2 15B require 19.1 GB, while Qwen2.5 14.7B uses 19.5 GB under Q8_0. These configurations maximize output quality within your hardware limits.
When a model is too large for the 24 GB limit, you can offload layers to your system RAM. Assuming you have 32 GB of system RAM, you can run OTel 2.0 LLM 31B IT at Q4_K_M, which needs 27.5 GB of memory and 29.5 GB of system RAM. DeepSeek-Coder 33B and WizardCoder 33B require 24.2 GB of memory and 26.2 GB of system RAM. Yi 1.5 34B, Granite Code 34B, LLaVA 1.5 / 1.6 34B, and Ovis 2 need 24.9 GB of memory and 26.9 GB of system RAM. Qwen3.6-35B-A3B and Command R 35B require 25.6 GB of memory and 27.6 GB of system RAM, while Seed-OSS 36B needs 26.4 GB of memory and 28.4 GB of system RAM.
All memory calculations in this guide assume a standard 4k context window. The context window is the amount of text the model can process at one time. If you increase the context window beyond 4k tokens, the memory usage will rise. This extra memory demand can push a model past the 24 GB limit of your graphics card and force slow CPU offloading.