Best local AI models for NVIDIA RTX 4090 D

24 GB GDDR6X. At a 4k context, 173 of the 233 models in our catalog with verified parameter counts fit fully, up to Qwen3 8B / 14B / 32B at 32B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 173 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Qwen3 8B / 14B / 32B32BQ4_K_M23.4 GB
Qwen3.5 (dense variants)32BQ4_K_M23.4 GB
Aya Expanse 8B / 32B32BQ4_K_M23.4 GB
Granite 4.0 Small/Tiny32BQ4_K_M23.4 GB
Qwen2.5-Coder 0.5B to 32B32BQ4_K_M23.4 GB
Qwen3-30B-A3B30BQ4_K_M22 GB
Qwen3-Coder 30B-A3B30BQ4_K_M22 GB
Gemma 3 27B27BQ5_K_M23 GB
Gemma 3 4B/12B/27B (vision)27BQ5_K_M23 GB
Wan 2.2 / 2.527BQ5_K_M23 GB
Gemma 4 26B-A4B26BQ5_K_M22.2 GB
Gemma 4 (all sizes)26BQ5_K_M22.2 GB
Aria25BQ5_K_M21.3 GB
Mistral Small 3.224BQ6_K23.6 GB
Magistral Small24BQ6_K23.6 GB
Devstral Small 1.124BQ6_K23.6 GB
Solar Pro22BQ6_K21.6 GB
Codestral 22B22BQ6_K21.6 GB
gpt-oss-20b21BQ6_K20.7 GB
Reka Flash 321BQ6_K20.7 GB
Qwen-Image20BQ6_K19.7 GB
Qwen-Image-Edit20BQ6_K19.7 GB
CogVLM219BQ6_K18.7 GB
HunyuanImage 2.1 / 3.017BQ8_021.6 GB
Ling-Coder-Lite16.8BQ8_021.4 GB
DeepSeek-Coder-V2 16B / 236B16BQ8_020.4 GB
Kimi-VL A3B16BQ8_020.4 GB
Apriel-1.5-15B-Thinker15BQ8_019.1 GB
StarCoder2 3B / 7B / 15B15BQ8_019.1 GB
Qwen2.5 14B14.7BQ8_019.5 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at Q4_K_MSystem RAM at 4k
OTel 2.0 LLM 31B IT32.1B27.5 GB needed29.5 GB
DeepSeek-Coder 1.3B / 6.7B / 33B33B24.2 GB needed26.2 GB
WizardCoder 33B33B24.2 GB needed26.2 GB
Yi 1.5 9B / 34B34B24.9 GB needed26.9 GB
Granite Code 3B to 34B34B24.9 GB needed26.9 GB
LLaVA 1.5 / 1.6 (7B to 34B)34B24.9 GB needed26.9 GB
Ovis 234B24.9 GB needed26.9 GB
Qwen3.6-35B-A3B35B25.6 GB needed27.6 GB
Command R (35B)35B25.6 GB needed27.6 GB
Seed-OSS 36B36B26.4 GB needed28.4 GB

How to read this

The NVIDIA RTX 4090 D graphics card features 24 GB of GDDR6X onboard memory. This dedicated memory determines the maximum size of the artificial intelligence models you can run locally. To run a model entirely on the graphics hardware, the model files and the active processing data must fit within this 24 GB limit. If a model exceeds this capacity, the system must use slower system memory, which reduces processing speeds.

Quantization is a method that compresses model files to save space. The quantization column shows the best compression level that fits within your hardware limits. For example, Qwen3 32B, Qwen3.5 32B, Aya Expanse 32B, Granite 4.0 32B, and Qwen2.5-Coder 32B fit using the Q4_K_M quantization, which uses 23.4 GB of memory. Similarly, Qwen3-30B-A3B and Qwen3-Coder 30B-A3B fit at Q4_K_M while using 22 GB of memory.

Higher precision quantizations are possible with slightly smaller models. Gemma 3 27B, Gemma 3 27B vision, and Wan 2.2 / 2.5 fit using the Q5_K_M quantization, which uses 23 GB of memory. Gemma 4 26B-A4B and Gemma 4 26B-A4B all sizes also use Q5_K_M and consume 22.2 GB of memory. Aria uses Q5_K_M and requires 21.3 GB of memory. Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 can run at the higher Q6_K quantization, using 23.6 GB of memory.

Other models fit comfortably within the 24 GB limit at high precision. Solar Pro and Codestral 22B use 21.6 GB of memory at Q6_K. The gpt-oss-20b and Reka Flash 3 models use 20.7 GB at Q6_K. Qwen-Image and Qwen-Image-Edit use 19.7 GB at Q6_K, while CogVLM2 uses 18.7 GB at Q6_K. For maximum precision, HunyuanImage 2.1 / 3.0 uses 21.6 GB at Q8_0, Ling-Coder-Lite uses 21.4 GB at Q8_0, DeepSeek-Coder-V2 16B uses 20.4 GB at Q8_0, Kimi-VL A3B uses 20.4 GB at Q8_0, Apriel-1.5-15B-Thinker uses 19.1 GB at Q8_0, StarCoder2 15B uses 19.1 GB at Q8_0, and Qwen2.5 14B uses 19.5 GB at Q8_0.

When a model is too large for the graphics card, you can offload parts of it to your system RAM. This offload process requires a system with 32 GB of system RAM. For instance, OTel 2.0 LLM 31B IT needs 27.5 GB at Q4_K_M and uses 29.5 GB of system RAM. DeepSeek-Coder 33B and WizardCoder 33B need 24.2 GB at Q4_K_M and use 26.2 GB of system RAM. Yi 1.5 34B, Granite Code 34B, LLaVA 1.5 / 1.6 34B, and Ovis 2 need 24.9 GB at Q4_K_M and use 26.9 GB of system RAM. Qwen3.6-35B-A3B and Command R 35B need 25.6 GB at Q4_K_M and use 27.6 GB of system RAM, while Seed-OSS 36B needs 26.4 GB at Q4_K_M and uses 28.4 GB of system RAM.

Offloading models to system RAM comes with a performance cost. Sharing data between the graphics card and system memory is much slower than keeping everything on the graphics card. Additionally, these memory calculations are based on a standard 4k context window. If you increase the context window to process longer texts, the model will require significantly more memory, which may force you to use smaller models or lower quantization levels.