Best local AI models for NVIDIA RTX 4000 Ada Generation

20 GB GDDR6. At a 4k context, 166 of the 233 models in our catalog with verified parameter counts fit fully, up to Gemma 3 27B at 27B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 166 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Gemma 3 27B27BQ4_K_M19.8 GB
Gemma 3 4B/12B/27B (vision)27BQ4_K_M19.8 GB
Wan 2.2 / 2.527BQ4_K_M19.8 GB
Gemma 4 26B-A4B26BQ4_K_M19 GB
Gemma 4 (all sizes)26BQ4_K_M19 GB
Aria25BQ4_K_M18.3 GB
Mistral Small 3.224BQ4_K_M17.6 GB
Magistral Small24BQ4_K_M17.6 GB
Devstral Small 1.124BQ4_K_M17.6 GB
Solar Pro22BQ5_K_M18.7 GB
Codestral 22B22BQ5_K_M18.7 GB
gpt-oss-20b21BQ5_K_M17.9 GB
Reka Flash 321BQ5_K_M17.9 GB
Qwen-Image20BQ6_K19.7 GB
Qwen-Image-Edit20BQ6_K19.7 GB
CogVLM219BQ6_K18.7 GB
HunyuanImage 2.1 / 3.017BQ6_K16.7 GB
Ling-Coder-Lite16.8BQ6_K16.5 GB
DeepSeek-Coder-V2 16B / 236B16BQ6_K15.7 GB
Kimi-VL A3B16BQ6_K15.7 GB
Apriel-1.5-15B-Thinker15BQ8_019.1 GB
StarCoder2 3B / 7B / 15B15BQ8_019.1 GB
Qwen2.5 14B14.7BQ8_019.5 GB
Phi-3 Medium14BQ8_017.8 GB
Phi-414BQ8_017.8 GB
Phi-4-reasoning / -plus14BQ8_017.8 GB
Wan 2.2 T2I14BQ8_017.8 GB
Wan 2.1 (1.3B / 14B)14BQ8_017.8 GB
SkyReels V214BQ8_017.8 GB
Vicuna 13B13BQ8_016.5 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at Q4_K_MSystem RAM at 4k
Qwen3-30B-A3B30B22 GB needed24 GB
Qwen3-Coder 30B-A3B30B22 GB needed24 GB
Qwen3 8B / 14B / 32B32B23.4 GB needed25.4 GB
Qwen3.5 (dense variants)32B23.4 GB needed25.4 GB
Aya Expanse 8B / 32B32B23.4 GB needed25.4 GB
Granite 4.0 Small/Tiny32B23.4 GB needed25.4 GB
Qwen2.5-Coder 0.5B to 32B32B23.4 GB needed25.4 GB
OTel 2.0 LLM 31B IT32.1B27.5 GB needed29.5 GB
DeepSeek-Coder 1.3B / 6.7B / 33B33B24.2 GB needed26.2 GB
WizardCoder 33B33B24.2 GB needed26.2 GB

How to read this

The NVIDIA RTX 4000 Ada Generation workstation graphics card features 20 GB GDDR6 of dedicated video memory. This onboard memory capacity determines the size of the artificial intelligence models you can run entirely on the graphics hardware. Keeping the entire model inside the video memory ensures the fastest possible processing speeds for your local inference tasks.

The quantization column indicates the compression level applied to each model. Quantization reduces the precision of model weights to save memory. For this hardware, larger models like Gemma 3 27B, Gemma 4 26B-A4B, and Wan 2.2 use the Q4_K_M quantization to fit within 19.8 GB or 19 GB of video memory. Medium models like Solar Pro 22B and Codestral 22B run at Q5_K_M quantization using 18.7 GB. Smaller models like Qwen2.5 14B and Phi-4 can run at the higher quality Q8_0 quantization using 19.5 GB and 17.8 GB respectively.

When a model size exceeds the 20 GB limit of your graphics card, you can use CPU offload. This technique splits the model layers between your video memory and your system memory. For example, running Qwen3 32B or Aya Expanse 32B requires 23.4 GB of memory at Q4_K_M quantization. This setup requires 25.4 GB of system RAM to function. Similarly, DeepSeek-Coder 33B requires 24.2 GB at Q4_K_M quantization and needs 26.2 GB of system RAM.

CPU offload comes with a significant performance cost. Transferring data between the system RAM and the graphics card over the system bus is much slower than reading directly from the onboard GDDR6 memory. While offloading allows you to run larger models like OTel 2.0 LLM 31B IT, your generation speed will drop noticeably compared to models that fit entirely on the graphics card.

You must also consider the memory cost of the context window. The memory figures listed for these models assume a standard 4k context window. If you increase the context length to process longer documents or chat histories, the system will require additional video memory to store the active attention cache. This extra memory usage can push a model that normally fits on the card into CPU offload territory.