Best local AI models for NVIDIA RTX 4000 Ada Generation
20 GB GDDR6. At a 4k context, 166 of the 233 models in our catalog with verified parameter counts fit fully, up to Gemma 3 27B at 27B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 166 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Gemma 3 27B | 27B | Q4_K_M | 19.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | Q4_K_M | 19.8 GB |
| Wan 2.2 / 2.5 | 27B | Q4_K_M | 19.8 GB |
| Gemma 4 26B-A4B | 26B | Q4_K_M | 19 GB |
| Gemma 4 (all sizes) | 26B | Q4_K_M | 19 GB |
| Aria | 25B | Q4_K_M | 18.3 GB |
| Mistral Small 3.2 | 24B | Q4_K_M | 17.6 GB |
| Magistral Small | 24B | Q4_K_M | 17.6 GB |
| Devstral Small 1.1 | 24B | Q4_K_M | 17.6 GB |
| Solar Pro | 22B | Q5_K_M | 18.7 GB |
| Codestral 22B | 22B | Q5_K_M | 18.7 GB |
| gpt-oss-20b | 21B | Q5_K_M | 17.9 GB |
| Reka Flash 3 | 21B | Q5_K_M | 17.9 GB |
| Qwen-Image | 20B | Q6_K | 19.7 GB |
| Qwen-Image-Edit | 20B | Q6_K | 19.7 GB |
| CogVLM2 | 19B | Q6_K | 18.7 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q6_K | 16.7 GB |
| Ling-Coder-Lite | 16.8B | Q6_K | 16.5 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q6_K | 15.7 GB |
| Kimi-VL A3B | 16B | Q6_K | 15.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q8_0 | 19.1 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q8_0 | 19.1 GB |
| Qwen2.5 14B | 14.7B | Q8_0 | 19.5 GB |
| Phi-3 Medium | 14B | Q8_0 | 17.8 GB |
| Phi-4 | 14B | Q8_0 | 17.8 GB |
| Phi-4-reasoning / -plus | 14B | Q8_0 | 17.8 GB |
| Wan 2.2 T2I | 14B | Q8_0 | 17.8 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q8_0 | 17.8 GB |
| SkyReels V2 | 14B | Q8_0 | 17.8 GB |
| Vicuna 13B | 13B | Q8_0 | 16.5 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Qwen3-30B-A3B | 30B | 22 GB needed | 24 GB |
| Qwen3-Coder 30B-A3B | 30B | 22 GB needed | 24 GB |
| Qwen3 8B / 14B / 32B | 32B | 23.4 GB needed | 25.4 GB |
| Qwen3.5 (dense variants) | 32B | 23.4 GB needed | 25.4 GB |
| Aya Expanse 8B / 32B | 32B | 23.4 GB needed | 25.4 GB |
| Granite 4.0 Small/Tiny | 32B | 23.4 GB needed | 25.4 GB |
| Qwen2.5-Coder 0.5B to 32B | 32B | 23.4 GB needed | 25.4 GB |
| OTel 2.0 LLM 31B IT | 32.1B | 27.5 GB needed | 29.5 GB |
| DeepSeek-Coder 1.3B / 6.7B / 33B | 33B | 24.2 GB needed | 26.2 GB |
| WizardCoder 33B | 33B | 24.2 GB needed | 26.2 GB |
How to read this
The NVIDIA RTX 4000 Ada Generation workstation graphics card features 20 GB GDDR6 of dedicated video memory. This onboard memory capacity determines the size of the artificial intelligence models you can run entirely on the graphics hardware. Keeping the entire model inside the video memory ensures the fastest possible processing speeds for your local inference tasks.
The quantization column indicates the compression level applied to each model. Quantization reduces the precision of model weights to save memory. For this hardware, larger models like Gemma 3 27B, Gemma 4 26B-A4B, and Wan 2.2 use the Q4_K_M quantization to fit within 19.8 GB or 19 GB of video memory. Medium models like Solar Pro 22B and Codestral 22B run at Q5_K_M quantization using 18.7 GB. Smaller models like Qwen2.5 14B and Phi-4 can run at the higher quality Q8_0 quantization using 19.5 GB and 17.8 GB respectively.
When a model size exceeds the 20 GB limit of your graphics card, you can use CPU offload. This technique splits the model layers between your video memory and your system memory. For example, running Qwen3 32B or Aya Expanse 32B requires 23.4 GB of memory at Q4_K_M quantization. This setup requires 25.4 GB of system RAM to function. Similarly, DeepSeek-Coder 33B requires 24.2 GB at Q4_K_M quantization and needs 26.2 GB of system RAM.
CPU offload comes with a significant performance cost. Transferring data between the system RAM and the graphics card over the system bus is much slower than reading directly from the onboard GDDR6 memory. While offloading allows you to run larger models like OTel 2.0 LLM 31B IT, your generation speed will drop noticeably compared to models that fit entirely on the graphics card.
You must also consider the memory cost of the context window. The memory figures listed for these models assume a standard 4k context window. If you increase the context length to process longer documents or chat histories, the system will require additional video memory to store the active attention cache. This extra memory usage can push a model that normally fits on the card into CPU offload territory.