Best local AI models for NVIDIA RTX 4000 SFF Ada Generation

20 GB GDDR6. At a 4k context, 166 of the 233 models in our catalog with verified parameter counts fit fully, up to Gemma 3 27B at 27B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 166 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Gemma 3 27B27BQ4_K_M19.8 GB
Gemma 3 4B/12B/27B (vision)27BQ4_K_M19.8 GB
Wan 2.2 / 2.527BQ4_K_M19.8 GB
Gemma 4 26B-A4B26BQ4_K_M19 GB
Gemma 4 (all sizes)26BQ4_K_M19 GB
Aria25BQ4_K_M18.3 GB
Mistral Small 3.224BQ4_K_M17.6 GB
Magistral Small24BQ4_K_M17.6 GB
Devstral Small 1.124BQ4_K_M17.6 GB
Solar Pro22BQ5_K_M18.7 GB
Codestral 22B22BQ5_K_M18.7 GB
gpt-oss-20b21BQ5_K_M17.9 GB
Reka Flash 321BQ5_K_M17.9 GB
Qwen-Image20BQ6_K19.7 GB
Qwen-Image-Edit20BQ6_K19.7 GB
CogVLM219BQ6_K18.7 GB
HunyuanImage 2.1 / 3.017BQ6_K16.7 GB
Ling-Coder-Lite16.8BQ6_K16.5 GB
DeepSeek-Coder-V2 16B / 236B16BQ6_K15.7 GB
Kimi-VL A3B16BQ6_K15.7 GB
Apriel-1.5-15B-Thinker15BQ8_019.1 GB
StarCoder2 3B / 7B / 15B15BQ8_019.1 GB
Qwen2.5 14B14.7BQ8_019.5 GB
Phi-3 Medium14BQ8_017.8 GB
Phi-414BQ8_017.8 GB
Phi-4-reasoning / -plus14BQ8_017.8 GB
Wan 2.2 T2I14BQ8_017.8 GB
Wan 2.1 (1.3B / 14B)14BQ8_017.8 GB
SkyReels V214BQ8_017.8 GB
Vicuna 13B13BQ8_016.5 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at Q4_K_MSystem RAM at 4k
Qwen3-30B-A3B30B22 GB needed24 GB
Qwen3-Coder 30B-A3B30B22 GB needed24 GB
Qwen3 8B / 14B / 32B32B23.4 GB needed25.4 GB
Qwen3.5 (dense variants)32B23.4 GB needed25.4 GB
Aya Expanse 8B / 32B32B23.4 GB needed25.4 GB
Granite 4.0 Small/Tiny32B23.4 GB needed25.4 GB
Qwen2.5-Coder 0.5B to 32B32B23.4 GB needed25.4 GB
OTel 2.0 LLM 31B IT32.1B27.5 GB needed29.5 GB
DeepSeek-Coder 1.3B / 6.7B / 33B33B24.2 GB needed26.2 GB
WizardCoder 33B33B24.2 GB needed26.2 GB

How to read this

The NVIDIA RTX 4000 SFF Ada Generation is a low profile workstation graphics card. It features 20 GB of GDDR6 dedicated video memory. This memory size determines which artificial intelligence models you can run entirely on the graphics hardware. Keeping the model files inside this 20 GB limit ensures fast processing speeds and low latency.

The quantization column indicates the compression level of each model. Quantization reduces the precision of model weights to save space. For example, Gemma 3 27B, Gemma 3 27B vision, and Wan 2.2 / 2.5 can run at the Q4_K_M quantization level using 19.8 GB of video memory. Gemma 4 26B-A4B and Gemma 4 all sizes also run at Q4_K_M using 19 GB of video memory. Aria fits at Q4_K_M using 18.3 GB of video memory.

Smaller models can run at higher precision levels. Apriel-1.5-15B-Thinker and StarCoder2 15B fit at the Q8_0 quantization level using 19.1 GB of video memory. Qwen2.5 14B fits at Q8_0 using 19.5 GB of video memory. Phi-3 Medium, Phi-4, Phi-4-reasoning / -plus, Wan 2.2 T2I, Wan 2.1 14B, and SkyReels V2 all fit at Q8_0 using 17.8 GB of video memory. Vicuna 13B fits at Q8_0 using 16.5 GB of video memory.

When a model is too large for the 20 GB video memory, you must offload parts of it to the system RAM. This offload process allows you to run larger models but slows down the generation speed significantly. For these setups, we assume a system with 32 GB of system RAM. Qwen3 32B, Qwen3.5 dense variants, Aya Expanse 32B, Granite 4.0 Small/Tiny, and Qwen2.5-Coder 32B need 23.4 GB of memory at Q4_K_M, which requires 25.4 GB of system RAM.

Other offload options include Qwen3-30B-A3B and Qwen3-Coder 30B-A3B, which need 22 GB at Q4_K_M and require 24 GB of system RAM. OTel 2.0 LLM 31B IT needs 27.5 GB at Q4_K_M and requires 29.5 GB of system RAM. DeepSeek-Coder 33B and WizardCoder 33B need 24.2 GB at Q4_K_M and require 26.2 GB of system RAM.

You must also consider the context window size when loading these models. The memory usage figures listed here are calculated using a baseline 4k context window. If you increase the context window to process longer documents, the system will require additional video memory. This extra memory requirement might force you to use a lower quantization level or offload more layers to system RAM.