Best local AI models for NVIDIA RTX A4000H
16 GB GDDR6. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| gpt-oss-20b | 21B | Q4_K_M | 15.4 GB |
| Reka Flash 3 | 21B | Q4_K_M | 15.4 GB |
| Qwen-Image | 20B | Q4_K_M | 14.6 GB |
| Qwen-Image-Edit | 20B | Q4_K_M | 14.6 GB |
| CogVLM2 | 19B | Q4_K_M | 13.9 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q5_K_M | 14.5 GB |
| Ling-Coder-Lite | 16.8B | Q5_K_M | 14.3 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q6_K | 15.7 GB |
| Kimi-VL A3B | 16B | Q6_K | 15.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q6_K | 14.8 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q6_K | 14.8 GB |
| Qwen2.5 14B | 14.7B | Q6_K | 15.3 GB |
| Phi-3 Medium | 14B | Q6_K | 13.8 GB |
| Phi-4 | 14B | Q6_K | 13.8 GB |
| Phi-4-reasoning / -plus | 14B | Q6_K | 13.8 GB |
| Wan 2.2 T2I | 14B | Q6_K | 13.8 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q6_K | 13.8 GB |
| SkyReels V2 | 14B | Q6_K | 13.8 GB |
| Vicuna 13B | 13B | Q6_K | 12.8 GB |
| HunyuanVideo | 13B | Q6_K | 12.8 GB |
| HunyuanVideo-Avatar | 13B | Q6_K | 12.8 GB |
| LTX-Video / LTX-2 | 13B | Q6_K | 12.8 GB |
| FramePack | 13B | Q6_K | 12.8 GB |
| FLUX.1 dev | 12B | FP8 / optimized | 14.4 GB |
| Gemma 3 12B | 12B | Q8_0 | 15.3 GB |
| Gemma 4 12B | 12B | Q8_0 | 15.3 GB |
| Mistral NeMo 12B | 12B | Q8_0 | 15.3 GB |
| Pixtral 12B | 12B | Q8_0 | 15.3 GB |
| FLUX.1 schnell | 12B | Q8_0 | 15.3 GB |
| FLUX.1 Kontext dev | 12B | Q8_0 | 15.3 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
| Mistral Small 3.2 | 24B | 17.6 GB needed | 19.6 GB |
| Magistral Small | 24B | 17.6 GB needed | 19.6 GB |
| Devstral Small 1.1 | 24B | 17.6 GB needed | 19.6 GB |
| Aria | 25B | 18.3 GB needed | 20.3 GB |
| Gemma 4 26B-A4B | 26B | 19 GB needed | 21 GB |
| Gemma 4 (all sizes) | 26B | 19 GB needed | 21 GB |
| Gemma 3 27B | 27B | 19.8 GB needed | 21.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | 19.8 GB needed | 21.8 GB |
How to read this
The NVIDIA RTX A4000H graphics card features 16 GB of GDDR6 memory. This dedicated video memory determines the maximum size of the artificial intelligence models you can run locally. To load a model entirely on the graphics card, the model files and the active memory must fit within this 16 GB limit. Running models fully in video memory ensures the fastest processing speeds.
The quantization column shows the compression level used to fit these models into memory. Quantization reduces the precision of model weights to save space. For example, Q4_K_M represents a four bit quantization that balances size and quality. Higher quantizations like Q6_K or Q8_0 offer better accuracy but require more memory. The gpt-oss-20b and Reka Flash 3 models utilize Q4_K_M to fit their 21B parameters into 15.4 GB of video memory.
For models that exceed 16 GB, you can use CPU offloading if your computer has at least 32 GB of system RAM. This technique splits the model between the graphics card and your system memory. Solar Pro and Codestral 22B require 16.1 GB of memory at Q4_K_M, which needs 18.1 GB of system RAM. Larger models like Gemma 3 27B require 19.8 GB at Q4_K_M, which needs 21.8 GB of system RAM.
CPU offloading allows you to run larger models like Mistral Small 3.2, Magistral Small, and Devstral Small 1.1. These 24B models need 17.6 GB at Q4_K_M and require 19.6 GB of system RAM. Aria needs 18.3 GB at Q4_K_M and requires 20.3 GB of system RAM. While offloading enables these larger models to run, it reduces processing speed because system RAM is slower than GDDR6 video memory.
You can run several highly capable models entirely within the 16 GB video memory limit. DeepSeek-Coder-V2 16B and Kimi-VL A3B fit into 15.7 GB using Q6_K. Qwen2.5 14B fits into 15.3 GB using Q6_K. Image and video models like FLUX.1 dev fit into 14.4 GB using FP8 or optimized settings. Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, and Pixtral 12B fit into 15.3 GB using Q8_0.
Memory calculations must also account for the context window. The listed memory figures assume a standard 4k context window. If you increase the context window to process longer documents or chat histories, the memory usage will rise. This extra memory requirement may force you to use a lower quantization level or switch to CPU offloading to prevent out of memory errors.