Best local AI models for NVIDIA RTX A4500
20 GB GDDR6. At a 4k context, 166 of the 233 models in our catalog with verified parameter counts fit fully, up to Gemma 3 27B at 27B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 166 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Gemma 3 27B | 27B | Q4_K_M | 19.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | Q4_K_M | 19.8 GB |
| Wan 2.2 / 2.5 | 27B | Q4_K_M | 19.8 GB |
| Gemma 4 26B-A4B | 26B | Q4_K_M | 19 GB |
| Gemma 4 (all sizes) | 26B | Q4_K_M | 19 GB |
| Aria | 25B | Q4_K_M | 18.3 GB |
| Mistral Small 3.2 | 24B | Q4_K_M | 17.6 GB |
| Magistral Small | 24B | Q4_K_M | 17.6 GB |
| Devstral Small 1.1 | 24B | Q4_K_M | 17.6 GB |
| Solar Pro | 22B | Q5_K_M | 18.7 GB |
| Codestral 22B | 22B | Q5_K_M | 18.7 GB |
| gpt-oss-20b | 21B | Q5_K_M | 17.9 GB |
| Reka Flash 3 | 21B | Q5_K_M | 17.9 GB |
| Qwen-Image | 20B | Q6_K | 19.7 GB |
| Qwen-Image-Edit | 20B | Q6_K | 19.7 GB |
| CogVLM2 | 19B | Q6_K | 18.7 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q6_K | 16.7 GB |
| Ling-Coder-Lite | 16.8B | Q6_K | 16.5 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q6_K | 15.7 GB |
| Kimi-VL A3B | 16B | Q6_K | 15.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q8_0 | 19.1 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q8_0 | 19.1 GB |
| Qwen2.5 14B | 14.7B | Q8_0 | 19.5 GB |
| Phi-3 Medium | 14B | Q8_0 | 17.8 GB |
| Phi-4 | 14B | Q8_0 | 17.8 GB |
| Phi-4-reasoning / -plus | 14B | Q8_0 | 17.8 GB |
| Wan 2.2 T2I | 14B | Q8_0 | 17.8 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q8_0 | 17.8 GB |
| SkyReels V2 | 14B | Q8_0 | 17.8 GB |
| Vicuna 13B | 13B | Q8_0 | 16.5 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Qwen3-30B-A3B | 30B | 22 GB needed | 24 GB |
| Qwen3-Coder 30B-A3B | 30B | 22 GB needed | 24 GB |
| Qwen3 8B / 14B / 32B | 32B | 23.4 GB needed | 25.4 GB |
| Qwen3.5 (dense variants) | 32B | 23.4 GB needed | 25.4 GB |
| Aya Expanse 8B / 32B | 32B | 23.4 GB needed | 25.4 GB |
| Granite 4.0 Small/Tiny | 32B | 23.4 GB needed | 25.4 GB |
| Qwen2.5-Coder 0.5B to 32B | 32B | 23.4 GB needed | 25.4 GB |
| OTel 2.0 LLM 31B IT | 32.1B | 27.5 GB needed | 29.5 GB |
| DeepSeek-Coder 1.3B / 6.7B / 33B | 33B | 24.2 GB needed | 26.2 GB |
| WizardCoder 33B | 33B | 24.2 GB needed | 26.2 GB |
How to read this
The NVIDIA RTX A4500 graphics card features 20 GB of GDDR6 dedicated video memory. This memory capacity determines the maximum size of the artificial intelligence models you can run locally. To run a model entirely on your hardware, the model files and the active context data must fit within this 20 GB limit.
The quantization column indicates the compression level used to shrink the model. Uncompressed models are too large for local consumer hardware. Quantization formats like Q4_K_M, Q5_K_M, Q6_K, and Q8_0 reduce the precision of the model weights. This compression allows larger models to fit into the 20 GB memory space of your card while preserving most of their reasoning capabilities.
For maximum performance, you can run models up to 27B parameters completely within your video memory. The Gemma 3 27B model, the Gemma 3 27B vision model, and the Wan 2.2 / 2.5 models fit on the card using the Q4_K_M quantization, which consumes 19.8 GB of memory. Gemma 4 26B-A4B and Gemma 4 all sizes also run locally at Q4_K_M quantization while using 19 GB of memory.
Other models fit comfortably on the card with higher precision levels. You can run the Apriel-1.5-15B-Thinker or StarCoder2 15B models at Q8_0 quantization using 19.1 GB of memory. The Qwen2.5 14B model fits at Q8_0 quantization using 19.5 GB of memory. Smaller models like the Vicuna 13B model use 16.5 GB of memory at Q8_0 quantization.
If you want to run larger models, you must use CPU offloading. This process splits the model between your 20 GB video memory and your system RAM. For example, running the Qwen3 32B, Qwen3.5 dense variants, Aya Expanse 32B, Granite 4.0 Small/Tiny, or Qwen2.5-Coder 32B models requires 23.4 GB of memory at Q4_K_M quantization. This setup requires 25.4 GB of system RAM to handle the offloaded portions.
CPU offloading allows you to run massive models like the DeepSeek-Coder 33B or WizardCoder 33B, which need 24.2 GB of memory at Q4_K_M quantization and 26.2 GB of system RAM. However, offloading comes with a significant speed penalty. Moving data between the system RAM and the graphics card slows down the generation speed compared to running models entirely on the graphics card.
All memory calculations assume a standard 4k context window. As your conversation history grows, the context window consumes additional video memory. If you generate very long responses or upload large documents, the memory usage will exceed the listed figures, which may cause the model to fail or force your system to offload data to the slower system RAM.