Best local AI models for NVIDIA RTX A5500
24 GB GDDR6. At a 4k context, 173 of the 233 models in our catalog with verified parameter counts fit fully, up to Qwen3 8B / 14B / 32B at 32B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 173 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Qwen3 8B / 14B / 32B | 32B | Q4_K_M | 23.4 GB |
| Qwen3.5 (dense variants) | 32B | Q4_K_M | 23.4 GB |
| Aya Expanse 8B / 32B | 32B | Q4_K_M | 23.4 GB |
| Granite 4.0 Small/Tiny | 32B | Q4_K_M | 23.4 GB |
| Qwen2.5-Coder 0.5B to 32B | 32B | Q4_K_M | 23.4 GB |
| Qwen3-30B-A3B | 30B | Q4_K_M | 22 GB |
| Qwen3-Coder 30B-A3B | 30B | Q4_K_M | 22 GB |
| Gemma 3 27B | 27B | Q5_K_M | 23 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | Q5_K_M | 23 GB |
| Wan 2.2 / 2.5 | 27B | Q5_K_M | 23 GB |
| Gemma 4 26B-A4B | 26B | Q5_K_M | 22.2 GB |
| Gemma 4 (all sizes) | 26B | Q5_K_M | 22.2 GB |
| Aria | 25B | Q5_K_M | 21.3 GB |
| Mistral Small 3.2 | 24B | Q6_K | 23.6 GB |
| Magistral Small | 24B | Q6_K | 23.6 GB |
| Devstral Small 1.1 | 24B | Q6_K | 23.6 GB |
| Solar Pro | 22B | Q6_K | 21.6 GB |
| Codestral 22B | 22B | Q6_K | 21.6 GB |
| gpt-oss-20b | 21B | Q6_K | 20.7 GB |
| Reka Flash 3 | 21B | Q6_K | 20.7 GB |
| Qwen-Image | 20B | Q6_K | 19.7 GB |
| Qwen-Image-Edit | 20B | Q6_K | 19.7 GB |
| CogVLM2 | 19B | Q6_K | 18.7 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q8_0 | 21.6 GB |
| Ling-Coder-Lite | 16.8B | Q8_0 | 21.4 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q8_0 | 20.4 GB |
| Kimi-VL A3B | 16B | Q8_0 | 20.4 GB |
| Apriel-1.5-15B-Thinker | 15B | Q8_0 | 19.1 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q8_0 | 19.1 GB |
| Qwen2.5 14B | 14.7B | Q8_0 | 19.5 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| OTel 2.0 LLM 31B IT | 32.1B | 27.5 GB needed | 29.5 GB |
| DeepSeek-Coder 1.3B / 6.7B / 33B | 33B | 24.2 GB needed | 26.2 GB |
| WizardCoder 33B | 33B | 24.2 GB needed | 26.2 GB |
| Yi 1.5 9B / 34B | 34B | 24.9 GB needed | 26.9 GB |
| Granite Code 3B to 34B | 34B | 24.9 GB needed | 26.9 GB |
| LLaVA 1.5 / 1.6 (7B to 34B) | 34B | 24.9 GB needed | 26.9 GB |
| Ovis 2 | 34B | 24.9 GB needed | 26.9 GB |
| Qwen3.6-35B-A3B | 35B | 25.6 GB needed | 27.6 GB |
| Command R (35B) | 35B | 25.6 GB needed | 27.6 GB |
| Seed-OSS 36B | 36B | 26.4 GB needed | 28.4 GB |
How to read this
The NVIDIA RTX A5500 is a professional workstation graphics card equipped with 24 GB of GDDR6 memory. This dedicated memory determines the maximum size of the artificial intelligence models you can run entirely on the hardware. When a model fits completely within the onboard memory, it processes tokens at maximum speed. If a model exceeds this capacity, you must offload parts of it to your system memory.
To fit larger models into the available space, developers use quantization. The quantization column shows the compression format used to reduce the size of the model weights. For example, the Q4_K_M format compresses weights to approximately four bits. This allows you to run a 32B model like Qwen3, Qwen3.5, Aya Expanse, Granite 4.0, or Qwen2.5-Coder on your hardware. These models use 23.4 GB of memory at this quantization level.
Other models can run at higher precision levels because they have fewer parameters. The 27B models like Gemma 3, Gemma 3 vision, and Wan 2.2 fit within 23 GB using the Q5_K_M quantization. You can run 24B models like Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 at the Q6_K level using 23.6 GB. Smaller models like Qwen2.5 14B can run at the Q8_0 level using 19.5 GB of memory.
When a model requires more than 24 GB of memory, you must offload layers to your system RAM. This offload process allows you to run larger models but reduces processing speed. For these setups, we assume your workstation has 32 GB of system RAM. Under this configuration, you can run a 34B model like Yi 1.5, Granite Code, LLaVA 1.5, LLaVA 1.6, or Ovis 2. These models require 24.9 GB of memory at Q4_K_M and use 26.9 GB of system RAM.
Memory calculations in this guide assume a standard 4k context window. The context window is the active memory used for your current conversation history. If you increase the context window beyond 4k tokens, the model will require more memory. This extra memory usage might force you to use a lower quantization level or offload more layers to your system RAM.