Best local AI models for NVIDIA RTX A5000

24 GB GDDR6. At a 4k context, 173 of the 233 models in our catalog with verified parameter counts fit fully, up to Qwen3 8B / 14B / 32B at 32B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 173 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Qwen3 8B / 14B / 32B32BQ4_K_M23.4 GB
Qwen3.5 (dense variants)32BQ4_K_M23.4 GB
Aya Expanse 8B / 32B32BQ4_K_M23.4 GB
Granite 4.0 Small/Tiny32BQ4_K_M23.4 GB
Qwen2.5-Coder 0.5B to 32B32BQ4_K_M23.4 GB
Qwen3-30B-A3B30BQ4_K_M22 GB
Qwen3-Coder 30B-A3B30BQ4_K_M22 GB
Gemma 3 27B27BQ5_K_M23 GB
Gemma 3 4B/12B/27B (vision)27BQ5_K_M23 GB
Wan 2.2 / 2.527BQ5_K_M23 GB
Gemma 4 26B-A4B26BQ5_K_M22.2 GB
Gemma 4 (all sizes)26BQ5_K_M22.2 GB
Aria25BQ5_K_M21.3 GB
Mistral Small 3.224BQ6_K23.6 GB
Magistral Small24BQ6_K23.6 GB
Devstral Small 1.124BQ6_K23.6 GB
Solar Pro22BQ6_K21.6 GB
Codestral 22B22BQ6_K21.6 GB
gpt-oss-20b21BQ6_K20.7 GB
Reka Flash 321BQ6_K20.7 GB
Qwen-Image20BQ6_K19.7 GB
Qwen-Image-Edit20BQ6_K19.7 GB
CogVLM219BQ6_K18.7 GB
HunyuanImage 2.1 / 3.017BQ8_021.6 GB
Ling-Coder-Lite16.8BQ8_021.4 GB
DeepSeek-Coder-V2 16B / 236B16BQ8_020.4 GB
Kimi-VL A3B16BQ8_020.4 GB
Apriel-1.5-15B-Thinker15BQ8_019.1 GB
StarCoder2 3B / 7B / 15B15BQ8_019.1 GB
Qwen2.5 14B14.7BQ8_019.5 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at Q4_K_MSystem RAM at 4k
OTel 2.0 LLM 31B IT32.1B27.5 GB needed29.5 GB
DeepSeek-Coder 1.3B / 6.7B / 33B33B24.2 GB needed26.2 GB
WizardCoder 33B33B24.2 GB needed26.2 GB
Yi 1.5 9B / 34B34B24.9 GB needed26.9 GB
Granite Code 3B to 34B34B24.9 GB needed26.9 GB
LLaVA 1.5 / 1.6 (7B to 34B)34B24.9 GB needed26.9 GB
Ovis 234B24.9 GB needed26.9 GB
Qwen3.6-35B-A3B35B25.6 GB needed27.6 GB
Command R (35B)35B25.6 GB needed27.6 GB
Seed-OSS 36B36B26.4 GB needed28.4 GB

How to read this

The NVIDIA RTX A5000 is equipped with 24 GB of GDDR6 memory. This dedicated VRAM determines the maximum size of the artificial intelligence models you can run locally. To run a model entirely on the graphics card, the model files and the active context data must fit inside this 24 GB limit. If a model exceeds this capacity, it will fail to load or require system memory offloading which slows down processing speeds.

The quantization column indicates the compression level applied to each model. Raw models are often too large for local hardware so they are compressed into smaller formats like Q4_K_M, Q5_K_M, Q6_K, or Q8_0. A lower quantization level like Q4_K_M reduces the memory footprint significantly but may slightly degrade output quality. A higher quantization level like Q8_0 preserves original model accuracy but demands much more VRAM.

For models that fit entirely within the 24 GB VRAM of the NVIDIA RTX A5000, you can run large architectures. The Qwen3 32B, Qwen3.5 dense variants 32B, Aya Expanse 32B, Granite 4.0 32B, and Qwen2.5-Coder 32B all fit using the Q4_K_M quantization which consumes 23.4 GB of memory. You can also run the Gemma 3 27B, Gemma 3 vision 27B, and Wan 2.2 or 2.5 27B models at the Q5_K_M quantization level using 23 GB of VRAM.

Other high performance options include the Mistral Small 3.2 24B, Magistral Small 24B, and Devstral Small 1.1 24B models which run at the Q6_K quantization level using 23.6 GB of VRAM. If you prefer higher precision, you can run the Qwen2.5 14B model at the Q8_0 quantization level which uses 19.5 GB of VRAM. This leaves a comfortable buffer for system operations and basic context windows.

When a model is too large for the 24 GB VRAM, you must offload parts of it to your system RAM. This requires a system with at least 32 GB of system RAM. For example, the Command R 35B model needs 25.6 GB of memory at Q4_K_M quantization and requires 27.6 GB of system RAM to run. Similarly, the Yi 1.5 34B model needs 24.9 GB of memory at Q4_K_M quantization and requires 26.9 GB of system RAM. Offloading allows you to run these larger models but it introduces a severe performance cost because system RAM is much slower than GDDR6 VRAM.

You must also consider the 4k context caveat when planning your local deployments. The memory figures listed for these models assume a standard 4k context window. If you increase the context length to process longer documents or extended conversations, the memory requirements will rise. This extra memory demand can push a model that normally fits inside the 24 GB VRAM limit into system RAM offloading.