Best local AI models for NVIDIA A10G
24 GB GDDR6. At a 4k context, 173 of the 233 models in our catalog with verified parameter counts fit fully, up to Qwen3 8B / 14B / 32B at 32B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 173 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Qwen3 8B / 14B / 32B | 32B | Q4_K_M | 23.4 GB |
| Qwen3.5 (dense variants) | 32B | Q4_K_M | 23.4 GB |
| Aya Expanse 8B / 32B | 32B | Q4_K_M | 23.4 GB |
| Granite 4.0 Small/Tiny | 32B | Q4_K_M | 23.4 GB |
| Qwen2.5-Coder 0.5B to 32B | 32B | Q4_K_M | 23.4 GB |
| Qwen3-30B-A3B | 30B | Q4_K_M | 22 GB |
| Qwen3-Coder 30B-A3B | 30B | Q4_K_M | 22 GB |
| Gemma 3 27B | 27B | Q5_K_M | 23 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | Q5_K_M | 23 GB |
| Wan 2.2 / 2.5 | 27B | Q5_K_M | 23 GB |
| Gemma 4 26B-A4B | 26B | Q5_K_M | 22.2 GB |
| Gemma 4 (all sizes) | 26B | Q5_K_M | 22.2 GB |
| Aria | 25B | Q5_K_M | 21.3 GB |
| Mistral Small 3.2 | 24B | Q6_K | 23.6 GB |
| Magistral Small | 24B | Q6_K | 23.6 GB |
| Devstral Small 1.1 | 24B | Q6_K | 23.6 GB |
| Solar Pro | 22B | Q6_K | 21.6 GB |
| Codestral 22B | 22B | Q6_K | 21.6 GB |
| gpt-oss-20b | 21B | Q6_K | 20.7 GB |
| Reka Flash 3 | 21B | Q6_K | 20.7 GB |
| Qwen-Image | 20B | Q6_K | 19.7 GB |
| Qwen-Image-Edit | 20B | Q6_K | 19.7 GB |
| CogVLM2 | 19B | Q6_K | 18.7 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q8_0 | 21.6 GB |
| Ling-Coder-Lite | 16.8B | Q8_0 | 21.4 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q8_0 | 20.4 GB |
| Kimi-VL A3B | 16B | Q8_0 | 20.4 GB |
| Apriel-1.5-15B-Thinker | 15B | Q8_0 | 19.1 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q8_0 | 19.1 GB |
| Qwen2.5 14B | 14.7B | Q8_0 | 19.5 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| OTel 2.0 LLM 31B IT | 32.1B | 27.5 GB needed | 29.5 GB |
| DeepSeek-Coder 1.3B / 6.7B / 33B | 33B | 24.2 GB needed | 26.2 GB |
| WizardCoder 33B | 33B | 24.2 GB needed | 26.2 GB |
| Yi 1.5 9B / 34B | 34B | 24.9 GB needed | 26.9 GB |
| Granite Code 3B to 34B | 34B | 24.9 GB needed | 26.9 GB |
| LLaVA 1.5 / 1.6 (7B to 34B) | 34B | 24.9 GB needed | 26.9 GB |
| Ovis 2 | 34B | 24.9 GB needed | 26.9 GB |
| Qwen3.6-35B-A3B | 35B | 25.6 GB needed | 27.6 GB |
| Command R (35B) | 35B | 25.6 GB needed | 27.6 GB |
| Seed-OSS 36B | 36B | 26.4 GB needed | 28.4 GB |
How to read this
The NVIDIA A10G graphics card features 24 GB of GDDR6 memory. This physical memory size determines the maximum size of the artificial intelligence models you can run locally. To fit larger models into this memory limit, developers use quantization. Quantization reduces the precision of model weights to make the files smaller while keeping most of the model intelligence.
The best quant column shows the optimal quantization level for each model on this hardware. For models up to 32B parameters, a Q4_K_M quantization is common. This level uses 4 bit precision and fits models like Qwen3 32B, Qwen3.5 dense variants 32B, Aya Expanse 32B, Granite 4.0 Small or Tiny 32B, and Qwen2.5-Coder 32B into 23.4 GB of memory. Other models like Qwen3-30B-A3B and Qwen3-Coder 30B-A3B fit at Q4_K_M using 22 GB of memory.
For slightly smaller models, you can use higher precision quantization levels like Q5_K_M or Q6_K. Gemma 3 27B, Gemma 3 vision 27B, and Wan 2.2 or 2.5 27B use Q5_K_M quantization and consume 23 GB of memory. Gemma 4 26B-A4B and Gemma 4 all sizes 26B use Q5_K_M to fit in 22.2 GB of memory. Aria 25B uses Q5_K_M and consumes 21.3 GB. Mistral Small 3.2 24B, Magistral Small 24B, and Devstral Small 1.1 24B use Q6_K quantization and consume 23.6 GB of memory. Solar Pro 22B and Codestral 22B use Q6_K to fit in 21.6 GB. Models like gpt-oss-20b 21B and Reka Flash 3 21B use Q6_K and consume 20.7 GB. Qwen-Image 20B and Qwen-Image-Edit 20B use Q6_K and consume 19.7 GB. CogVLM2 19B uses Q6_K and consumes 18.7 GB.
The highest precision level is Q8_0 quantization. HunyuanImage 2.1 or 3.0 17B uses Q8_0 and consumes 21.6 GB of memory. Ling-Coder-Lite 16.8B uses Q8_0 and consumes 21.4 GB. DeepSeek-Coder-V2 16B and Kimi-VL A3B 16B use Q8_0 and consume 20.4 GB. Apriel-1.5-15B-Thinker 15B and StarCoder2 15B use Q8_0 and consume 19.1 GB. Qwen2.5 14B is a 14.7B model that uses Q8_0 and consumes 19.5 GB.
When a model is too large for the 24 GB graphics memory, you can offload parts of it to your system RAM. This CPU offload process requires a system with at least 32 GB of RAM. Offloading allows you to run larger models, but it costs performance because system RAM is much slower than graphics memory. For example, OTel 2.0 LLM 31B IT is a 32.1B model that needs 27.5 GB at Q4_K_M and requires 29.5 GB of system RAM. DeepSeek-Coder 33B and WizardCoder 33B need 24.2 GB at Q4_K_M and require 26.2 GB of system RAM. Yi 1.5 34B, Granite Code 34B, LLaVA 1.5 or 1.6 34B, and Ovis 2 34B need 24.9 GB at Q4_K_M and require 26.9 GB of system RAM. Qwen3.6-35B-A3B 35B and Command R 35B need 25.6 GB at Q4_K_M and require 27.6 GB of system RAM. Seed-OSS 36B needs 26.4 GB at Q4_K_M and requires 28.4 GB of system RAM.
You must also consider the context window size. The memory numbers listed here assume a standard 4k context window. If you increase the context window to process longer documents, the system will require more memory for the attention cache. This extra memory requirement might force you to use a lower quantization level or offload more layers to system RAM.