Best local AI models for AMD AI PRO R9700
32 GB GDDR6. At a 4k context, 183 of the 233 models in our catalog with verified parameter counts fit fully, up to Seed-OSS 36B at 36B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 183 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Seed-OSS 36B | 36B | Q5_K_M | 30.7 GB |
| Qwen3.6-35B-A3B | 35B | Q5_K_M | 29.8 GB |
| Command R (35B) | 35B | Q5_K_M | 29.8 GB |
| Yi 1.5 9B / 34B | 34B | Q5_K_M | 29 GB |
| Granite Code 3B to 34B | 34B | Q5_K_M | 29 GB |
| LLaVA 1.5 / 1.6 (7B to 34B) | 34B | Q5_K_M | 29 GB |
| Ovis 2 | 34B | Q5_K_M | 29 GB |
| DeepSeek-Coder 1.3B / 6.7B / 33B | 33B | Q5_K_M | 28.1 GB |
| WizardCoder 33B | 33B | Q5_K_M | 28.1 GB |
| OTel 2.0 LLM 31B IT | 32.1B | Q5_K_M | 31.4 GB |
| Qwen3 8B / 14B / 32B | 32B | Q6_K | 31.5 GB |
| Qwen3.5 (dense variants) | 32B | Q6_K | 31.5 GB |
| Aya Expanse 8B / 32B | 32B | Q6_K | 31.5 GB |
| Granite 4.0 Small/Tiny | 32B | Q6_K | 31.5 GB |
| Qwen2.5-Coder 0.5B to 32B | 32B | Q6_K | 31.5 GB |
| Qwen3-30B-A3B | 30B | Q6_K | 29.5 GB |
| Qwen3-Coder 30B-A3B | 30B | Q6_K | 29.5 GB |
| Gemma 3 27B | 27B | Q6_K | 26.6 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | Q6_K | 26.6 GB |
| Wan 2.2 / 2.5 | 27B | Q6_K | 26.6 GB |
| Gemma 4 26B-A4B | 26B | Q6_K | 25.6 GB |
| Gemma 4 (all sizes) | 26B | Q6_K | 25.6 GB |
| Aria | 25B | Q8_0 | 31.8 GB |
| Mistral Small 3.2 | 24B | Q8_0 | 30.5 GB |
| Magistral Small | 24B | Q8_0 | 30.5 GB |
| Devstral Small 1.1 | 24B | Q8_0 | 30.5 GB |
| Solar Pro | 22B | Q8_0 | 28 GB |
| Codestral 22B | 22B | Q8_0 | 28 GB |
| gpt-oss-20b | 21B | Q8_0 | 26.7 GB |
| Reka Flash 3 | 21B | Q8_0 | 26.7 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Mixtral 8x7B | 47B | 34.4 GB needed | 36.4 GB |
| Llama 3.1 Nemotron 51B | 51B | 37.3 GB needed | 39.3 GB |
| Jamba 1.5 Mini / Large | 52B | 38.1 GB needed | 40.1 GB |
How to read this
The AMD AI PRO R9700 graphics card features 32 GB GDDR6 memory. This dedicated video memory determines the maximum size of the artificial intelligence models you can run locally. To load a model entirely onto the hardware for fast processing, the model files and the active context memory must fit within this 32 GB limit.
Model files are typically compressed using quantization to save space. The best quant column indicates the highest quality quantization level that fits comfortably inside the available memory. For example, the Seed-OSS 36B model fits at the Q5_K_M quantization level using 30.7 GB of memory. Similarly, Qwen3.6-35B-A3B and Command R (35B) both utilize the Q5_K_M quantization level and consume 29.8 GB of memory.
Many popular model families fit well within this hardware envelope. The Yi 1.5 9B / 34B, Granite Code 3B to 34B, LLaVA 1.5 / 1.6 (7B to 34B), and Ovis 2 models all run at the Q5_K_M quantization level using 29 GB of memory. DeepSeek-Coder 1.3B / 6.7B / 33B and WizardCoder 33B also run at Q5_K_M using 28.1 GB of memory. OTel 2.0 LLM 31B IT fits at Q5_K_M using 31.4 GB of memory.
Other models can run at the higher quality Q6_K quantization level. Qwen3 8B / 14B / 32B, Qwen3.5 (dense variants), Aya Expanse 8B / 32B, Granite 4.0 Small/Tiny, and Qwen2.5-Coder 0.5B to 32B all use 31.5 GB of memory at Q6_K. The Qwen3-30B-A3B and Qwen3-Coder 30B-A3B models use 29.5 GB at Q6_K. Gemma 3 27B, Gemma 3 4B/12B/27B (vision), and Wan 2.2 / 2.5 use 26.6 GB at Q6_K. Gemma 4 26B-A4B and Gemma 4 (all sizes) use 25.6 GB at Q6_K.
For maximum precision, some models can run at the Q8_0 quantization level. Aria uses 31.8 GB of memory at Q8_0. Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 use 30.5 GB at Q8_0. Solar Pro and Codestral 22B use 28 GB at Q8_0. The gpt-oss-20b and Reka Flash 3 models use 26.7 GB at Q8_0.
When a model is too large for the 32 GB video memory, you can offload parts of it to your system CPU. This offloading process requires sufficient system RAM but slows down processing speeds. For example, Mixtral 8x7B needs 34.4 GB of memory at Q4_K_M and requires 36.4 GB of system RAM. Llama 3.1 Nemotron 51B needs 37.3 GB of memory at Q4_K_M and requires 39.3 GB of system RAM. Jamba 1.5 Mini / Large needs 38.1 GB of memory at Q4_K_M and requires 40.1 GB of system RAM.
All memory calculations listed here assume a standard 4k context window. If you increase the context window to process longer documents, the system will require more memory. Running close to the 32 GB limit with a larger context window can cause the system to run out of memory.