Best local AI models for AMD RX 7700
16 GB GDDR6. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| gpt-oss-20b | 21B | Q4_K_M | 15.4 GB |
| Reka Flash 3 | 21B | Q4_K_M | 15.4 GB |
| Qwen-Image | 20B | Q4_K_M | 14.6 GB |
| Qwen-Image-Edit | 20B | Q4_K_M | 14.6 GB |
| CogVLM2 | 19B | Q4_K_M | 13.9 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q5_K_M | 14.5 GB |
| Ling-Coder-Lite | 16.8B | Q5_K_M | 14.3 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q6_K | 15.7 GB |
| Kimi-VL A3B | 16B | Q6_K | 15.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q6_K | 14.8 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q6_K | 14.8 GB |
| Qwen2.5 14B | 14.7B | Q6_K | 15.3 GB |
| Phi-3 Medium | 14B | Q6_K | 13.8 GB |
| Phi-4 | 14B | Q6_K | 13.8 GB |
| Phi-4-reasoning / -plus | 14B | Q6_K | 13.8 GB |
| Wan 2.2 T2I | 14B | Q6_K | 13.8 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q6_K | 13.8 GB |
| SkyReels V2 | 14B | Q6_K | 13.8 GB |
| Vicuna 13B | 13B | Q6_K | 12.8 GB |
| HunyuanVideo | 13B | Q6_K | 12.8 GB |
| HunyuanVideo-Avatar | 13B | Q6_K | 12.8 GB |
| LTX-Video / LTX-2 | 13B | Q6_K | 12.8 GB |
| FramePack | 13B | Q6_K | 12.8 GB |
| FLUX.1 dev | 12B | FP8 / optimized | 14.4 GB |
| Gemma 3 12B | 12B | Q8_0 | 15.3 GB |
| Gemma 4 12B | 12B | Q8_0 | 15.3 GB |
| Mistral NeMo 12B | 12B | Q8_0 | 15.3 GB |
| Pixtral 12B | 12B | Q8_0 | 15.3 GB |
| FLUX.1 schnell | 12B | Q8_0 | 15.3 GB |
| FLUX.1 Kontext dev | 12B | Q8_0 | 15.3 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
| Mistral Small 3.2 | 24B | 17.6 GB needed | 19.6 GB |
| Magistral Small | 24B | 17.6 GB needed | 19.6 GB |
| Devstral Small 1.1 | 24B | 17.6 GB needed | 19.6 GB |
| Aria | 25B | 18.3 GB needed | 20.3 GB |
| Gemma 4 26B-A4B | 26B | 19 GB needed | 21 GB |
| Gemma 4 (all sizes) | 26B | 19 GB needed | 21 GB |
| Gemma 3 27B | 27B | 19.8 GB needed | 21.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | 19.8 GB needed | 21.8 GB |
How to read this
The AMD Radeon RX 7700 graphics card features 16 GB of GDDR6 memory. This onboard memory determines the maximum size of the artificial intelligence models you can run locally. To run a model entirely on your graphics hardware, the model files and the active context data must fit inside this 16 GB limit.
Raw model files are often too large for consumer hardware. Quantization compresses these models to make them smaller. The best quant column shows the highest quality quantization level that still fits within your hardware limits. For example, gpt-oss-20b and Reka Flash 3 are 21B models that fit using the Q4_K_M quantization, which uses 15.4 GB of memory. Qwen-Image and Qwen-Image-Edit are 20B models that use 14.6 GB of memory at the Q4_K_M quantization level. CogVLM2 is a 19B model that fits at Q4_K_M, using 13.9 GB of memory.
Models with slightly smaller parameter counts can run at higher quantization levels for better accuracy. HunyuanImage 2.1 / 3.0 is a 17B model that fits at Q5_K_M, using 14.5 GB of memory. Ling-Coder-Lite is a 16.8B model that fits at Q5_K_M, using 14.3 GB of memory. DeepSeek-Coder-V2 16B / 236B and Kimi-VL A3B are 16B models that run at Q6_K, using 15.7 GB of memory. Apriel-1.5-15B-Thinker and StarCoder2 3B / 7B / 15B are 15B models that use 14.8 GB of memory at Q6_K. Qwen2.5 14B uses 15.3 GB at Q6_K. Several 14B models like Phi-3 Medium, Phi-4, Phi-4-reasoning / -plus, Wan 2.2 T2I, Wan 2.1 (1.3B / 14B), and SkyReels V2 use 13.8 GB of memory at Q6_K.
You can also run models at the high quality Q8_0 or FP8 quantization levels if the parameter count is lower. Vicuna 13B, HunyuanVideo, HunyuanVideo-Avatar, LTX-Video / LTX-2, and FramePack are 13B models that use 12.8 GB at Q6_K. FLUX.1 dev is a 12B model that uses 14.4 GB at the FP8 / optimized quantization level. Other 12B models like Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, and FLUX.1 Kontext dev use 15.3 GB of memory at the Q8_0 quantization level.
When a model is too large for the 16 GB graphics memory, you can use CPU offload if you have 32 GB of system RAM. This method splits the model between your graphics card and your system memory, which slows down processing speed. Solar Pro and Codestral 22B are 22B models that need 16.1 GB at Q4_K_M and require 18.1 GB of system RAM. Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 are 24B models that need 17.6 GB at Q4_K_M and require 19.6 GB of system RAM. Aria is a 25B model needing 18.3 GB at Q4_K_M and requiring 20.3 GB of system RAM. Gemma 4 26B-A4B and Gemma 4 (all sizes) are 26B models needing 19 GB at Q4_K_M and requiring 21 GB of system RAM. Gemma 3 27B and Gemma 3 4B/12B/27B (vision) are 27B models needing 19.8 GB at Q4_K_M and requiring 21.8 GB of system RAM.
All memory calculations for these models assume a standard 4k context window. If you increase the context window to process longer documents or larger chat histories, the memory usage will increase. This extra memory demand may force you to use a lower quantization level or switch to CPU offload to avoid running out of memory.