Best local AI models for AMD PRO W7700
16 GB GDDR6. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| gpt-oss-20b | 21B | Q4_K_M | 15.4 GB |
| Reka Flash 3 | 21B | Q4_K_M | 15.4 GB |
| Qwen-Image | 20B | Q4_K_M | 14.6 GB |
| Qwen-Image-Edit | 20B | Q4_K_M | 14.6 GB |
| CogVLM2 | 19B | Q4_K_M | 13.9 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q5_K_M | 14.5 GB |
| Ling-Coder-Lite | 16.8B | Q5_K_M | 14.3 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q6_K | 15.7 GB |
| Kimi-VL A3B | 16B | Q6_K | 15.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q6_K | 14.8 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q6_K | 14.8 GB |
| Qwen2.5 14B | 14.7B | Q6_K | 15.3 GB |
| Phi-3 Medium | 14B | Q6_K | 13.8 GB |
| Phi-4 | 14B | Q6_K | 13.8 GB |
| Phi-4-reasoning / -plus | 14B | Q6_K | 13.8 GB |
| Wan 2.2 T2I | 14B | Q6_K | 13.8 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q6_K | 13.8 GB |
| SkyReels V2 | 14B | Q6_K | 13.8 GB |
| Vicuna 13B | 13B | Q6_K | 12.8 GB |
| HunyuanVideo | 13B | Q6_K | 12.8 GB |
| HunyuanVideo-Avatar | 13B | Q6_K | 12.8 GB |
| LTX-Video / LTX-2 | 13B | Q6_K | 12.8 GB |
| FramePack | 13B | Q6_K | 12.8 GB |
| FLUX.1 dev | 12B | FP8 / optimized | 14.4 GB |
| Gemma 3 12B | 12B | Q8_0 | 15.3 GB |
| Gemma 4 12B | 12B | Q8_0 | 15.3 GB |
| Mistral NeMo 12B | 12B | Q8_0 | 15.3 GB |
| Pixtral 12B | 12B | Q8_0 | 15.3 GB |
| FLUX.1 schnell | 12B | Q8_0 | 15.3 GB |
| FLUX.1 Kontext dev | 12B | Q8_0 | 15.3 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
| Mistral Small 3.2 | 24B | 17.6 GB needed | 19.6 GB |
| Magistral Small | 24B | 17.6 GB needed | 19.6 GB |
| Devstral Small 1.1 | 24B | 17.6 GB needed | 19.6 GB |
| Aria | 25B | 18.3 GB needed | 20.3 GB |
| Gemma 4 26B-A4B | 26B | 19 GB needed | 21 GB |
| Gemma 4 (all sizes) | 26B | 19 GB needed | 21 GB |
| Gemma 3 27B | 27B | 19.8 GB needed | 21.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | 19.8 GB needed | 21.8 GB |
How to read this
The AMD Radeon PRO W7700 workstation graphics card features 16 GB of GDDR6 memory. This dedicated memory determines the maximum size of the artificial intelligence models you can run entirely on the hardware. When a model fits completely within this 16 GB boundary, the GPU processes tokens at maximum speed. If a model exceeds this limit, you must offload parts of the workload to your system memory.
To fit larger models into the onboard memory, developers use quantization. The quantization column shows the compression level applied to the model weights. For example, the 21B gpt-oss-20b and Reka Flash 3 models fit within 15.4 GB of memory when compressed to the Q4_K_M format. Similarly, the 20B Qwen-Image and Qwen-Image-Edit models require 14.6 GB of space using the same Q4_K_M quantization level.
As model sizes decrease, you can use higher quality quantization levels. The 17B HunyuanImage 2.1 / 3.0 fits within 14.5 GB using Q5_K_M compression. The 16B DeepSeek-Coder-V2 16B / 236B and Kimi-VL A3B models run at the Q6_K level using 15.7 GB of memory. Popular 14B models like Phi-4, Phi-4-reasoning / -plus, Wan 2.2 T2I, and SkyReels V2 also run at the Q6_K level using 13.8 GB of memory.
For smaller models, you can deploy the highest quality Q8_0 quantization. The 12B Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, and Pixtral 12B models all utilize 15.3 GB of memory at Q8_0. The FLUX.1 dev model can run at 12B using an optimized FP8 quantization that fits within 14.4 GB of the onboard graphics memory.
When a model is too large for the 16 GB graphics card, you can offload layers to your system RAM. This process requires at least 32 GB of system memory. For instance, running the 22B Solar Pro or Codestral 22B at Q4_K_M requires 16.1 GB of memory and 18.1 GB of system RAM. The 27B Gemma 3 27B requires 19.8 GB of memory and 21.8 GB of system RAM. Offloading allows you to run these larger models but reduces processing speeds.
All memory calculations assume a standard 4k context window. If you increase the context window to process longer documents, the memory requirements will rise. This extra demand might force you to use a lower quantization level or offload more data to the slower system RAM.