Best local AI models for AMD Vega Frontier Edition
16 GB HBM2. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| gpt-oss-20b | 21B | Q4_K_M | 15.4 GB |
| Reka Flash 3 | 21B | Q4_K_M | 15.4 GB |
| Qwen-Image | 20B | Q4_K_M | 14.6 GB |
| Qwen-Image-Edit | 20B | Q4_K_M | 14.6 GB |
| CogVLM2 | 19B | Q4_K_M | 13.9 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q5_K_M | 14.5 GB |
| Ling-Coder-Lite | 16.8B | Q5_K_M | 14.3 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q6_K | 15.7 GB |
| Kimi-VL A3B | 16B | Q6_K | 15.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q6_K | 14.8 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q6_K | 14.8 GB |
| Qwen2.5 14B | 14.7B | Q6_K | 15.3 GB |
| Phi-3 Medium | 14B | Q6_K | 13.8 GB |
| Phi-4 | 14B | Q6_K | 13.8 GB |
| Phi-4-reasoning / -plus | 14B | Q6_K | 13.8 GB |
| Wan 2.2 T2I | 14B | Q6_K | 13.8 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q6_K | 13.8 GB |
| SkyReels V2 | 14B | Q6_K | 13.8 GB |
| Vicuna 13B | 13B | Q6_K | 12.8 GB |
| HunyuanVideo | 13B | Q6_K | 12.8 GB |
| HunyuanVideo-Avatar | 13B | Q6_K | 12.8 GB |
| LTX-Video / LTX-2 | 13B | Q6_K | 12.8 GB |
| FramePack | 13B | Q6_K | 12.8 GB |
| FLUX.1 dev | 12B | FP8 / optimized | 14.4 GB |
| Gemma 3 12B | 12B | Q8_0 | 15.3 GB |
| Gemma 4 12B | 12B | Q8_0 | 15.3 GB |
| Mistral NeMo 12B | 12B | Q8_0 | 15.3 GB |
| Pixtral 12B | 12B | Q8_0 | 15.3 GB |
| FLUX.1 schnell | 12B | Q8_0 | 15.3 GB |
| FLUX.1 Kontext dev | 12B | Q8_0 | 15.3 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
| Mistral Small 3.2 | 24B | 17.6 GB needed | 19.6 GB |
| Magistral Small | 24B | 17.6 GB needed | 19.6 GB |
| Devstral Small 1.1 | 24B | 17.6 GB needed | 19.6 GB |
| Aria | 25B | 18.3 GB needed | 20.3 GB |
| Gemma 4 26B-A4B | 26B | 19 GB needed | 21 GB |
| Gemma 4 (all sizes) | 26B | 19 GB needed | 21 GB |
| Gemma 3 27B | 27B | 19.8 GB needed | 21.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | 19.8 GB needed | 21.8 GB |
How to read this
The AMD Vega Frontier Edition features 16 GB of HBM2 memory. This high speed onboard memory determines the maximum size of the artificial intelligence models you can run locally. To fit a model entirely within this VRAM limit, you must balance the parameter count of the model against its quantization level. Running models completely inside the 16 GB VRAM ensures the fastest possible processing speeds.
The quantization column indicates the compression level applied to each model. For example, the 21B models gpt-oss-20b and Reka Flash 3 fit within 15.4 GB of VRAM when using the Q4_K_M quantization. Vision models like Qwen-Image and Qwen-Image-Edit require 14.6 GB of VRAM at the same Q4_K_M level. As parameter sizes decrease, you can use higher quality quantizations. The 15B StarCoder2 and the 14.7B Qwen2.5 models run at the sharper Q6_K quantization, consuming 14.8 GB and 15.3 GB of VRAM respectively.
Popular 12B models can run at the high quality Q8_0 quantization level. Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, and FLUX.1 Kontext dev all fit within 15.3 GB of VRAM at Q8_0. If you want to run the FLUX.1 dev model, it requires 14.4 GB of VRAM using an FP8 or optimized quantization. These options allow you to maximize output accuracy without exceeding the physical memory limit of your graphics hardware.
When a model exceeds the 16 GB HBM2 capacity, you must offload some layers to your system RAM. This offloading process allows you to run larger models but decreases generation speed. For these setups, we assume a system with 32 GB of system RAM. Under this configuration, Solar Pro and Codestral 22B require 16.1 GB of memory at Q4_K_M, which utilizes 18.1 GB of system RAM. Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 require 17.6 GB at Q4_K_M, using 19.6 GB of system RAM.
Even larger models can be run using this system RAM offload method. Aria requires 18.3 GB at Q4_K_M, which uses 20.3 GB of system RAM. Gemma 4 26B-A4B and Gemma 4 all sizes require 19 GB at Q4_K_M, consuming 21 GB of system RAM. The Gemma 3 27B model and the Gemma 3 4B/12B/27B vision model require 19.8 GB at Q4_K_M, which utilizes 21.8 GB of system RAM.
You must also consider the memory cost of the context window. The VRAM usage figures listed here are calculated using a standard 4k context window. If you increase the context window to process longer documents or larger chat histories, the memory requirement will rise. This extra memory usage might force you to use a lower quantization level or offload more layers to system RAM to prevent out of memory errors.