Best local AI models for AMD Pro SSG

16 GB HBM2. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
gpt-oss-20b21BQ4_K_M15.4 GB
Reka Flash 321BQ4_K_M15.4 GB
Qwen-Image20BQ4_K_M14.6 GB
Qwen-Image-Edit20BQ4_K_M14.6 GB
CogVLM219BQ4_K_M13.9 GB
HunyuanImage 2.1 / 3.017BQ5_K_M14.5 GB
Ling-Coder-Lite16.8BQ5_K_M14.3 GB
DeepSeek-Coder-V2 16B / 236B16BQ6_K15.7 GB
Kimi-VL A3B16BQ6_K15.7 GB
Apriel-1.5-15B-Thinker15BQ6_K14.8 GB
StarCoder2 3B / 7B / 15B15BQ6_K14.8 GB
Qwen2.5 14B14.7BQ6_K15.3 GB
Phi-3 Medium14BQ6_K13.8 GB
Phi-414BQ6_K13.8 GB
Phi-4-reasoning / -plus14BQ6_K13.8 GB
Wan 2.2 T2I14BQ6_K13.8 GB
Wan 2.1 (1.3B / 14B)14BQ6_K13.8 GB
SkyReels V214BQ6_K13.8 GB
Vicuna 13B13BQ6_K12.8 GB
HunyuanVideo13BQ6_K12.8 GB
HunyuanVideo-Avatar13BQ6_K12.8 GB
LTX-Video / LTX-213BQ6_K12.8 GB
FramePack13BQ6_K12.8 GB
FLUX.1 dev12BFP8 / optimized14.4 GB
Gemma 3 12B12BQ8_015.3 GB
Gemma 4 12B12BQ8_015.3 GB
Mistral NeMo 12B12BQ8_015.3 GB
Pixtral 12B12BQ8_015.3 GB
FLUX.1 schnell12BQ8_015.3 GB
FLUX.1 Kontext dev12BQ8_015.3 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at Q4_K_MSystem RAM at 4k
Solar Pro22B16.1 GB needed18.1 GB
Codestral 22B22B16.1 GB needed18.1 GB
Mistral Small 3.224B17.6 GB needed19.6 GB
Magistral Small24B17.6 GB needed19.6 GB
Devstral Small 1.124B17.6 GB needed19.6 GB
Aria25B18.3 GB needed20.3 GB
Gemma 4 26B-A4B26B19 GB needed21 GB
Gemma 4 (all sizes)26B19 GB needed21 GB
Gemma 3 27B27B19.8 GB needed21.8 GB
Gemma 3 4B/12B/27B (vision)27B19.8 GB needed21.8 GB

How to read this

The AMD Radeon Pro SSG features 16 GB of high bandwidth HBM2 memory. This dedicated memory determines the size of the artificial intelligence models you can run locally. To run a model entirely on the hardware, the model files and the active context data must fit within this 16 GB boundary.

The quantization column indicates the compression level used to fit these models into memory. For example, the 21B gpt-oss-20b and Reka Flash 3 models fit using a Q4_K_M quantization, which uses 15.4 GB. Vision models like Qwen-Image and Qwen-Image-Edit at 20B fit with Q4_K_M quantization using 14.6 GB, while CogVLM2 at 19B uses 13.9 GB. Models like HunyuanImage 2.1 / 3.0 at 17B and Ling-Coder-Lite at 16.8B utilize Q5_K_M quantization, consuming 14.5 GB and 14.3 GB respectively.

Smaller models can run at higher precision levels. The 16B DeepSeek-Coder-V2 16B / 236B and Kimi-VL A3B models run at Q6_K quantization using 15.7 GB. Apriel-1.5-15B-Thinker and StarCoder2 3B / 7B / 15B at 15B use 14.8 GB at Q6_K. Qwen2.5 14B uses 15.3 GB at Q6_K. Several 14B models, including Phi-3 Medium, Phi-4, Phi-4-reasoning / -plus, Wan 2.2 T2I, Wan 2.1 (1.3B / 14B), and SkyReels V2, use 13.8 GB at Q6_K. Vicuna 13B, HunyuanVideo, HunyuanVideo-Avatar, LTX-Video / LTX-2, and FramePack at 13B use 12.8 GB at Q6_K. FLUX.1 dev at 12B uses 14.4 GB at FP8 / optimized. Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, and FLUX.1 Kontext dev at 12B run at Q8_0 quantization using 15.3 GB.

When a model exceeds the 16 GB HBM2 limit, you must offload layers to the system RAM. This offloading process allows you to run larger models but reduces processing speed because system RAM is slower than HBM2. For these cases, we assume your workstation has 32 GB of system RAM.

Using CPU offload, you can run Solar Pro and Codestral 22B at Q4_K_M, which requires 16.1 GB of model space and 18.1 GB of system RAM. Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 at 24B require 17.6 GB at Q4_K_M and 19.6 GB of system RAM. Aria at 25B requires 18.3 GB at Q4_K_M and 20.3 GB of system RAM. Gemma 4 26B-A4B and Gemma 4 (all sizes) at 26B require 19 GB at Q4_K_M and 21 GB of system RAM. Gemma 3 27B and Gemma 3 4B/12B/27B (vision) at 27B require 19.8 GB at Q4_K_M and 21.8 GB of system RAM.

All listed memory figures are calculated using a standard 4k context window. If you increase the context window to process longer documents, the memory usage will grow. This extra memory demand may force you to use a lower quantization level or offload more layers to system RAM.