Best local AI models for AMD RX 7900 GRE

16 GB GDDR6. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
gpt-oss-20b21BQ4_K_M15.4 GB
Reka Flash 321BQ4_K_M15.4 GB
Qwen-Image20BQ4_K_M14.6 GB
Qwen-Image-Edit20BQ4_K_M14.6 GB
CogVLM219BQ4_K_M13.9 GB
HunyuanImage 2.1 / 3.017BQ5_K_M14.5 GB
Ling-Coder-Lite16.8BQ5_K_M14.3 GB
DeepSeek-Coder-V2 16B / 236B16BQ6_K15.7 GB
Kimi-VL A3B16BQ6_K15.7 GB
Apriel-1.5-15B-Thinker15BQ6_K14.8 GB
StarCoder2 3B / 7B / 15B15BQ6_K14.8 GB
Qwen2.5 14B14.7BQ6_K15.3 GB
Phi-3 Medium14BQ6_K13.8 GB
Phi-414BQ6_K13.8 GB
Phi-4-reasoning / -plus14BQ6_K13.8 GB
Wan 2.2 T2I14BQ6_K13.8 GB
Wan 2.1 (1.3B / 14B)14BQ6_K13.8 GB
SkyReels V214BQ6_K13.8 GB
Vicuna 13B13BQ6_K12.8 GB
HunyuanVideo13BQ6_K12.8 GB
HunyuanVideo-Avatar13BQ6_K12.8 GB
LTX-Video / LTX-213BQ6_K12.8 GB
FramePack13BQ6_K12.8 GB
FLUX.1 dev12BFP8 / optimized14.4 GB
Gemma 3 12B12BQ8_015.3 GB
Gemma 4 12B12BQ8_015.3 GB
Mistral NeMo 12B12BQ8_015.3 GB
Pixtral 12B12BQ8_015.3 GB
FLUX.1 schnell12BQ8_015.3 GB
FLUX.1 Kontext dev12BQ8_015.3 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at Q4_K_MSystem RAM at 4k
Solar Pro22B16.1 GB needed18.1 GB
Codestral 22B22B16.1 GB needed18.1 GB
Mistral Small 3.224B17.6 GB needed19.6 GB
Magistral Small24B17.6 GB needed19.6 GB
Devstral Small 1.124B17.6 GB needed19.6 GB
Aria25B18.3 GB needed20.3 GB
Gemma 4 26B-A4B26B19 GB needed21 GB
Gemma 4 (all sizes)26B19 GB needed21 GB
Gemma 3 27B27B19.8 GB needed21.8 GB
Gemma 3 4B/12B/27B (vision)27B19.8 GB needed21.8 GB

How to read this

The AMD Radeon RX 7900 GRE features 16 GB of GDDR6 VRAM. This memory capacity determines the size of the artificial intelligence models you can run locally. When a model fits entirely within this 16 GB frame, it executes directly on the graphics hardware for maximum speed. If a model exceeds this limit, you must offload parts of it to your system memory.

The quantization column indicates the compression level applied to each model. Quantization reduces the precision of model weights to save space. For example, the 21B models gpt-oss-20b and Reka Flash 3 fit within 15.4 GB of VRAM using the Q4_K_M quantization. Smaller models like Mistral NeMo 12B, Pixtral 12B, and Gemma 4 12B can run at a higher Q8_0 precision while using 15.3 GB of VRAM.

Vision and image generation models also fit well on this hardware. The HunyuanImage 2.1 / 3.0 model uses 14.5 GB of VRAM at the Q5_K_M quantization. FLUX.1 dev runs at 14.4 GB using an FP8 / optimized quantization. Video generation models like HunyuanVideo and LTX-Video / LTX-2 use 12.8 GB of VRAM at the Q6_K quantization level.

Running larger models requires CPU offloading to your system RAM. This process allows you to run models like Mistral Small 3.2 or Gemma 3 27B, but it reduces processing speed significantly. For instance, Mistral Small 3.2 requires 17.6 GB at Q4_K_M quantization and needs 19.6 GB of system RAM. Gemma 3 27B requires 19.8 GB at Q4_K_M quantization and needs 21.8 GB of system RAM.

Memory calculations assume a standard 4k context window. As your conversation history grows, the context window consumes additional VRAM. If you use long prompts or extended chat sessions, the model might exceed the 16 GB limit of the RX 7900 GRE. You should choose a slightly smaller model size if you plan to use large context windows.