Best local AI models for AMD RX Vega M GH

4 GB HBM2. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Lumina-Next / Lumina-Image 2.05BQ4_K_M3.7 GB
CogVideoX 2B / 5B5BQ4_K_M3.7 GB
DeepSeek-VL24.5BQ5_K_M3.8 GB
DeepFloyd IF4.3BQ5_K_M3.7 GB
Phi-3.5-vision4.2BQ5_K_M3.6 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-4-mini-instruct3.8BQ6_K3.7 GB
Phi-3.5 Mini3.8BQ6_K3.7 GB
OmniGen / OmniGen23.8BQ6_K3.7 GB
SD Cascade (Würstchen v3)3.6BQ6_K3.5 GB
SDXL Turbo3.5BQ6_K3.4 GB
SDXL Lightning3.5BQ6_K3.4 GB
ACE-Step3.5BQ6_K3.4 GB
MusicGen small/medium/large3.3BQ6_K3.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB
Open-Sora Plan2.7BQ8_03.4 GB
LFM2 1.2B / 2.6B2.6BQ8_03.3 GB
Playground v2.52.6BQ8_03.3 GB
Stable Diffusion 3.5 Medium2.5BQ8_03.2 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB

How to read this

The AMD RX Vega M GH graphics processor features 4 GB of high bandwidth HBM2 memory. This dedicated memory determines which artificial intelligence models can run entirely on your hardware. When a model fits inside this limit, it executes quickly because the processor accesses the parameters directly from the fast HBM2 memory pool.

The best quantization column indicates the optimal compression format for each model. Quantization reduces the precision of model weights to save space. For example, the Q6_K format allows a 4B model like Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, Danube 3 4B, or Fish Speech 1.5 / OpenAudio S1 to fit within 3.9 GB of memory. Larger models like Lumina-Next / Lumina-Image 2.0 5B and CogVideoX 2B / 5B 5B require a Q4_K_M quantization to fit inside 3.7 GB of memory.

Models with smaller parameter counts can run at higher precision levels. The 3B models such as SmolLM3 3B, Replit Code v1.5 3B, Kandinsky 3.1, Voxtral Mini / Small, Orpheus TTS, and Higgs Audio v2 use the Q8_0 quantization which occupies 3.8 GB of memory. Other options like DeepSeek-VL2 4.5B and DeepFloyd IF 4.3B utilize Q5_K_M quantization to stay under the limit at 3.8 GB and 3.7 GB respectively.

When a model exceeds the 4 GB limit of your graphics hardware, you must use CPU offload. This technique splits the model layers between your graphics card and your system RAM. We assume your computer has 32 GB of system RAM for these scenarios. Offloading allows you to run larger models, but it reduces processing speed because system RAM is slower than dedicated HBM2 memory.

Several popular models require this offload setup. Mistral 7B needs 5.7 GB of memory at Q4_K_M quantization and uses 7.7 GB of system RAM. The Qwen2.5 0.5B / 1.5B / 3B / 7B 7B model, OLMo 2 1B / 7B 7B model, Falcon 3 1B / 3B / 7B 7B model, Command R7B 7B model, and OpenHermes 2.5 7B model all require 5.1 GB at Q4_K_M quantization and use 7.1 GB of system RAM. Stable Diffusion XL requires 4.1 GB at FP8 or optimized settings and uses 6.1 GB of system RAM.

Be aware of the context window limits when running these models. The memory calculations on this page are based on a standard 4k context window. If you increase the context window to process longer documents or chat histories, the memory usage will rise. This extra memory demand can push a model past the 4 GB limit and trigger automatic CPU offloading.