Best local AI models for AMD RX Vega M GH
4 GB HBM2. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Lumina-Next / Lumina-Image 2.0 | 5B | Q4_K_M | 3.7 GB |
| CogVideoX 2B / 5B | 5B | Q4_K_M | 3.7 GB |
| DeepSeek-VL2 | 4.5B | Q5_K_M | 3.8 GB |
| DeepFloyd IF | 4.3B | Q5_K_M | 3.7 GB |
| Phi-3.5-vision | 4.2B | Q5_K_M | 3.6 GB |
| Qwen3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 4 E4B | 4B | Q6_K | 3.9 GB |
| MiniCPM 3 4B | 4B | Q6_K | 3.9 GB |
| Danube 3 4B | 4B | Q6_K | 3.9 GB |
| Fish Speech 1.5 / OpenAudio S1 | 4B | Q6_K | 3.9 GB |
| Phi-4-mini-instruct | 3.8B | Q6_K | 3.7 GB |
| Phi-3.5 Mini | 3.8B | Q6_K | 3.7 GB |
| OmniGen / OmniGen2 | 3.8B | Q6_K | 3.7 GB |
| SD Cascade (Würstchen v3) | 3.6B | Q6_K | 3.5 GB |
| SDXL Turbo | 3.5B | Q6_K | 3.4 GB |
| SDXL Lightning | 3.5B | Q6_K | 3.4 GB |
| ACE-Step | 3.5B | Q6_K | 3.4 GB |
| MusicGen small/medium/large | 3.3B | Q6_K | 3.2 GB |
| SmolLM3 3B | 3B | Q8_0 | 3.8 GB |
| Replit Code v1.5 3B | 3B | Q8_0 | 3.8 GB |
| Kandinsky 3.1 | 3B | Q8_0 | 3.8 GB |
| Voxtral Mini / Small | 3B | Q8_0 | 3.8 GB |
| Orpheus TTS | 3B | Q8_0 | 3.8 GB |
| Higgs Audio v2 | 3B | Q8_0 | 3.8 GB |
| Allegro | 2.8B | Q8_0 | 3.6 GB |
| Open-Sora Plan | 2.7B | Q8_0 | 3.4 GB |
| LFM2 1.2B / 2.6B | 2.6B | Q8_0 | 3.3 GB |
| Playground v2.5 | 2.6B | Q8_0 | 3.3 GB |
| Stable Diffusion 3.5 Medium | 2.5B | Q8_0 | 3.2 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| Stable Diffusion XL | 3.417B | 4.1 GB needed | 6.1 GB |
| Phi-3 Mini | 3.8B | 4.4 GB needed | 6.4 GB |
| Phi-4-multimodal | 5.6B | 4.1 GB needed | 6.1 GB |
| Magicoder-S-DS 6.7B | 6.7B | 4.9 GB needed | 6.9 GB |
| Mistral 7B | 7B | 5.7 GB needed | 7.7 GB |
| Qwen2.5 0.5B / 1.5B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| OLMo 2 1B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Falcon 3 1B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Command R7B | 7B | 5.1 GB needed | 7.1 GB |
| OpenHermes 2.5 | 7B | 5.1 GB needed | 7.1 GB |
How to read this
The AMD RX Vega M GH graphics processor features 4 GB of high bandwidth HBM2 memory. This dedicated memory determines which artificial intelligence models can run entirely on your hardware. When a model fits inside this limit, it executes quickly because the processor accesses the parameters directly from the fast HBM2 memory pool.
The best quantization column indicates the optimal compression format for each model. Quantization reduces the precision of model weights to save space. For example, the Q6_K format allows a 4B model like Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, Danube 3 4B, or Fish Speech 1.5 / OpenAudio S1 to fit within 3.9 GB of memory. Larger models like Lumina-Next / Lumina-Image 2.0 5B and CogVideoX 2B / 5B 5B require a Q4_K_M quantization to fit inside 3.7 GB of memory.
Models with smaller parameter counts can run at higher precision levels. The 3B models such as SmolLM3 3B, Replit Code v1.5 3B, Kandinsky 3.1, Voxtral Mini / Small, Orpheus TTS, and Higgs Audio v2 use the Q8_0 quantization which occupies 3.8 GB of memory. Other options like DeepSeek-VL2 4.5B and DeepFloyd IF 4.3B utilize Q5_K_M quantization to stay under the limit at 3.8 GB and 3.7 GB respectively.
When a model exceeds the 4 GB limit of your graphics hardware, you must use CPU offload. This technique splits the model layers between your graphics card and your system RAM. We assume your computer has 32 GB of system RAM for these scenarios. Offloading allows you to run larger models, but it reduces processing speed because system RAM is slower than dedicated HBM2 memory.
Several popular models require this offload setup. Mistral 7B needs 5.7 GB of memory at Q4_K_M quantization and uses 7.7 GB of system RAM. The Qwen2.5 0.5B / 1.5B / 3B / 7B 7B model, OLMo 2 1B / 7B 7B model, Falcon 3 1B / 3B / 7B 7B model, Command R7B 7B model, and OpenHermes 2.5 7B model all require 5.1 GB at Q4_K_M quantization and use 7.1 GB of system RAM. Stable Diffusion XL requires 4.1 GB at FP8 or optimized settings and uses 6.1 GB of system RAM.
Be aware of the context window limits when running these models. The memory calculations on this page are based on a standard 4k context window. If you increase the context window to process longer documents or chat histories, the memory usage will rise. This extra memory demand can push a model past the 4 GB limit and trigger automatic CPU offloading.