Best local AI models for AMD R9 M290X
4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Lumina-Next / Lumina-Image 2.0 | 5B | Q4_K_M | 3.7 GB |
| CogVideoX 2B / 5B | 5B | Q4_K_M | 3.7 GB |
| DeepSeek-VL2 | 4.5B | Q5_K_M | 3.8 GB |
| DeepFloyd IF | 4.3B | Q5_K_M | 3.7 GB |
| Phi-3.5-vision | 4.2B | Q5_K_M | 3.6 GB |
| Qwen3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 4 E4B | 4B | Q6_K | 3.9 GB |
| MiniCPM 3 4B | 4B | Q6_K | 3.9 GB |
| Danube 3 4B | 4B | Q6_K | 3.9 GB |
| Fish Speech 1.5 / OpenAudio S1 | 4B | Q6_K | 3.9 GB |
| Phi-4-mini-instruct | 3.8B | Q6_K | 3.7 GB |
| Phi-3.5 Mini | 3.8B | Q6_K | 3.7 GB |
| OmniGen / OmniGen2 | 3.8B | Q6_K | 3.7 GB |
| SD Cascade (Würstchen v3) | 3.6B | Q6_K | 3.5 GB |
| SDXL Turbo | 3.5B | Q6_K | 3.4 GB |
| SDXL Lightning | 3.5B | Q6_K | 3.4 GB |
| ACE-Step | 3.5B | Q6_K | 3.4 GB |
| MusicGen small/medium/large | 3.3B | Q6_K | 3.2 GB |
| SmolLM3 3B | 3B | Q8_0 | 3.8 GB |
| Replit Code v1.5 3B | 3B | Q8_0 | 3.8 GB |
| Kandinsky 3.1 | 3B | Q8_0 | 3.8 GB |
| Voxtral Mini / Small | 3B | Q8_0 | 3.8 GB |
| Orpheus TTS | 3B | Q8_0 | 3.8 GB |
| Higgs Audio v2 | 3B | Q8_0 | 3.8 GB |
| Allegro | 2.8B | Q8_0 | 3.6 GB |
| Open-Sora Plan | 2.7B | Q8_0 | 3.4 GB |
| LFM2 1.2B / 2.6B | 2.6B | Q8_0 | 3.3 GB |
| Playground v2.5 | 2.6B | Q8_0 | 3.3 GB |
| Stable Diffusion 3.5 Medium | 2.5B | Q8_0 | 3.2 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| Stable Diffusion XL | 3.417B | 4.1 GB needed | 6.1 GB |
| Phi-3 Mini | 3.8B | 4.4 GB needed | 6.4 GB |
| Phi-4-multimodal | 5.6B | 4.1 GB needed | 6.1 GB |
| Magicoder-S-DS 6.7B | 6.7B | 4.9 GB needed | 6.9 GB |
| Mistral 7B | 7B | 5.7 GB needed | 7.7 GB |
| Qwen2.5 0.5B / 1.5B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| OLMo 2 1B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Falcon 3 1B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Command R7B | 7B | 5.1 GB needed | 7.1 GB |
| OpenHermes 2.5 | 7B | 5.1 GB needed | 7.1 GB |
How to read this
The AMD R9 M290X graphics card features 4 GB of GDDR5 memory. This physical memory size determines the maximum size of the local AI models you can run directly on the hardware. To run a model entirely on this GPU, the model files and the active memory must fit within this 4 GB limit. If a model exceeds this limit, the system will fail to run it or experience severe performance drops.
The quantization column shows the compression level used to fit these models into your hardware. Quantization reduces the precision of model weights to save space. For example, a Q4_K_M quant uses 4 bit quantization to fit larger models like Lumina-Next 5B or CogVideoX 5B into 3.7 GB of memory. A Q6_K quant offers higher precision for models like Qwen3 4B or Phi-4-mini-instruct. A Q8_0 quant provides the highest precision for smaller models like SmolLM3 3B or Kandinsky 3.1.
When a model is too large for the 4 GB GPU memory, you can use CPU offloading. This process splits the model layers between your GPU memory and your system RAM. We assume your system has 32 GB of system RAM for these scenarios. Offloading allows you to run larger models like Mistral 7B or Falcon 3 7B, but it comes with a cost. Moving data between the system RAM and the GPU slows down the generation speed significantly.
For fully local GPU execution, you can run models up to 5B parameters if they are highly compressed. DeepSeek-VL2 4.5B fits at Q5_K_M using 3.8 GB. Phi-3.5-vision 4.2B fits at Q5_K_M using 3.6 GB. Popular image generation models like SDXL Turbo and SDXL Lightning fit at Q6_K using 3.4 GB. Smaller 3B models like Allegro or Playground v2.5 fit at Q8_0 using 3.6 GB and 3.3 GB respectively.
CPU offloading expands your options to 7B models. Mistral 7B requires 5.7 GB at Q4_K_M and needs 7.7 GB of system RAM. Qwen2.5 7B, OLMo 2 7B, and Command R7B each require 5.1 GB at Q4_K_M and need 7.1 GB of system RAM. Even Stable Diffusion XL can run via offloading, requiring 4.1 GB at FP8 and 6.1 GB of system RAM.
You must consider the context window when running these models. The memory figures listed are calculated using a baseline 4k context window. If you increase the context window to process longer documents or chat histories, the memory usage will grow. Running close to the 4 GB limit on models like DeepSeek-VL2 or Qwen3 4B leaves very little room for context expansion.