Best local AI models for AMD R9 M290X

4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Lumina-Next / Lumina-Image 2.05BQ4_K_M3.7 GB
CogVideoX 2B / 5B5BQ4_K_M3.7 GB
DeepSeek-VL24.5BQ5_K_M3.8 GB
DeepFloyd IF4.3BQ5_K_M3.7 GB
Phi-3.5-vision4.2BQ5_K_M3.6 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-4-mini-instruct3.8BQ6_K3.7 GB
Phi-3.5 Mini3.8BQ6_K3.7 GB
OmniGen / OmniGen23.8BQ6_K3.7 GB
SD Cascade (Würstchen v3)3.6BQ6_K3.5 GB
SDXL Turbo3.5BQ6_K3.4 GB
SDXL Lightning3.5BQ6_K3.4 GB
ACE-Step3.5BQ6_K3.4 GB
MusicGen small/medium/large3.3BQ6_K3.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB
Open-Sora Plan2.7BQ8_03.4 GB
LFM2 1.2B / 2.6B2.6BQ8_03.3 GB
Playground v2.52.6BQ8_03.3 GB
Stable Diffusion 3.5 Medium2.5BQ8_03.2 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB

How to read this

The AMD R9 M290X graphics card features 4 GB of GDDR5 memory. This physical memory size determines the maximum size of the local AI models you can run directly on the hardware. To run a model entirely on this GPU, the model files and the active memory must fit within this 4 GB limit. If a model exceeds this limit, the system will fail to run it or experience severe performance drops.

The quantization column shows the compression level used to fit these models into your hardware. Quantization reduces the precision of model weights to save space. For example, a Q4_K_M quant uses 4 bit quantization to fit larger models like Lumina-Next 5B or CogVideoX 5B into 3.7 GB of memory. A Q6_K quant offers higher precision for models like Qwen3 4B or Phi-4-mini-instruct. A Q8_0 quant provides the highest precision for smaller models like SmolLM3 3B or Kandinsky 3.1.

When a model is too large for the 4 GB GPU memory, you can use CPU offloading. This process splits the model layers between your GPU memory and your system RAM. We assume your system has 32 GB of system RAM for these scenarios. Offloading allows you to run larger models like Mistral 7B or Falcon 3 7B, but it comes with a cost. Moving data between the system RAM and the GPU slows down the generation speed significantly.

For fully local GPU execution, you can run models up to 5B parameters if they are highly compressed. DeepSeek-VL2 4.5B fits at Q5_K_M using 3.8 GB. Phi-3.5-vision 4.2B fits at Q5_K_M using 3.6 GB. Popular image generation models like SDXL Turbo and SDXL Lightning fit at Q6_K using 3.4 GB. Smaller 3B models like Allegro or Playground v2.5 fit at Q8_0 using 3.6 GB and 3.3 GB respectively.

CPU offloading expands your options to 7B models. Mistral 7B requires 5.7 GB at Q4_K_M and needs 7.7 GB of system RAM. Qwen2.5 7B, OLMo 2 7B, and Command R7B each require 5.1 GB at Q4_K_M and need 7.1 GB of system RAM. Even Stable Diffusion XL can run via offloading, requiring 4.1 GB at FP8 and 6.1 GB of system RAM.

You must consider the context window when running these models. The memory figures listed are calculated using a baseline 4k context window. If you increase the context window to process longer documents or chat histories, the memory usage will grow. Running close to the 4 GB limit on models like DeepSeek-VL2 or Qwen3 4B leaves very little room for context expansion.