Best local AI models for AMD R9 M295X

4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Lumina-Next / Lumina-Image 2.05BQ4_K_M3.7 GB
CogVideoX 2B / 5B5BQ4_K_M3.7 GB
DeepSeek-VL24.5BQ5_K_M3.8 GB
DeepFloyd IF4.3BQ5_K_M3.7 GB
Phi-3.5-vision4.2BQ5_K_M3.6 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-4-mini-instruct3.8BQ6_K3.7 GB
Phi-3.5 Mini3.8BQ6_K3.7 GB
OmniGen / OmniGen23.8BQ6_K3.7 GB
SD Cascade (Würstchen v3)3.6BQ6_K3.5 GB
SDXL Turbo3.5BQ6_K3.4 GB
SDXL Lightning3.5BQ6_K3.4 GB
ACE-Step3.5BQ6_K3.4 GB
MusicGen small/medium/large3.3BQ6_K3.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB
Open-Sora Plan2.7BQ8_03.4 GB
LFM2 1.2B / 2.6B2.6BQ8_03.3 GB
Playground v2.52.6BQ8_03.3 GB
Stable Diffusion 3.5 Medium2.5BQ8_03.2 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB

How to read this

The AMD R9 M295X graphics card features 4 GB of GDDR5 memory. This dedicated memory size determines which local AI models can run entirely on the hardware. When a model fits inside this 4 GB limit, it processes data much faster because the graphics processor can access the model weights directly without waiting for the system memory.

To fit larger models into the 4 GB limit, we use quantized versions. The quant column shows the compression level used for each model. For example, Q4_K_M represents a medium four bit quantization, while Q6_K and Q8_0 represent six bit and eight bit quantization. Higher quantization numbers preserve more model accuracy but require more memory space.

Several models fit completely within the graphics memory. The Lumina-Next or Lumina-Image 2.0 model at 5B parameters uses 3.7 GB of memory with the Q4_K_M quant. The CogVideoX 5B model also uses 3.7 GB with the Q4_K_M quant. DeepSeek-VL2 at 4.5B parameters fits using 3.8 GB with the Q5_K_M quant. Phi-3.5-vision at 4.2B parameters uses 3.6 GB with the Q5_K_M quant.

For text and audio tasks, Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, Danube 3 4B, and Fish Speech 1.5 or OpenAudio S1 all use 3.9 GB of memory with the Q6_K quant. Phi-4-mini-instruct and Phi-3.5 Mini use 3.7 GB with the Q6_K quant. Under the Q8_0 quant, SmolLM3 3B and Kandinsky 3.1 use 3.8 GB of memory, while Stable Diffusion 3.5 Medium uses 3.2 GB.

When a model exceeds the 4 GB graphics memory, you must use CPU offload. This method splits the model between your graphics card and your system RAM. Assuming you have 32 GB of system RAM, you can run Mistral 7B with the Q4_K_M quant, which needs 5.7 GB of memory and 7.7 GB of system RAM. Similarly, Qwen2.5 7B, OLMo 2 7B, Falcon 3 7B, Command R7B, and OpenHermes 2.5 need 5.1 GB of memory and 7.1 GB of system RAM.

CPU offload comes with a performance cost. Moving data between the graphics card and system RAM slows down processing speeds significantly. Additionally, these memory calculations assume a standard 4k context window. If you increase the context window to process longer texts, the memory usage will rise and may exceed your available limits.