Best local AI models for AMD R9 M280X
4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Lumina-Next / Lumina-Image 2.0 | 5B | Q4_K_M | 3.7 GB |
| CogVideoX 2B / 5B | 5B | Q4_K_M | 3.7 GB |
| DeepSeek-VL2 | 4.5B | Q5_K_M | 3.8 GB |
| DeepFloyd IF | 4.3B | Q5_K_M | 3.7 GB |
| Phi-3.5-vision | 4.2B | Q5_K_M | 3.6 GB |
| Qwen3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 4 E4B | 4B | Q6_K | 3.9 GB |
| MiniCPM 3 4B | 4B | Q6_K | 3.9 GB |
| Danube 3 4B | 4B | Q6_K | 3.9 GB |
| Fish Speech 1.5 / OpenAudio S1 | 4B | Q6_K | 3.9 GB |
| Phi-4-mini-instruct | 3.8B | Q6_K | 3.7 GB |
| Phi-3.5 Mini | 3.8B | Q6_K | 3.7 GB |
| OmniGen / OmniGen2 | 3.8B | Q6_K | 3.7 GB |
| SD Cascade (Würstchen v3) | 3.6B | Q6_K | 3.5 GB |
| SDXL Turbo | 3.5B | Q6_K | 3.4 GB |
| SDXL Lightning | 3.5B | Q6_K | 3.4 GB |
| ACE-Step | 3.5B | Q6_K | 3.4 GB |
| MusicGen small/medium/large | 3.3B | Q6_K | 3.2 GB |
| SmolLM3 3B | 3B | Q8_0 | 3.8 GB |
| Replit Code v1.5 3B | 3B | Q8_0 | 3.8 GB |
| Kandinsky 3.1 | 3B | Q8_0 | 3.8 GB |
| Voxtral Mini / Small | 3B | Q8_0 | 3.8 GB |
| Orpheus TTS | 3B | Q8_0 | 3.8 GB |
| Higgs Audio v2 | 3B | Q8_0 | 3.8 GB |
| Allegro | 2.8B | Q8_0 | 3.6 GB |
| Open-Sora Plan | 2.7B | Q8_0 | 3.4 GB |
| LFM2 1.2B / 2.6B | 2.6B | Q8_0 | 3.3 GB |
| Playground v2.5 | 2.6B | Q8_0 | 3.3 GB |
| Stable Diffusion 3.5 Medium | 2.5B | Q8_0 | 3.2 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| Stable Diffusion XL | 3.417B | 4.1 GB needed | 6.1 GB |
| Phi-3 Mini | 3.8B | 4.4 GB needed | 6.4 GB |
| Phi-4-multimodal | 5.6B | 4.1 GB needed | 6.1 GB |
| Magicoder-S-DS 6.7B | 6.7B | 4.9 GB needed | 6.9 GB |
| Mistral 7B | 7B | 5.7 GB needed | 7.7 GB |
| Qwen2.5 0.5B / 1.5B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| OLMo 2 1B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Falcon 3 1B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Command R7B | 7B | 5.1 GB needed | 7.1 GB |
| OpenHermes 2.5 | 7B | 5.1 GB needed | 7.1 GB |
How to read this
The AMD Radeon R9 M280X is equipped with 4 GB of GDDR5 graphics memory. This physical memory limit determines which artificial intelligence models can run entirely on the graphics hardware. When a model fits completely within this 4 GB frame buffer, it benefits from the direct memory bandwidth of the graphics card. This direct placement ensures the fastest possible processing speeds for local inference tasks.
To fit larger models into this hardware, developers use quantization. The quantization column shows the compression level applied to the model weights. For example, a Q4_K_M quant represents a four bit compression level, while Q6_K and Q8_0 represent six bit and eight bit compressions. Higher quantization levels like Q8_0 preserve more original model accuracy but require more memory space. Lower quantization levels like Q4_K_M reduce the memory footprint so larger parameter models can fit.
Under the strict 4 GB limit, several models can run entirely on the graphics card. The Lumina-Next or Lumina-Image 2.0 model at 5B parameters fits using a Q4_K_M quant which consumes 3.7 GB of memory. Similarly, CogVideoX 2B or 5B fits at 5B parameters using a Q4_K_M quant for 3.7 GB of usage. For vision tasks, DeepSeek-VL2 at 4.5B parameters uses 3.8 GB with a Q5_K_M quant, and Phi-3.5-vision at 4.2B parameters uses 3.6 GB with a Q5_K_M quant. Text models like Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, and Danube 3 4B all fit at 4B parameters using a Q6_K quant which consumes 3.9 GB.
When a model exceeds the 4 GB graphics memory limit, you must use CPU offload. This technique splits the model layers between your graphics card and your system RAM. CPU offload allows you to run larger models like Mistral 7B, which needs 5.7 GB at Q4_K_M and requires 7.7 GB of system RAM. Other 7B models like Qwen2.5, OLMo 2, Falcon 3, Command R7B, and OpenHermes 2.5 require 5.1 GB at Q4_K_M and 7.1 GB of system RAM. While CPU offload enables these larger models to run, the transfer of data between system RAM and graphics memory reduces processing speeds significantly.
Users must also consider the context window size when loading these models. The memory usage figures listed for these models assume a standard base context window. If you increase the context length to 4k tokens or higher, the memory required for the active conversation history will grow. This extra memory usage can push a model that normally fits at 3.9 GB over the physical 4 GB limit of the graphics card, which will trigger slow system RAM fallback.