Best local AI models for AMD RX 560X
4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Lumina-Next / Lumina-Image 2.0 | 5B | Q4_K_M | 3.7 GB |
| CogVideoX 2B / 5B | 5B | Q4_K_M | 3.7 GB |
| DeepSeek-VL2 | 4.5B | Q5_K_M | 3.8 GB |
| DeepFloyd IF | 4.3B | Q5_K_M | 3.7 GB |
| Phi-3.5-vision | 4.2B | Q5_K_M | 3.6 GB |
| Qwen3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 4 E4B | 4B | Q6_K | 3.9 GB |
| MiniCPM 3 4B | 4B | Q6_K | 3.9 GB |
| Danube 3 4B | 4B | Q6_K | 3.9 GB |
| Fish Speech 1.5 / OpenAudio S1 | 4B | Q6_K | 3.9 GB |
| Phi-4-mini-instruct | 3.8B | Q6_K | 3.7 GB |
| Phi-3.5 Mini | 3.8B | Q6_K | 3.7 GB |
| OmniGen / OmniGen2 | 3.8B | Q6_K | 3.7 GB |
| SD Cascade (Würstchen v3) | 3.6B | Q6_K | 3.5 GB |
| SDXL Turbo | 3.5B | Q6_K | 3.4 GB |
| SDXL Lightning | 3.5B | Q6_K | 3.4 GB |
| ACE-Step | 3.5B | Q6_K | 3.4 GB |
| MusicGen small/medium/large | 3.3B | Q6_K | 3.2 GB |
| SmolLM3 3B | 3B | Q8_0 | 3.8 GB |
| Replit Code v1.5 3B | 3B | Q8_0 | 3.8 GB |
| Kandinsky 3.1 | 3B | Q8_0 | 3.8 GB |
| Voxtral Mini / Small | 3B | Q8_0 | 3.8 GB |
| Orpheus TTS | 3B | Q8_0 | 3.8 GB |
| Higgs Audio v2 | 3B | Q8_0 | 3.8 GB |
| Allegro | 2.8B | Q8_0 | 3.6 GB |
| Open-Sora Plan | 2.7B | Q8_0 | 3.4 GB |
| LFM2 1.2B / 2.6B | 2.6B | Q8_0 | 3.3 GB |
| Playground v2.5 | 2.6B | Q8_0 | 3.3 GB |
| Stable Diffusion 3.5 Medium | 2.5B | Q8_0 | 3.2 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| Stable Diffusion XL | 3.417B | 4.1 GB needed | 6.1 GB |
| Phi-3 Mini | 3.8B | 4.4 GB needed | 6.4 GB |
| Phi-4-multimodal | 5.6B | 4.1 GB needed | 6.1 GB |
| Magicoder-S-DS 6.7B | 6.7B | 4.9 GB needed | 6.9 GB |
| Mistral 7B | 7B | 5.7 GB needed | 7.7 GB |
| Qwen2.5 0.5B / 1.5B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| OLMo 2 1B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Falcon 3 1B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Command R7B | 7B | 5.1 GB needed | 7.1 GB |
| OpenHermes 2.5 | 7B | 5.1 GB needed | 7.1 GB |
How to read this
The AMD RX 560X graphics card features 4 GB of GDDR5 video memory. This memory limit determines which artificial intelligence models can run directly on your hardware. To fit inside this 4 GB space, models must undergo quantization. Quantization reduces the precision of model weights to save space. The quant column shows the best quality level that fits your hardware. A higher quant like Q8_0 offers better accuracy than Q6_K or Q4_K_M but uses more memory per parameter.
For maximum performance, you can run models entirely within the onboard memory. The largest fitting models include Lumina-Next or Lumina-Image 2.0 and CogVideoX 2B or 5B. Both are 5B parameter models that use 3.7 GB of memory at the Q4_K_M quant. DeepSeek-VL2 is a 4.5B model using 3.8 GB at Q5_K_M. DeepFloyd IF uses 3.7 GB at Q5_K_M. Phi-3.5-vision is a 4.2B model using 3.6 GB at Q5_K_M.
Several 4B models run at the high quality Q6_K quant. Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, Danube 3 4B, and Fish Speech 1.5 or OpenAudio S1 all use 3.9 GB of memory. Phi-4-mini-instruct, Phi-3.5 Mini, and OmniGen or OmniGen2 are 3.8B models using 3.7 GB at Q6_K. Image generation models also fit well. SD Cascade (Würstchen v3) uses 3.5 GB, while SDXL Turbo, SDXL Lightning, and ACE-Step use 3.4 GB at Q6_K.
Audio and smaller language models can run at the excellent Q8_0 quant. MusicGen small/medium/large uses 3.2 GB at Q6_K. SmolLM3 3B, Replit Code v1.5 3B, Kandinsky 3.1, Voxtral Mini or Small, Orpheus TTS, and Higgs Audio v2 are 3B models using 3.8 GB at Q8_0. Allegro uses 3.6 GB, Open-Sora Plan uses 3.4 GB, and LFM2 1.2B or 2.6B and Playground v2.5 use 3.3 GB at Q8_0. Stable Diffusion 3.5 Medium uses 3.2 GB at Q8_0.
When a model is too large for the 4 GB video memory, you can use CPU offload. This process splits the model between your graphics card and your system RAM. Offloading allows you to run larger models but reduces processing speed. For these cases, we assume your computer has 32 GB of system RAM. Stable Diffusion XL needs 4.1 GB at FP8 or optimized settings and requires 6.1 GB of system RAM. Phi-3 Mini needs 4.4 GB at Q4_K_M and 6.4 GB of system RAM.
Other offload options include Phi-4-multimodal which needs 4.1 GB at Q4_K_M and 6.1 GB of system RAM. Magicoder-S-DS 6.7B needs 4.9 GB at Q4_K_M and 6.9 GB of system RAM. Mistral 7B needs 5.7 GB at Q4_K_M and 7.7 GB of system RAM. Qwen2.5 0.5B or 1.5B or 3B or 7B, OLMo 2 1B or 7B, Falcon 3 1B or 3B or 7B, Command R7B, and OpenHermes 2.5 all need 5.1 GB at Q4_K_M and 7.1 GB of system RAM.
Be aware of the 4k context limit when running local models. Generating longer text or processing large prompts increases memory usage. If you exceed the 4k context window, the model will require more memory than listed. This extra demand can cause the model to spill over into system RAM or fail to run entirely.