Best local AI models for AMD RX 560

4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Lumina-Next / Lumina-Image 2.05BQ4_K_M3.7 GB
CogVideoX 2B / 5B5BQ4_K_M3.7 GB
DeepSeek-VL24.5BQ5_K_M3.8 GB
DeepFloyd IF4.3BQ5_K_M3.7 GB
Phi-3.5-vision4.2BQ5_K_M3.6 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-4-mini-instruct3.8BQ6_K3.7 GB
Phi-3.5 Mini3.8BQ6_K3.7 GB
OmniGen / OmniGen23.8BQ6_K3.7 GB
SD Cascade (Würstchen v3)3.6BQ6_K3.5 GB
SDXL Turbo3.5BQ6_K3.4 GB
SDXL Lightning3.5BQ6_K3.4 GB
ACE-Step3.5BQ6_K3.4 GB
MusicGen small/medium/large3.3BQ6_K3.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB
Open-Sora Plan2.7BQ8_03.4 GB
LFM2 1.2B / 2.6B2.6BQ8_03.3 GB
Playground v2.52.6BQ8_03.3 GB
Stable Diffusion 3.5 Medium2.5BQ8_03.2 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB

How to read this

The AMD Radeon RX 560 is an entry level graphics card equipped with 4 GB of GDDR5 memory. When running artificial intelligence models locally, this hardware memory limit is the most critical factor. The model weights must reside in this memory for the graphics processor to compute them quickly. If a model exceeds this capacity, it cannot run entirely on the hardware processor.

To fit larger models into the 4 GB limit, developers use quantization. The quant column shows the compression level applied to each model. For example, the Q4_K_M quant uses four bit compression to fit the Lumina-Next 5B model at 3.7 GB of memory. Higher quants like Q6_K and Q8_0 offer better precision but require more space. A Q8_0 quant allows the 3B parameter SmolLM3 to fit at 3.8 GB of memory.

You can run models that exceed the onboard memory by offloading layers to your system RAM. This setup assumes you have 32 GB of system RAM available. For instance, running Mistral 7B at Q4_K_M requires 5.7 GB of total memory, which uses 7.7 GB of system RAM. Offloading allows you to run larger models like Command R7B or Falcon 3 7B, but it costs processing speed because system RAM is much slower than GDDR5.

When selecting a model, you must also consider the context window. Running a model at its maximum context length of 4k tokens or higher increases memory consumption during generation. If you use the maximum context, the active memory may exceed the limits of your card. For the tightest fits like DeepSeek-VL2 at 3.8 GB, you should keep your context short to avoid running out of memory.

The RX 560 can run diverse tasks within its 4 GB boundary. You can run image generation with SDXL Turbo at Q6_K using 3.4 GB of memory. For vision tasks, Phi-3.5-vision fits at Q5_K_M using 3.6 GB of memory. Audio generation is also possible with Orpheus TTS at Q8_0 using 3.8 GB of memory. Selecting the correct quant ensures these models run entirely on your graphics hardware.