Best local AI models for AMD Pro 560X

4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Lumina-Next / Lumina-Image 2.05BQ4_K_M3.7 GB
CogVideoX 2B / 5B5BQ4_K_M3.7 GB
DeepSeek-VL24.5BQ5_K_M3.8 GB
DeepFloyd IF4.3BQ5_K_M3.7 GB
Phi-3.5-vision4.2BQ5_K_M3.6 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-4-mini-instruct3.8BQ6_K3.7 GB
Phi-3.5 Mini3.8BQ6_K3.7 GB
OmniGen / OmniGen23.8BQ6_K3.7 GB
SD Cascade (Würstchen v3)3.6BQ6_K3.5 GB
SDXL Turbo3.5BQ6_K3.4 GB
SDXL Lightning3.5BQ6_K3.4 GB
ACE-Step3.5BQ6_K3.4 GB
MusicGen small/medium/large3.3BQ6_K3.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB
Open-Sora Plan2.7BQ8_03.4 GB
LFM2 1.2B / 2.6B2.6BQ8_03.3 GB
Playground v2.52.6BQ8_03.3 GB
Stable Diffusion 3.5 Medium2.5BQ8_03.2 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB

How to read this

The AMD Radeon Pro 560X is a mobile graphics card equipped with 4 GB of GDDR5 dedicated video memory. When running artificial intelligence models locally, this hardware boundary determines which models can execute entirely on your graphics processor. Keeping the model files within this memory limit ensures faster processing speeds because the system does not need to transfer data back and forth from your system memory.

To fit inside the 4 GB limit, models use quantization. Quantization is a compression method that reduces the precision of model weights to save space. The quant column shows the best balance of size and quality for each model. For example, the Q4_K_M quant represents a four bit compression level, while Q6_K and Q8_0 represent higher quality six bit and eight bit compressions that require more memory space.

Several highly capable models can run entirely within your graphics memory. The largest fitting options include Lumina-Next or Lumina-Image 2.0 and CogVideoX 2B or 5B, which both use 3.7 GB of memory at the Q4_K_M quant. DeepSeek-VL2 fits at Q5_K_M using 3.8 GB. For text and vision tasks, Phi-3.5-vision uses 3.6 GB at Q5_K_M, while Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, and Danube 3 4B all utilize 3.9 GB at the Q6_K quant.

Audio and image generation models also fit comfortably on this hardware. Fish Speech 1.5 or OpenAudio S1 uses 3.9 GB at Q6_K. SDXL Turbo and SDXL Lightning both require 3.4 GB at Q6_K. If you want higher precision, smaller models like SmolLM3 3B, Replit Code v1.5 3B, and Kandinsky 3.1 can run at the Q8_0 quant, using 3.8 GB of video memory.

When a model exceeds 4 GB, you must use CPU offload. This process splits the model layers between your graphics card and your system RAM, assuming you have 32 GB of system RAM. Offloading allows you to run larger models, but it costs performance because system RAM is much slower than video memory. For instance, Mistral 7B needs 5.7 GB at Q4_K_M, which requires 7.7 GB of system RAM. Similarly, Qwen2.5 7B, OLMo 2 7B, Falcon 3 7B, Command R7B, and OpenHermes 2.5 all need 5.1 GB at Q4_K_M, requiring 7.1 GB of system RAM.

You must also consider the context window caveat. The memory figures listed are calculated using a baseline 4k context window. If you increase the context length to process longer documents or conversations, the memory requirements will rise. This extra memory usage might push a fitting model over the 4 GB limit, which will force the system to offload layers to your system RAM and slow down generation speeds.