Best local AI models for AMD Pro 560X
4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Lumina-Next / Lumina-Image 2.0 | 5B | Q4_K_M | 3.7 GB |
| CogVideoX 2B / 5B | 5B | Q4_K_M | 3.7 GB |
| DeepSeek-VL2 | 4.5B | Q5_K_M | 3.8 GB |
| DeepFloyd IF | 4.3B | Q5_K_M | 3.7 GB |
| Phi-3.5-vision | 4.2B | Q5_K_M | 3.6 GB |
| Qwen3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 4 E4B | 4B | Q6_K | 3.9 GB |
| MiniCPM 3 4B | 4B | Q6_K | 3.9 GB |
| Danube 3 4B | 4B | Q6_K | 3.9 GB |
| Fish Speech 1.5 / OpenAudio S1 | 4B | Q6_K | 3.9 GB |
| Phi-4-mini-instruct | 3.8B | Q6_K | 3.7 GB |
| Phi-3.5 Mini | 3.8B | Q6_K | 3.7 GB |
| OmniGen / OmniGen2 | 3.8B | Q6_K | 3.7 GB |
| SD Cascade (Würstchen v3) | 3.6B | Q6_K | 3.5 GB |
| SDXL Turbo | 3.5B | Q6_K | 3.4 GB |
| SDXL Lightning | 3.5B | Q6_K | 3.4 GB |
| ACE-Step | 3.5B | Q6_K | 3.4 GB |
| MusicGen small/medium/large | 3.3B | Q6_K | 3.2 GB |
| SmolLM3 3B | 3B | Q8_0 | 3.8 GB |
| Replit Code v1.5 3B | 3B | Q8_0 | 3.8 GB |
| Kandinsky 3.1 | 3B | Q8_0 | 3.8 GB |
| Voxtral Mini / Small | 3B | Q8_0 | 3.8 GB |
| Orpheus TTS | 3B | Q8_0 | 3.8 GB |
| Higgs Audio v2 | 3B | Q8_0 | 3.8 GB |
| Allegro | 2.8B | Q8_0 | 3.6 GB |
| Open-Sora Plan | 2.7B | Q8_0 | 3.4 GB |
| LFM2 1.2B / 2.6B | 2.6B | Q8_0 | 3.3 GB |
| Playground v2.5 | 2.6B | Q8_0 | 3.3 GB |
| Stable Diffusion 3.5 Medium | 2.5B | Q8_0 | 3.2 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| Stable Diffusion XL | 3.417B | 4.1 GB needed | 6.1 GB |
| Phi-3 Mini | 3.8B | 4.4 GB needed | 6.4 GB |
| Phi-4-multimodal | 5.6B | 4.1 GB needed | 6.1 GB |
| Magicoder-S-DS 6.7B | 6.7B | 4.9 GB needed | 6.9 GB |
| Mistral 7B | 7B | 5.7 GB needed | 7.7 GB |
| Qwen2.5 0.5B / 1.5B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| OLMo 2 1B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Falcon 3 1B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Command R7B | 7B | 5.1 GB needed | 7.1 GB |
| OpenHermes 2.5 | 7B | 5.1 GB needed | 7.1 GB |
How to read this
The AMD Radeon Pro 560X is a mobile graphics card equipped with 4 GB of GDDR5 dedicated video memory. When running artificial intelligence models locally, this hardware boundary determines which models can execute entirely on your graphics processor. Keeping the model files within this memory limit ensures faster processing speeds because the system does not need to transfer data back and forth from your system memory.
To fit inside the 4 GB limit, models use quantization. Quantization is a compression method that reduces the precision of model weights to save space. The quant column shows the best balance of size and quality for each model. For example, the Q4_K_M quant represents a four bit compression level, while Q6_K and Q8_0 represent higher quality six bit and eight bit compressions that require more memory space.
Several highly capable models can run entirely within your graphics memory. The largest fitting options include Lumina-Next or Lumina-Image 2.0 and CogVideoX 2B or 5B, which both use 3.7 GB of memory at the Q4_K_M quant. DeepSeek-VL2 fits at Q5_K_M using 3.8 GB. For text and vision tasks, Phi-3.5-vision uses 3.6 GB at Q5_K_M, while Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, and Danube 3 4B all utilize 3.9 GB at the Q6_K quant.
Audio and image generation models also fit comfortably on this hardware. Fish Speech 1.5 or OpenAudio S1 uses 3.9 GB at Q6_K. SDXL Turbo and SDXL Lightning both require 3.4 GB at Q6_K. If you want higher precision, smaller models like SmolLM3 3B, Replit Code v1.5 3B, and Kandinsky 3.1 can run at the Q8_0 quant, using 3.8 GB of video memory.
When a model exceeds 4 GB, you must use CPU offload. This process splits the model layers between your graphics card and your system RAM, assuming you have 32 GB of system RAM. Offloading allows you to run larger models, but it costs performance because system RAM is much slower than video memory. For instance, Mistral 7B needs 5.7 GB at Q4_K_M, which requires 7.7 GB of system RAM. Similarly, Qwen2.5 7B, OLMo 2 7B, Falcon 3 7B, Command R7B, and OpenHermes 2.5 all need 5.1 GB at Q4_K_M, requiring 7.1 GB of system RAM.
You must also consider the context window caveat. The memory figures listed are calculated using a baseline 4k context window. If you increase the context length to process longer documents or conversations, the memory requirements will rise. This extra memory usage might push a fitting model over the 4 GB limit, which will force the system to offload layers to your system RAM and slow down generation speeds.