Best local AI models for AMD RX 570

4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Lumina-Next / Lumina-Image 2.05BQ4_K_M3.7 GB
CogVideoX 2B / 5B5BQ4_K_M3.7 GB
DeepSeek-VL24.5BQ5_K_M3.8 GB
DeepFloyd IF4.3BQ5_K_M3.7 GB
Phi-3.5-vision4.2BQ5_K_M3.6 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-4-mini-instruct3.8BQ6_K3.7 GB
Phi-3.5 Mini3.8BQ6_K3.7 GB
OmniGen / OmniGen23.8BQ6_K3.7 GB
SD Cascade (Würstchen v3)3.6BQ6_K3.5 GB
SDXL Turbo3.5BQ6_K3.4 GB
SDXL Lightning3.5BQ6_K3.4 GB
ACE-Step3.5BQ6_K3.4 GB
MusicGen small/medium/large3.3BQ6_K3.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB
Open-Sora Plan2.7BQ8_03.4 GB
LFM2 1.2B / 2.6B2.6BQ8_03.3 GB
Playground v2.52.6BQ8_03.3 GB
Stable Diffusion 3.5 Medium2.5BQ8_03.2 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB

How to read this

The AMD RX 570 graphics card features 4 GB GDDR5 of onboard video memory. This memory size determines which local artificial intelligence models can run directly on your hardware. To fit within this limit, models must use quantization. Quantization is a compression method that reduces the precision of model weights to save space. The quant column shows the best format for each model to balance quality and memory usage.

For fully local execution, your model must fit entirely within the 4 GB GDDR5 limit. The largest fitting models include Lumina-Next or Lumina-Image 2.0 at 5B parameters using the Q4_K_M quant which consumes 3.7 GB of memory. CogVideoX 2B or 5B also fits at 5B parameters using the Q4_K_M quant with 3.7 GB used. DeepSeek-VL2 at 4.5B parameters fits using the Q5_K_M quant with 3.8 GB used. DeepFloyd IF at 4.3B parameters fits using the Q5_K_M quant with 3.7 GB used. Phi-3.5-vision at 4.2B parameters fits using the Q5_K_M quant with 3.6 GB used.

Several 4B parameter models can run using the Q6_K quant which uses 3.9 GB of memory. These models include Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, Danube 3 4B, and Fish Speech 1.5 or OpenAudio S1. You can also run Phi-4-mini-instruct, Phi-3.5 Mini, and OmniGen or OmniGen2 at 3.8B parameters using the Q6_K quant with 3.7 GB used. SD Cascade (Würstchen v3) at 3.6B parameters uses 3.5 GB with the Q6_K quant. SDXL Turbo, SDXL Lightning, and ACE-Step at 3.5B parameters use 3.4 GB with the Q6_K quant. MusicGen small/medium/large at 3.3B parameters uses 3.2 GB with the Q6_K quant.

Models at 3B parameters can run using the Q8_0 quant which uses 3.8 GB of memory. This group includes SmolLM3 3B, Replit Code v1.5 3B, Kandinsky 3.1, Voxtral Mini or Small, Orpheus TTS, and Higgs Audio v2. Smaller models like Allegro at 2.8B parameters use 3.6 GB with the Q8_0 quant. Open-Sora Plan at 2.7B parameters uses 3.4 GB with the Q8_0 quant. LFM2 1.2B or 2.6B and Playground v2.5 at 2.6B parameters use 3.3 GB with the Q8_0 quant. Stable Diffusion 3.5 Medium at 2.5B parameters uses 3.2 GB with the Q8_0 quant.

When a model is too large for the 4 GB GDDR5 memory, you must use CPU offload. This process splits the model layers between your graphics card and your system RAM. Offload allows you to run larger models but it costs performance because system RAM is much slower than GDDR5. For these cases, we assume your computer has 32 GB of system RAM. Stable Diffusion XL at 3.417B parameters needs 4.1 GB at FP8 or optimized settings and 6.1 GB of system RAM. Phi-3 Mini at 3.8B parameters needs 4.4 GB at Q4_K_M and 6.4 GB of system RAM. Phi-4-multimodal at 5.6B parameters needs 4.1 GB at Q4_K_M and 6.1 GB of system RAM. Magicoder-S-DS 6.7B needs 4.9 GB at Q4_K_M and 6.9 GB of system RAM.

Other offload options include Mistral 7B which needs 5.7 GB at Q4_K_M and 7.7 GB of system RAM. Several 7B models need 5.1 GB at Q4_K_M and 7.1 GB of system RAM. These models include Qwen2.5 0.5B or 1.5B or 3B or 7B, OLMo 2 1B or 7B, Falcon 3 1B or 3B or 7B, Command R7B, and OpenHermes 2.5. Be aware that these memory figures are calculated using a basic 4k context window. Running longer text contexts will increase memory usage and may cause the model to exceed the limits.