Best local AI models for AMD RX 580 2048SP

4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Lumina-Next / Lumina-Image 2.05BQ4_K_M3.7 GB
CogVideoX 2B / 5B5BQ4_K_M3.7 GB
DeepSeek-VL24.5BQ5_K_M3.8 GB
DeepFloyd IF4.3BQ5_K_M3.7 GB
Phi-3.5-vision4.2BQ5_K_M3.6 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-4-mini-instruct3.8BQ6_K3.7 GB
Phi-3.5 Mini3.8BQ6_K3.7 GB
OmniGen / OmniGen23.8BQ6_K3.7 GB
SD Cascade (Würstchen v3)3.6BQ6_K3.5 GB
SDXL Turbo3.5BQ6_K3.4 GB
SDXL Lightning3.5BQ6_K3.4 GB
ACE-Step3.5BQ6_K3.4 GB
MusicGen small/medium/large3.3BQ6_K3.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB
Open-Sora Plan2.7BQ8_03.4 GB
LFM2 1.2B / 2.6B2.6BQ8_03.3 GB
Playground v2.52.6BQ8_03.3 GB
Stable Diffusion 3.5 Medium2.5BQ8_03.2 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB

How to read this

The AMD RX 580 2048SP graphics card features 4 GB of GDDR5 memory. This memory size determines which artificial intelligence models can run entirely on your hardware. When a model fits completely inside this 4 GB limit, it runs at the fastest possible speed. If a model exceeds this limit, you must use system memory to help run it.

The quant column shows the quantization level used to compress each model. Quantization reduces the size of a model so it can fit into smaller memory spaces. For example, the 5B Lumina-Next model fits inside 3.7 GB of memory when using the Q4_K_M quant. Other models like the 4B Qwen3 and Gemma 3 use the Q6_K quant to fit inside 3.9 GB of memory. Smaller models like the 3B SmolLM3 can use the larger Q8_0 quant and fit inside 3.8 GB of memory.

Running models with a 4k context window requires extra memory. The context window is the memory used to remember the conversation history. When you use a 4k context window, you must leave enough free space in your 4 GB of video memory. If the model and the context window together exceed 4 GB, your system will slow down.

You can run larger models by offloading some work to your system RAM. This process is called CPU offload. It requires a system with 32 GB of system RAM. Offloading allows you to run models that are too large for your graphics card alone. However, offloading costs speed because system RAM is much slower than the GDDR5 memory on your graphics card.

Several popular models can run using CPU offload on this system. The Mistral 7B model needs 5.7 GB at the Q4_K_M quant and uses 7.7 GB of system RAM. The Qwen2.5 7B, OLMo 2 7B, Falcon 3 7B, Command R7B, and OpenHermes 2.5 models each need 5.1 GB at the Q4_K_M quant and use 7.1 GB of system RAM. The Phi-4-multimodal model needs 4.1 GB at the Q4_K_M quant and uses 6.1 GB of system RAM.