Best local AI models for AMD RX 5500M
4 GB GDDR6. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Lumina-Next / Lumina-Image 2.0 | 5B | Q4_K_M | 3.7 GB |
| CogVideoX 2B / 5B | 5B | Q4_K_M | 3.7 GB |
| DeepSeek-VL2 | 4.5B | Q5_K_M | 3.8 GB |
| DeepFloyd IF | 4.3B | Q5_K_M | 3.7 GB |
| Phi-3.5-vision | 4.2B | Q5_K_M | 3.6 GB |
| Qwen3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 4 E4B | 4B | Q6_K | 3.9 GB |
| MiniCPM 3 4B | 4B | Q6_K | 3.9 GB |
| Danube 3 4B | 4B | Q6_K | 3.9 GB |
| Fish Speech 1.5 / OpenAudio S1 | 4B | Q6_K | 3.9 GB |
| Phi-4-mini-instruct | 3.8B | Q6_K | 3.7 GB |
| Phi-3.5 Mini | 3.8B | Q6_K | 3.7 GB |
| OmniGen / OmniGen2 | 3.8B | Q6_K | 3.7 GB |
| SD Cascade (Würstchen v3) | 3.6B | Q6_K | 3.5 GB |
| SDXL Turbo | 3.5B | Q6_K | 3.4 GB |
| SDXL Lightning | 3.5B | Q6_K | 3.4 GB |
| ACE-Step | 3.5B | Q6_K | 3.4 GB |
| MusicGen small/medium/large | 3.3B | Q6_K | 3.2 GB |
| SmolLM3 3B | 3B | Q8_0 | 3.8 GB |
| Replit Code v1.5 3B | 3B | Q8_0 | 3.8 GB |
| Kandinsky 3.1 | 3B | Q8_0 | 3.8 GB |
| Voxtral Mini / Small | 3B | Q8_0 | 3.8 GB |
| Orpheus TTS | 3B | Q8_0 | 3.8 GB |
| Higgs Audio v2 | 3B | Q8_0 | 3.8 GB |
| Allegro | 2.8B | Q8_0 | 3.6 GB |
| Open-Sora Plan | 2.7B | Q8_0 | 3.4 GB |
| LFM2 1.2B / 2.6B | 2.6B | Q8_0 | 3.3 GB |
| Playground v2.5 | 2.6B | Q8_0 | 3.3 GB |
| Stable Diffusion 3.5 Medium | 2.5B | Q8_0 | 3.2 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| Stable Diffusion XL | 3.417B | 4.1 GB needed | 6.1 GB |
| Phi-3 Mini | 3.8B | 4.4 GB needed | 6.4 GB |
| Phi-4-multimodal | 5.6B | 4.1 GB needed | 6.1 GB |
| Magicoder-S-DS 6.7B | 6.7B | 4.9 GB needed | 6.9 GB |
| Mistral 7B | 7B | 5.7 GB needed | 7.7 GB |
| Qwen2.5 0.5B / 1.5B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| OLMo 2 1B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Falcon 3 1B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Command R7B | 7B | 5.1 GB needed | 7.1 GB |
| OpenHermes 2.5 | 7B | 5.1 GB needed | 7.1 GB |
How to read this
The AMD Radeon RX 5500M is a mobile graphics card equipped with 4 GB of GDDR6 video memory. This dedicated memory size determines which artificial intelligence models can run entirely on your hardware. To run a model locally without slowdowns, the model files and active memory must fit within this 4 GB limit. If a model exceeds this capacity, your system must use slower memory types.
The quantization column shows the compression level applied to each model. Quantization reduces the size of model weights to save video memory. For example, a Q4_K_M quant uses fewer bits per weight than a Q6_K or Q8_0 quant. Smaller quants allow larger models to fit into the 4 GB limit, but they may slightly reduce output quality. Higher quants like Q8_0 offer better quality but require more memory.
Several capable models fit completely inside the 4 GB video memory of the RX 5500M. The Lumina-Next or Lumina-Image 2.0 model at 5B parameters fits using a Q4_K_M quant which uses 3.7 GB. CogVideoX 2B or 5B also fits at 5B parameters using Q4_K_M with 3.7 GB used. DeepSeek-VL2 at 4.5B parameters fits using a Q5_K_M quant which uses 3.8 GB. For text tasks, Phi-4-mini-instruct at 3.8B parameters fits using a Q6_K quant which uses 3.7 GB.
You can also run image generation and audio models directly on the card. SDXL Turbo and SDXL Lightning at 3.5B parameters fit using a Q6_K quant which uses 3.4 GB. Stable Diffusion 3.5 Medium at 2.5B parameters fits using a Q8_0 quant which uses 3.2 GB. For audio, Fish Speech 1.5 or OpenAudio S1 at 4B parameters fits using a Q6_K quant which uses 3.9 GB. MusicGen small/medium/large at 3.3B parameters fits using a Q6_K quant which uses 3.2 GB.
When a model is too large for the 4 GB video memory, you must use CPU offload. This process splits the model between your graphics card and your system RAM. For example, Mistral 7B at 7B parameters needs 5.7 GB at Q4_K_M and requires 7.7 GB of system RAM. Stable Diffusion XL at 3.417B parameters needs 4.1 GB at FP8 or optimized settings and requires 6.1 GB of system RAM. CPU offload allows you to run these larger models but reduces processing speed.
Memory calculations assume a standard 4k context window for text generation. Running longer conversations or processing larger documents increases memory usage. If you exceed the 4k context limit, the model will require more memory than the listed figures. This extra memory demand can push a fitting model over the 4 GB limit and trigger slow system RAM usage.