Best local AI models for AMD RX 5500 XT
4 GB GDDR6. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Lumina-Next / Lumina-Image 2.0 | 5B | Q4_K_M | 3.7 GB |
| CogVideoX 2B / 5B | 5B | Q4_K_M | 3.7 GB |
| DeepSeek-VL2 | 4.5B | Q5_K_M | 3.8 GB |
| DeepFloyd IF | 4.3B | Q5_K_M | 3.7 GB |
| Phi-3.5-vision | 4.2B | Q5_K_M | 3.6 GB |
| Qwen3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 4 E4B | 4B | Q6_K | 3.9 GB |
| MiniCPM 3 4B | 4B | Q6_K | 3.9 GB |
| Danube 3 4B | 4B | Q6_K | 3.9 GB |
| Fish Speech 1.5 / OpenAudio S1 | 4B | Q6_K | 3.9 GB |
| Phi-4-mini-instruct | 3.8B | Q6_K | 3.7 GB |
| Phi-3.5 Mini | 3.8B | Q6_K | 3.7 GB |
| OmniGen / OmniGen2 | 3.8B | Q6_K | 3.7 GB |
| SD Cascade (Würstchen v3) | 3.6B | Q6_K | 3.5 GB |
| SDXL Turbo | 3.5B | Q6_K | 3.4 GB |
| SDXL Lightning | 3.5B | Q6_K | 3.4 GB |
| ACE-Step | 3.5B | Q6_K | 3.4 GB |
| MusicGen small/medium/large | 3.3B | Q6_K | 3.2 GB |
| SmolLM3 3B | 3B | Q8_0 | 3.8 GB |
| Replit Code v1.5 3B | 3B | Q8_0 | 3.8 GB |
| Kandinsky 3.1 | 3B | Q8_0 | 3.8 GB |
| Voxtral Mini / Small | 3B | Q8_0 | 3.8 GB |
| Orpheus TTS | 3B | Q8_0 | 3.8 GB |
| Higgs Audio v2 | 3B | Q8_0 | 3.8 GB |
| Allegro | 2.8B | Q8_0 | 3.6 GB |
| Open-Sora Plan | 2.7B | Q8_0 | 3.4 GB |
| LFM2 1.2B / 2.6B | 2.6B | Q8_0 | 3.3 GB |
| Playground v2.5 | 2.6B | Q8_0 | 3.3 GB |
| Stable Diffusion 3.5 Medium | 2.5B | Q8_0 | 3.2 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| Stable Diffusion XL | 3.417B | 4.1 GB needed | 6.1 GB |
| Phi-3 Mini | 3.8B | 4.4 GB needed | 6.4 GB |
| Phi-4-multimodal | 5.6B | 4.1 GB needed | 6.1 GB |
| Magicoder-S-DS 6.7B | 6.7B | 4.9 GB needed | 6.9 GB |
| Mistral 7B | 7B | 5.7 GB needed | 7.7 GB |
| Qwen2.5 0.5B / 1.5B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| OLMo 2 1B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Falcon 3 1B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Command R7B | 7B | 5.1 GB needed | 7.1 GB |
| OpenHermes 2.5 | 7B | 5.1 GB needed | 7.1 GB |
How to read this
The AMD Radeon RX 5500 XT graphics card features 4 GB of GDDR6 video memory. This onboard memory size determines which artificial intelligence models can run directly on your hardware. For local execution without slowdowns, the model weights and active memory must fit inside this 4 GB limit. If a model exceeds this capacity, your system must use alternative execution strategies.
The quantization column indicates the compression level applied to each model. Quantization reduces the precision of model weights to save space. For example, a Q6_K quant represents a high quality compression level that fits a 4B model like Gemma 3 4B into 3.9 GB of video memory. A Q8_0 quant offers even higher fidelity but requires more space, allowing a 3B model like SmolLM3 3B to occupy 3.8 GB of video memory.
Several capable models fit entirely within the local video memory of your card. You can run Lumina-Next or Lumina-Image 2.0 at a Q4_K_M quant using 3.7 GB of memory. Visual models like DeepSeek-VL2 at Q5_K_M use 3.8 GB of memory. Text models such as Phi-4-mini-instruct and Phi-3.5 Mini run at Q6_K quant using 3.7 GB of memory. Image generation models like SDXL Turbo fit at Q6_K quant using 3.4 GB of memory.
When a model is too large for the 4 GB video memory, you must use CPU offload. This process splits the model layers between your graphics card and your system RAM. Assuming you have 32 GB of system RAM, you can run larger architectures. For example, Mistral 7B at Q4_K_M needs 5.7 GB of memory, which requires 7.7 GB of system RAM. Similarly, Qwen2.5 7B at Q4_K_M needs 5.1 GB of memory and requires 7.1 GB of system RAM.
CPU offload allows you to run advanced models, but it introduces a performance cost. Transferring data between system RAM and video memory is much slower than keeping data on the graphics card. This transfer delay reduces the speed of text generation or image creation. Models like Stable Diffusion XL at FP8 need 4.1 GB of memory and require 6.1 GB of system RAM, which will run slower than models that fit entirely on the card.
Memory consumption estimates are calculated using a standard 4k context window. Running longer conversations or processing larger documents increases memory usage. If you exceed the 4k context limit, the model will require more memory than the listed figures. This extra memory demand can cause a model that normally fits on the card to spill over into system RAM, which slows down performance.