Best local AI models for AMD RX 5300M

3 GB GDDR6. At a 4k context, 76 of the 233 models in our catalog with verified parameter counts fit fully, up to Qwen3 4B at 4B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 76 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Qwen3 4B4BQ4_K_M2.9 GB
Gemma 3 4B4BQ4_K_M2.9 GB
Gemma 4 E4B4BQ4_K_M2.9 GB
MiniCPM 3 4B4BQ4_K_M2.9 GB
Danube 3 4B4BQ4_K_M2.9 GB
Fish Speech 1.5 / OpenAudio S14BQ4_K_M2.9 GB
Phi-4-mini-instruct3.8BQ4_K_M2.8 GB
Phi-3.5 Mini3.8BQ4_K_M2.8 GB
OmniGen / OmniGen23.8BQ4_K_M2.8 GB
SD Cascade (Würstchen v3)3.6BQ4_K_M2.6 GB
SDXL Turbo3.5BQ5_K_M3 GB
SDXL Lightning3.5BQ5_K_M3 GB
ACE-Step3.5BQ5_K_M3 GB
MusicGen small/medium/large3.3BQ5_K_M2.8 GB
SmolLM3 3B3BQ6_K3 GB
Replit Code v1.5 3B3BQ6_K3 GB
Kandinsky 3.13BQ6_K3 GB
Voxtral Mini / Small3BQ6_K3 GB
Orpheus TTS3BQ6_K3 GB
Higgs Audio v23BQ6_K3 GB
Allegro2.8BQ6_K2.8 GB
Open-Sora Plan2.7BQ6_K2.7 GB
LFM2 1.2B / 2.6B2.6BQ6_K2.6 GB
Playground v2.52.6BQ6_K2.6 GB
Stable Diffusion 3.5 Medium2.5BQ6_K2.5 GB
Canary 1B / Qwen-2.5B2.5BQ6_K2.5 GB
SeamlessM4T v22.3BQ8_02.9 GB
Parler-TTS2.2BQ8_02.8 GB
Kimi K3 DSpark2.2BQ8_02.9 GB
SmolVLM 256M / 500M / 2B2BQ8_02.5 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-3.5-vision4.2B3.1 GB needed5.1 GB
DeepFloyd IF4.3B3.1 GB needed5.1 GB
DeepSeek-VL24.5B3.3 GB needed5.3 GB
Lumina-Next / Lumina-Image 2.05B3.7 GB needed5.7 GB
CogVideoX 2B / 5B5B3.7 GB needed5.7 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB

How to read this

The AMD Radeon RX 5300M is a mobile graphics card equipped with 3 GB of GDDR6 video memory. This memory capacity determines the size of the artificial intelligence models you can run entirely on the hardware. To run a model without slowdowns, the model files and working memory must fit within this 3 GB limit. If a model exceeds this boundary, your system must use slower system memory, which reduces generation speeds.

Quantization is a compression method that reduces model size so it fits into your video memory. In our listings, the best quant column shows the highest quality quantization level that fits the graphics card. For example, Q4_K_M represents a four bit quantization that balances model accuracy and memory usage. Higher levels like Q5_K_M, Q6_K, and Q8_0 offer better output quality but require more memory space.

Several 4B models fit within the hardware limits using the Q4_K_M quantization. Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, Danube 3 4B, and Fish Speech 1.5 / OpenAudio S1 each use 2.9 GB of video memory. You can also run Phi-4-mini-instruct, Phi-3.5 Mini, and OmniGen / OmniGen2 at 2.8 GB of memory usage. Image generators like SD Cascade (Würstchen v3) use 2.6 GB of memory with the same quantization.

If you use higher quantization levels, you must select smaller models to stay under the 3 GB limit. SDXL Turbo, SDXL Lightning, and ACE-Step use 3 GB of video memory at Q5_K_M. MusicGen small/medium/large fits at Q5_K_M using 2.8 GB. At the Q6_K level, SmolLM3 3B, Replit Code v1.5 3B, Kandinsky 3.1, Voxtral Mini / Small, Orpheus TTS, and Higgs Audio v2 use exactly 3 GB. Allegro uses 2.8 GB, Open-Sora Plan uses 2.7 GB, and both LFM2 1.2B / 2.6B and Playground v2.5 use 2.6 GB. Stable Diffusion 3.5 Medium and Canary 1B / Qwen-2.5B use 2.5 GB. At Q8_0, SeamlessM4T v2 uses 2.9 GB, Parler-TTS uses 2.8 GB, Kimi K3 DSpark uses 2.9 GB, and SmolVLM 256M / 500M / 2B uses 2.5 GB.

When a model is too large for the video memory, you can use CPU offload if your computer has 32 GB of system RAM. This process splits the workload between your graphics card and your system memory. Stable Diffusion XL needs 4.1 GB at FP8 and requires 6.1 GB of system RAM. Phi-3 Mini needs 4.4 GB at Q4_K_M and requires 6.4 GB of system RAM. Phi-3.5-vision and DeepFloyd IF need 3.1 GB at Q4_K_M and require 5.1 GB of system RAM. DeepSeek-VL2 needs 3.3 GB at Q4_K_M and requires 5.3 GB of system RAM.

Other larger models also run through CPU offload with 32 GB of system RAM. Lumina-Next / Lumina-Image 2.0 and CogVideoX 2B / 5B need 3.7 GB at Q4_K_M and require 5.7 GB of system RAM. Phi-4-multimodal needs 4.1 GB at Q4_K_M and requires 6.1 GB of system RAM. Magicoder-S-DS 6.7B needs 4.9 GB at Q4_K_M and requires 6.9 GB of system RAM. Mistral 7B needs 5.7 GB at Q4_K_M and requires 7.7 GB of system RAM. Offloading allows you to run these larger models but it reduces the processing speed significantly.

You must also consider the context window size when loading these models. The memory numbers listed here assume a standard 4k context window. If you increase the context window to process longer texts, the model will require more video memory. This extra memory demand can push a model over the 3 GB limit and force the system to slow down.