Best local AI models for AMD Pro 5300M
4 GB GDDR6. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Lumina-Next / Lumina-Image 2.0 | 5B | Q4_K_M | 3.7 GB |
| CogVideoX 2B / 5B | 5B | Q4_K_M | 3.7 GB |
| DeepSeek-VL2 | 4.5B | Q5_K_M | 3.8 GB |
| DeepFloyd IF | 4.3B | Q5_K_M | 3.7 GB |
| Phi-3.5-vision | 4.2B | Q5_K_M | 3.6 GB |
| Qwen3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 4 E4B | 4B | Q6_K | 3.9 GB |
| MiniCPM 3 4B | 4B | Q6_K | 3.9 GB |
| Danube 3 4B | 4B | Q6_K | 3.9 GB |
| Fish Speech 1.5 / OpenAudio S1 | 4B | Q6_K | 3.9 GB |
| Phi-4-mini-instruct | 3.8B | Q6_K | 3.7 GB |
| Phi-3.5 Mini | 3.8B | Q6_K | 3.7 GB |
| OmniGen / OmniGen2 | 3.8B | Q6_K | 3.7 GB |
| SD Cascade (Würstchen v3) | 3.6B | Q6_K | 3.5 GB |
| SDXL Turbo | 3.5B | Q6_K | 3.4 GB |
| SDXL Lightning | 3.5B | Q6_K | 3.4 GB |
| ACE-Step | 3.5B | Q6_K | 3.4 GB |
| MusicGen small/medium/large | 3.3B | Q6_K | 3.2 GB |
| SmolLM3 3B | 3B | Q8_0 | 3.8 GB |
| Replit Code v1.5 3B | 3B | Q8_0 | 3.8 GB |
| Kandinsky 3.1 | 3B | Q8_0 | 3.8 GB |
| Voxtral Mini / Small | 3B | Q8_0 | 3.8 GB |
| Orpheus TTS | 3B | Q8_0 | 3.8 GB |
| Higgs Audio v2 | 3B | Q8_0 | 3.8 GB |
| Allegro | 2.8B | Q8_0 | 3.6 GB |
| Open-Sora Plan | 2.7B | Q8_0 | 3.4 GB |
| LFM2 1.2B / 2.6B | 2.6B | Q8_0 | 3.3 GB |
| Playground v2.5 | 2.6B | Q8_0 | 3.3 GB |
| Stable Diffusion 3.5 Medium | 2.5B | Q8_0 | 3.2 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| Stable Diffusion XL | 3.417B | 4.1 GB needed | 6.1 GB |
| Phi-3 Mini | 3.8B | 4.4 GB needed | 6.4 GB |
| Phi-4-multimodal | 5.6B | 4.1 GB needed | 6.1 GB |
| Magicoder-S-DS 6.7B | 6.7B | 4.9 GB needed | 6.9 GB |
| Mistral 7B | 7B | 5.7 GB needed | 7.7 GB |
| Qwen2.5 0.5B / 1.5B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| OLMo 2 1B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Falcon 3 1B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Command R7B | 7B | 5.1 GB needed | 7.1 GB |
| OpenHermes 2.5 | 7B | 5.1 GB needed | 7.1 GB |
How to read this
The AMD Radeon Pro 5300M is a mobile graphics processor equipped with 4 GB of GDDR6 video memory. This dedicated memory capacity determines the maximum size of the artificial intelligence models you can run entirely on the hardware. To execute a model without performance loss, the model files and the active processing data must fit within this 4 GB limit. If a model exceeds this boundary, your system must transfer data to your system memory, which slows down execution speeds.
The quantization column indicates the compression level applied to each model. Quantization reduces the precision of model weights to make the files smaller. For example, Q4_K_M represents a medium four bit quantization, while Q6_K and Q8_0 represent six bit and eight bit formats. Higher quantization levels like Q8_0 preserve more original model accuracy but require more video memory. Lower levels like Q4_K_M allow larger models such as the 5B Lumina-Next or CogVideoX 2B to fit inside the 4 GB frame by using 3.7 GB of memory.
Several highly capable models fit entirely within the local video memory. You can run the 4.5B DeepSeek-VL2 at Q5_K_M using 3.8 GB of memory or the 4.2B Phi-3.5-vision at Q5_K_M using 3.6 GB of memory. Text models like Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, and Danube 3 4B fit at Q6_K quantization while using 3.9 GB of memory. Audio and speech models like Fish Speech 1.5, Orpheus TTS, and Higgs Audio v2 also run locally within these memory limits.
When a model is too large for the 4 GB video memory, you can use CPU offload. This technique splits the workload between your graphics card and your system RAM. For these offload cases, we assume your computer has 32 GB of system RAM. Running Mistral 7B at Q4_K_M requires 5.7 GB of video memory and 7.7 GB of system RAM. Similarly, Qwen2.5 7B, OLMo 2 7B, Falcon 3 7B, Command R7B, and OpenHermes 2.5 7B require 5.1 GB of video memory and 7.1 GB of system RAM.
Other offload options include the 5.6B Phi-4-multimodal which needs 4.1 GB of video memory and 6.1 GB of system RAM at Q4_K_M. The Magicoder-S-DS 6.7B needs 4.9 GB of video memory and 6.9 GB of system RAM. Image models like Stable Diffusion XL need 4.1 GB of video memory and 6.1 GB of system RAM at FP8 or optimized settings. While offloading enables you to run these larger models, the transfer of data between system RAM and video memory reduces generation speed.
You must also consider the memory cost of context length. The memory figures listed here represent the model at its base state. As you input longer text prompts and generate longer responses, the active memory usage increases. Running a model close to the 4 GB limit, such as Phi-4-mini-instruct at 3.7 GB, leaves very little room for context. If your conversation history exceeds the remaining memory, the system will slow down as it offloads the active context to system RAM.