Best local AI models for AMD Pro 5300
4 GB GDDR6. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Lumina-Next / Lumina-Image 2.0 | 5B | Q4_K_M | 3.7 GB |
| CogVideoX 2B / 5B | 5B | Q4_K_M | 3.7 GB |
| DeepSeek-VL2 | 4.5B | Q5_K_M | 3.8 GB |
| DeepFloyd IF | 4.3B | Q5_K_M | 3.7 GB |
| Phi-3.5-vision | 4.2B | Q5_K_M | 3.6 GB |
| Qwen3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 4 E4B | 4B | Q6_K | 3.9 GB |
| MiniCPM 3 4B | 4B | Q6_K | 3.9 GB |
| Danube 3 4B | 4B | Q6_K | 3.9 GB |
| Fish Speech 1.5 / OpenAudio S1 | 4B | Q6_K | 3.9 GB |
| Phi-4-mini-instruct | 3.8B | Q6_K | 3.7 GB |
| Phi-3.5 Mini | 3.8B | Q6_K | 3.7 GB |
| OmniGen / OmniGen2 | 3.8B | Q6_K | 3.7 GB |
| SD Cascade (Würstchen v3) | 3.6B | Q6_K | 3.5 GB |
| SDXL Turbo | 3.5B | Q6_K | 3.4 GB |
| SDXL Lightning | 3.5B | Q6_K | 3.4 GB |
| ACE-Step | 3.5B | Q6_K | 3.4 GB |
| MusicGen small/medium/large | 3.3B | Q6_K | 3.2 GB |
| SmolLM3 3B | 3B | Q8_0 | 3.8 GB |
| Replit Code v1.5 3B | 3B | Q8_0 | 3.8 GB |
| Kandinsky 3.1 | 3B | Q8_0 | 3.8 GB |
| Voxtral Mini / Small | 3B | Q8_0 | 3.8 GB |
| Orpheus TTS | 3B | Q8_0 | 3.8 GB |
| Higgs Audio v2 | 3B | Q8_0 | 3.8 GB |
| Allegro | 2.8B | Q8_0 | 3.6 GB |
| Open-Sora Plan | 2.7B | Q8_0 | 3.4 GB |
| LFM2 1.2B / 2.6B | 2.6B | Q8_0 | 3.3 GB |
| Playground v2.5 | 2.6B | Q8_0 | 3.3 GB |
| Stable Diffusion 3.5 Medium | 2.5B | Q8_0 | 3.2 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| Stable Diffusion XL | 3.417B | 4.1 GB needed | 6.1 GB |
| Phi-3 Mini | 3.8B | 4.4 GB needed | 6.4 GB |
| Phi-4-multimodal | 5.6B | 4.1 GB needed | 6.1 GB |
| Magicoder-S-DS 6.7B | 6.7B | 4.9 GB needed | 6.9 GB |
| Mistral 7B | 7B | 5.7 GB needed | 7.7 GB |
| Qwen2.5 0.5B / 1.5B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| OLMo 2 1B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Falcon 3 1B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Command R7B | 7B | 5.1 GB needed | 7.1 GB |
| OpenHermes 2.5 | 7B | 5.1 GB needed | 7.1 GB |
How to read this
The AMD Radeon Pro 5300 is an entry level workstation graphics card equipped with 4 GB of GDDR6 memory. This dedicated video memory determines the maximum size of the artificial intelligence models you can run entirely on the hardware. To run a model smoothly without system slowdowns, the model files and active memory states must fit within this 4 GB limit.
The quantization level represents the compression format of the model weights. Choosing a lower quantization like Q4_K_M or Q5_K_M reduces the memory footprint of larger models so they fit into the graphics memory. Higher quantizations like Q6_K or Q8_0 preserve more original model quality but require more memory space per parameter.
For maximum performance, several models fit completely inside the 4 GB limit of the graphics card. The largest options include Lumina-Next or Lumina-Image 2.0 and CogVideoX 2B or 5B, which both use 3.7 GB of memory at the Q4_K_M quantization. DeepSeek-VL2 fits at Q5_K_M using 3.8 GB. You can also run Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, Danube 3 4B, and Fish Speech 1.5 or OpenAudio S1 at Q6_K using 3.9 GB.
Other fully local options include Phi-4-mini-instruct and Phi-3.5 Mini, which both use 3.7 GB at Q6_K. For audio and image generation, SDXL Turbo and SDXL Lightning run at Q6_K using 3.4 GB. If you prefer higher precision, SmolLM3 3B, Replit Code v1.5 3B, and Kandinsky 3.1 fit at the Q8_0 quantization level using 3.8 GB of video memory.
When a model exceeds the 4 GB video memory, you must offload parts of the workload to your system RAM. Assuming your computer has 32 GB of system RAM, you can run larger models with a performance cost. For example, Mistral 7B requires 5.7 GB at Q4_K_M and uses 7.7 GB of system RAM. Qwen2.5 0.5B / 1.5B / 3B / 7B and Falcon 3 1B / 3B / 7B require 5.1 GB at Q4_K_M and use 7.1 GB of system RAM. Stable Diffusion XL requires 4.1 GB at FP8 or optimized settings and uses 6.1 GB of system RAM.
Running models close to the memory limit of the graphics card introduces a strict context window limitation. The standard 4k context window requires extra memory to store the active conversation history. If you generate long responses or input large documents, the active memory will spill over the 4 GB limit and cause severe processing delays.