Best local AI models for AMD R5 M320
4 GB DDR3. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Lumina-Next / Lumina-Image 2.0 | 5B | Q4_K_M | 3.7 GB |
| CogVideoX 2B / 5B | 5B | Q4_K_M | 3.7 GB |
| DeepSeek-VL2 | 4.5B | Q5_K_M | 3.8 GB |
| DeepFloyd IF | 4.3B | Q5_K_M | 3.7 GB |
| Phi-3.5-vision | 4.2B | Q5_K_M | 3.6 GB |
| Qwen3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 4 E4B | 4B | Q6_K | 3.9 GB |
| MiniCPM 3 4B | 4B | Q6_K | 3.9 GB |
| Danube 3 4B | 4B | Q6_K | 3.9 GB |
| Fish Speech 1.5 / OpenAudio S1 | 4B | Q6_K | 3.9 GB |
| Phi-4-mini-instruct | 3.8B | Q6_K | 3.7 GB |
| Phi-3.5 Mini | 3.8B | Q6_K | 3.7 GB |
| OmniGen / OmniGen2 | 3.8B | Q6_K | 3.7 GB |
| SD Cascade (Würstchen v3) | 3.6B | Q6_K | 3.5 GB |
| SDXL Turbo | 3.5B | Q6_K | 3.4 GB |
| SDXL Lightning | 3.5B | Q6_K | 3.4 GB |
| ACE-Step | 3.5B | Q6_K | 3.4 GB |
| MusicGen small/medium/large | 3.3B | Q6_K | 3.2 GB |
| SmolLM3 3B | 3B | Q8_0 | 3.8 GB |
| Replit Code v1.5 3B | 3B | Q8_0 | 3.8 GB |
| Kandinsky 3.1 | 3B | Q8_0 | 3.8 GB |
| Voxtral Mini / Small | 3B | Q8_0 | 3.8 GB |
| Orpheus TTS | 3B | Q8_0 | 3.8 GB |
| Higgs Audio v2 | 3B | Q8_0 | 3.8 GB |
| Allegro | 2.8B | Q8_0 | 3.6 GB |
| Open-Sora Plan | 2.7B | Q8_0 | 3.4 GB |
| LFM2 1.2B / 2.6B | 2.6B | Q8_0 | 3.3 GB |
| Playground v2.5 | 2.6B | Q8_0 | 3.3 GB |
| Stable Diffusion 3.5 Medium | 2.5B | Q8_0 | 3.2 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| Stable Diffusion XL | 3.417B | 4.1 GB needed | 6.1 GB |
| Phi-3 Mini | 3.8B | 4.4 GB needed | 6.4 GB |
| Phi-4-multimodal | 5.6B | 4.1 GB needed | 6.1 GB |
| Magicoder-S-DS 6.7B | 6.7B | 4.9 GB needed | 6.9 GB |
| Mistral 7B | 7B | 5.7 GB needed | 7.7 GB |
| Qwen2.5 0.5B / 1.5B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| OLMo 2 1B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Falcon 3 1B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Command R7B | 7B | 5.1 GB needed | 7.1 GB |
| OpenHermes 2.5 | 7B | 5.1 GB needed | 7.1 GB |
How to read this
The AMD Radeon R5 M320 is an entry level graphics card equipped with 4 GB of DDR3 video memory. This dedicated memory pool determines the size of the artificial intelligence models you can run entirely on the hardware. When running models locally, the model weights must fit inside this 4 GB space to avoid severe performance slowdowns. Keeping your model size below this threshold ensures that the graphics processor can access the data quickly.
To fit larger models into the limited video memory, developers use quantization. The quant column indicates the compression level applied to each model. For example, a Q4_K_M quant uses approximately four bits per parameter, while a Q8_0 quant uses eight bits per parameter. Lower quantization levels like Q4_K_M allow larger models like the 5B Lumina-Next or CogVideoX to fit into 3.7 GB of memory. Higher quantization levels like Q8_0 preserve more model accuracy but require more memory per parameter, limiting you to smaller models like the 3B SmolLM3.
If a model exceeds the 4 GB video memory limit, you must use CPU offloading. This technique splits the model layers between your graphics card and your system memory. We assume your computer has 32 GB of system RAM for these scenarios. For instance, running Mistral 7B requires 5.7 GB of video memory at Q4_K_M, which forces 7.7 GB of data into your system RAM. While offloading allows you to run larger models like Qwen2.5 7B or Falcon 3 7B, the slow DDR3 interface will cause a significant drop in generation speed.
You must also consider the memory cost of context length. The listed memory usage figures represent the base model size before you type any prompts. As you write longer conversations, the active memory usage increases. A standard 4k context window requires additional video memory to store the active tokens. If your base model already uses 3.9 GB of your 4 GB limit, like Gemma 3 4B at Q6_K, processing a long conversation will likely exceed your physical memory and trigger slow system RAM usage.
For the best balance of speed and quality on this hardware, choose models that leave a comfortable safety margin. Models under 3.5B parameters, such as SDXL Turbo at 3.4 GB or MusicGen at 3.2 GB, leave enough room for processing. If you need to run larger 7B models like Command R7B or OpenHermes 2.5, prepare to configure your software for CPU offloading and expect slower processing times due to the shared system memory bottleneck.