Best local AI models for AMD R5 M320

4 GB DDR3. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Lumina-Next / Lumina-Image 2.05BQ4_K_M3.7 GB
CogVideoX 2B / 5B5BQ4_K_M3.7 GB
DeepSeek-VL24.5BQ5_K_M3.8 GB
DeepFloyd IF4.3BQ5_K_M3.7 GB
Phi-3.5-vision4.2BQ5_K_M3.6 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-4-mini-instruct3.8BQ6_K3.7 GB
Phi-3.5 Mini3.8BQ6_K3.7 GB
OmniGen / OmniGen23.8BQ6_K3.7 GB
SD Cascade (Würstchen v3)3.6BQ6_K3.5 GB
SDXL Turbo3.5BQ6_K3.4 GB
SDXL Lightning3.5BQ6_K3.4 GB
ACE-Step3.5BQ6_K3.4 GB
MusicGen small/medium/large3.3BQ6_K3.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB
Open-Sora Plan2.7BQ8_03.4 GB
LFM2 1.2B / 2.6B2.6BQ8_03.3 GB
Playground v2.52.6BQ8_03.3 GB
Stable Diffusion 3.5 Medium2.5BQ8_03.2 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB

How to read this

The AMD Radeon R5 M320 is an entry level graphics card equipped with 4 GB of DDR3 video memory. This dedicated memory pool determines the size of the artificial intelligence models you can run entirely on the hardware. When running models locally, the model weights must fit inside this 4 GB space to avoid severe performance slowdowns. Keeping your model size below this threshold ensures that the graphics processor can access the data quickly.

To fit larger models into the limited video memory, developers use quantization. The quant column indicates the compression level applied to each model. For example, a Q4_K_M quant uses approximately four bits per parameter, while a Q8_0 quant uses eight bits per parameter. Lower quantization levels like Q4_K_M allow larger models like the 5B Lumina-Next or CogVideoX to fit into 3.7 GB of memory. Higher quantization levels like Q8_0 preserve more model accuracy but require more memory per parameter, limiting you to smaller models like the 3B SmolLM3.

If a model exceeds the 4 GB video memory limit, you must use CPU offloading. This technique splits the model layers between your graphics card and your system memory. We assume your computer has 32 GB of system RAM for these scenarios. For instance, running Mistral 7B requires 5.7 GB of video memory at Q4_K_M, which forces 7.7 GB of data into your system RAM. While offloading allows you to run larger models like Qwen2.5 7B or Falcon 3 7B, the slow DDR3 interface will cause a significant drop in generation speed.

You must also consider the memory cost of context length. The listed memory usage figures represent the base model size before you type any prompts. As you write longer conversations, the active memory usage increases. A standard 4k context window requires additional video memory to store the active tokens. If your base model already uses 3.9 GB of your 4 GB limit, like Gemma 3 4B at Q6_K, processing a long conversation will likely exceed your physical memory and trigger slow system RAM usage.

For the best balance of speed and quality on this hardware, choose models that leave a comfortable safety margin. Models under 3.5B parameters, such as SDXL Turbo at 3.4 GB or MusicGen at 3.2 GB, leave enough room for processing. If you need to run larger 7B models like Command R7B or OpenHermes 2.5, prepare to configure your software for CPU offloading and expect slower processing times due to the shared system memory bottleneck.