Best local AI models for AMD R5 M430
4 GB DDR3. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Lumina-Next / Lumina-Image 2.0 | 5B | Q4_K_M | 3.7 GB |
| CogVideoX 2B / 5B | 5B | Q4_K_M | 3.7 GB |
| DeepSeek-VL2 | 4.5B | Q5_K_M | 3.8 GB |
| DeepFloyd IF | 4.3B | Q5_K_M | 3.7 GB |
| Phi-3.5-vision | 4.2B | Q5_K_M | 3.6 GB |
| Qwen3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 4 E4B | 4B | Q6_K | 3.9 GB |
| MiniCPM 3 4B | 4B | Q6_K | 3.9 GB |
| Danube 3 4B | 4B | Q6_K | 3.9 GB |
| Fish Speech 1.5 / OpenAudio S1 | 4B | Q6_K | 3.9 GB |
| Phi-4-mini-instruct | 3.8B | Q6_K | 3.7 GB |
| Phi-3.5 Mini | 3.8B | Q6_K | 3.7 GB |
| OmniGen / OmniGen2 | 3.8B | Q6_K | 3.7 GB |
| SD Cascade (Würstchen v3) | 3.6B | Q6_K | 3.5 GB |
| SDXL Turbo | 3.5B | Q6_K | 3.4 GB |
| SDXL Lightning | 3.5B | Q6_K | 3.4 GB |
| ACE-Step | 3.5B | Q6_K | 3.4 GB |
| MusicGen small/medium/large | 3.3B | Q6_K | 3.2 GB |
| SmolLM3 3B | 3B | Q8_0 | 3.8 GB |
| Replit Code v1.5 3B | 3B | Q8_0 | 3.8 GB |
| Kandinsky 3.1 | 3B | Q8_0 | 3.8 GB |
| Voxtral Mini / Small | 3B | Q8_0 | 3.8 GB |
| Orpheus TTS | 3B | Q8_0 | 3.8 GB |
| Higgs Audio v2 | 3B | Q8_0 | 3.8 GB |
| Allegro | 2.8B | Q8_0 | 3.6 GB |
| Open-Sora Plan | 2.7B | Q8_0 | 3.4 GB |
| LFM2 1.2B / 2.6B | 2.6B | Q8_0 | 3.3 GB |
| Playground v2.5 | 2.6B | Q8_0 | 3.3 GB |
| Stable Diffusion 3.5 Medium | 2.5B | Q8_0 | 3.2 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| Stable Diffusion XL | 3.417B | 4.1 GB needed | 6.1 GB |
| Phi-3 Mini | 3.8B | 4.4 GB needed | 6.4 GB |
| Phi-4-multimodal | 5.6B | 4.1 GB needed | 6.1 GB |
| Magicoder-S-DS 6.7B | 6.7B | 4.9 GB needed | 6.9 GB |
| Mistral 7B | 7B | 5.7 GB needed | 7.7 GB |
| Qwen2.5 0.5B / 1.5B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| OLMo 2 1B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Falcon 3 1B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Command R7B | 7B | 5.1 GB needed | 7.1 GB |
| OpenHermes 2.5 | 7B | 5.1 GB needed | 7.1 GB |
How to read this
The AMD Radeon R5 M430 is an entry level laptop graphics card equipped with 4 GB of DDR3 memory. This dedicated video memory determines the maximum size of the neural network weights you can load directly onto the hardware. Because DDR3 memory has lower bandwidth than modern GDDR graphics memory, keeping the entire model within the local 4 GB limit is critical to prevent severe performance bottlenecks.
To fit models into this memory budget, you must use quantized versions. Quantization is a compression technique that reduces the precision of model weights. The quant column shows the optimal format for each model. For example, a Q4_K_M quant uses approximately four bits per weight, while a Q6_K or Q8_0 quant uses six or eight bits. Higher quantization levels like Q8_0 preserve more output quality but require more memory space.
Several capable models fit entirely within the 4 GB limit of your hardware. The largest options include Lumina-Next or Lumina-Image 2.0 and CogVideoX 2B or 5B, which both run at the Q4_K_M quant using 3.7 GB of memory. Vision models like DeepSeek-VL2 and Phi-3.5-vision fit at Q5_K_M quants, using 3.8 GB and 3.6 GB respectively. Text models like Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, and MiniCPM 3 4B can run at the higher quality Q6_K quant, using 3.9 GB of video memory.
For audio and image generation, you can run SDXL Turbo or SDXL Lightning at the Q6_K quant using 3.4 GB of memory. Audio options like MusicGen small/medium/large fit at Q6_K using 3.2 GB. If you prefer higher precision, smaller models like SmolLM3 3B, Replit Code v1.5 3B, and Kandinsky 3.1 run at the Q8_0 quant, using 3.8 GB of video memory. These options run entirely on the graphics processor for maximum possible speed.
When a model exceeds 4 GB, you must use CPU offload to share the workload with your system RAM. Assuming your computer has 32 GB of system RAM, you can run larger models by splitting the weights. For instance, Mistral 7B requires 5.7 GB at the Q4_K_M quant, which uses 7.7 GB of system RAM. Similarly, Qwen2.5 7B, OLMo 2 7B, Falcon 3 7B, Command R7B, and OpenHermes 2.5 require 5.1 GB at Q4_K_M, using 7.1 GB of system RAM. Offloading allows you to run these larger models, but it significantly reduces generation speed because data must travel over the slower system bus.
You must also consider the memory cost of context length. The memory figures listed here are calculated using a standard 4k context window. If you increase the context window to process longer documents or chat histories, the key-value cache will consume additional memory. On a 4 GB card, expanding the context beyond 4k will quickly exceed your video memory and force the system to slow down.