Best local AI models for AMD R5 M430

4 GB DDR3. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Lumina-Next / Lumina-Image 2.05BQ4_K_M3.7 GB
CogVideoX 2B / 5B5BQ4_K_M3.7 GB
DeepSeek-VL24.5BQ5_K_M3.8 GB
DeepFloyd IF4.3BQ5_K_M3.7 GB
Phi-3.5-vision4.2BQ5_K_M3.6 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-4-mini-instruct3.8BQ6_K3.7 GB
Phi-3.5 Mini3.8BQ6_K3.7 GB
OmniGen / OmniGen23.8BQ6_K3.7 GB
SD Cascade (Würstchen v3)3.6BQ6_K3.5 GB
SDXL Turbo3.5BQ6_K3.4 GB
SDXL Lightning3.5BQ6_K3.4 GB
ACE-Step3.5BQ6_K3.4 GB
MusicGen small/medium/large3.3BQ6_K3.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB
Open-Sora Plan2.7BQ8_03.4 GB
LFM2 1.2B / 2.6B2.6BQ8_03.3 GB
Playground v2.52.6BQ8_03.3 GB
Stable Diffusion 3.5 Medium2.5BQ8_03.2 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB

How to read this

The AMD Radeon R5 M430 is an entry level laptop graphics card equipped with 4 GB of DDR3 memory. This dedicated video memory determines the maximum size of the neural network weights you can load directly onto the hardware. Because DDR3 memory has lower bandwidth than modern GDDR graphics memory, keeping the entire model within the local 4 GB limit is critical to prevent severe performance bottlenecks.

To fit models into this memory budget, you must use quantized versions. Quantization is a compression technique that reduces the precision of model weights. The quant column shows the optimal format for each model. For example, a Q4_K_M quant uses approximately four bits per weight, while a Q6_K or Q8_0 quant uses six or eight bits. Higher quantization levels like Q8_0 preserve more output quality but require more memory space.

Several capable models fit entirely within the 4 GB limit of your hardware. The largest options include Lumina-Next or Lumina-Image 2.0 and CogVideoX 2B or 5B, which both run at the Q4_K_M quant using 3.7 GB of memory. Vision models like DeepSeek-VL2 and Phi-3.5-vision fit at Q5_K_M quants, using 3.8 GB and 3.6 GB respectively. Text models like Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, and MiniCPM 3 4B can run at the higher quality Q6_K quant, using 3.9 GB of video memory.

For audio and image generation, you can run SDXL Turbo or SDXL Lightning at the Q6_K quant using 3.4 GB of memory. Audio options like MusicGen small/medium/large fit at Q6_K using 3.2 GB. If you prefer higher precision, smaller models like SmolLM3 3B, Replit Code v1.5 3B, and Kandinsky 3.1 run at the Q8_0 quant, using 3.8 GB of video memory. These options run entirely on the graphics processor for maximum possible speed.

When a model exceeds 4 GB, you must use CPU offload to share the workload with your system RAM. Assuming your computer has 32 GB of system RAM, you can run larger models by splitting the weights. For instance, Mistral 7B requires 5.7 GB at the Q4_K_M quant, which uses 7.7 GB of system RAM. Similarly, Qwen2.5 7B, OLMo 2 7B, Falcon 3 7B, Command R7B, and OpenHermes 2.5 require 5.1 GB at Q4_K_M, using 7.1 GB of system RAM. Offloading allows you to run these larger models, but it significantly reduces generation speed because data must travel over the slower system bus.

You must also consider the memory cost of context length. The memory figures listed here are calculated using a standard 4k context window. If you increase the context window to process longer documents or chat histories, the key-value cache will consume additional memory. On a 4 GB card, expanding the context beyond 4k will quickly exceed your video memory and force the system to slow down.