Best local AI models for AMD R7 M350

4 GB DDR3. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Lumina-Next / Lumina-Image 2.05BQ4_K_M3.7 GB
CogVideoX 2B / 5B5BQ4_K_M3.7 GB
DeepSeek-VL24.5BQ5_K_M3.8 GB
DeepFloyd IF4.3BQ5_K_M3.7 GB
Phi-3.5-vision4.2BQ5_K_M3.6 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-4-mini-instruct3.8BQ6_K3.7 GB
Phi-3.5 Mini3.8BQ6_K3.7 GB
OmniGen / OmniGen23.8BQ6_K3.7 GB
SD Cascade (Würstchen v3)3.6BQ6_K3.5 GB
SDXL Turbo3.5BQ6_K3.4 GB
SDXL Lightning3.5BQ6_K3.4 GB
ACE-Step3.5BQ6_K3.4 GB
MusicGen small/medium/large3.3BQ6_K3.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB
Open-Sora Plan2.7BQ8_03.4 GB
LFM2 1.2B / 2.6B2.6BQ8_03.3 GB
Playground v2.52.6BQ8_03.3 GB
Stable Diffusion 3.5 Medium2.5BQ8_03.2 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB

How to read this

The AMD Radeon R7 M350 is an entry level mobile graphics card equipped with 4 GB of DDR3 dedicated video memory. This specific memory capacity dictates the maximum size of the artificial intelligence models you can run locally. Because DDR3 memory has lower bandwidth than modern GDDR graphics memory, selecting the correct model size and quantization level is critical to achieve acceptable generation speeds.

The quantization column indicates the compression level applied to the model weights. Running a model at a Q4_K_M or Q5_K_M quantization reduces the memory footprint significantly compared to the original 16 bit floating point format. For example, Lumina-Next and CogVideoX 2B or 5B fit within 3.7 GB of video memory using a Q4_K_M quantization. Models like DeepSeek-VL2 and DeepFloyd IF use a Q5_K_M quantization to fit within 3.8 GB and 3.7 GB of video memory respectively.

Smaller models can run at higher precision levels because they require less physical space. You can run Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, Danube 3 4B, and Fish Speech 1.5 or OpenAudio S1 at a Q6_K quantization which uses 3.9 GB of video memory. Similarly, Phi-4-mini-instruct, Phi-3.5 Mini, and OmniGen or OmniGen2 fit within 3.7 GB using Q6_K. Image generators like SD Cascade, SDXL Turbo, and SDXL Lightning also run well at Q6_K using between 3.4 GB and 3.5 GB of video memory.

For models under 3 billion parameters, you can use the Q8_0 quantization level for better output quality. SmolLM3 3B, Replit Code v1.5 3B, Kandinsky 3.1, Voxtral Mini or Small, Orpheus TTS, and Higgs Audio v2 all fit within 3.8 GB using Q8_0. Very small models like Allegro, Open-Sora Plan, LFM2 1.2B or 2.6B, Playground v2.5, and Stable Diffusion 3.5 Medium fit comfortably within 3.2 GB to 3.6 GB of video memory using the Q8_0 quantization.

When a model exceeds the 4 GB video memory limit, you must use CPU offloading. This process splits the model layers between your graphics card and your system RAM. For these cases, we assume a system with 32 GB of system RAM. Offloading allows you to run larger models like Mistral 7B, Qwen2.5 7B, OLMo 2 7B, Falcon 3 7B, Command R7B, and OpenHermes 2.5, which require 5.1 GB to 5.7 GB of memory at Q4_K_M and up to 7.7 GB of system RAM. However, offloading over the slow system bus will drastically reduce your generation speed.

You must also consider the context window size when calculating memory usage. The memory figures listed here assume a standard 4k context window. If you increase the context length to process longer documents or chat histories, the key value cache will consume more video memory. This extra memory usage can easily push a model over the 4 GB limit of your AMD R7 M350 and force slow CPU offloading.