Best local AI models for AMD R9 M365X

4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Lumina-Next / Lumina-Image 2.05BQ4_K_M3.7 GB
CogVideoX 2B / 5B5BQ4_K_M3.7 GB
DeepSeek-VL24.5BQ5_K_M3.8 GB
DeepFloyd IF4.3BQ5_K_M3.7 GB
Phi-3.5-vision4.2BQ5_K_M3.6 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-4-mini-instruct3.8BQ6_K3.7 GB
Phi-3.5 Mini3.8BQ6_K3.7 GB
OmniGen / OmniGen23.8BQ6_K3.7 GB
SD Cascade (Würstchen v3)3.6BQ6_K3.5 GB
SDXL Turbo3.5BQ6_K3.4 GB
SDXL Lightning3.5BQ6_K3.4 GB
ACE-Step3.5BQ6_K3.4 GB
MusicGen small/medium/large3.3BQ6_K3.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB
Open-Sora Plan2.7BQ8_03.4 GB
LFM2 1.2B / 2.6B2.6BQ8_03.3 GB
Playground v2.52.6BQ8_03.3 GB
Stable Diffusion 3.5 Medium2.5BQ8_03.2 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB

How to read this

The AMD Radeon R9 M365X graphics card features 4 GB of GDDR5 video memory. This dedicated memory size determines which artificial intelligence models can run entirely on your hardware. For local execution, the model files must fit within this 4 GB limit to maintain acceptable processing speeds. If a model exceeds this limit, your system must use slower system memory to process the remaining data.

The quantization column indicates the compression level used to shrink these models. Quantization formats like Q4_K_M, Q5_K_M, Q6_K, and Q8_0 reduce the precision of model weights to save space. A lower number like Q4_K_M uses less video memory but reduces output quality. A higher number like Q8_0 preserves more original quality but requires much more of your 4 GB video memory.

Several capable models fit completely within the 4 GB video memory limit of your AMD R9 M365X. The Lumina-Next or Lumina-Image 2.0 model at 5B parameters fits using a Q4_K_M quantization which consumes 3.7 GB of memory. The CogVideoX 5B model also fits at Q4_K_M quantization using 3.7 GB of memory. For vision tasks, DeepSeek-VL2 at 4.5B parameters fits with a Q5_K_M quantization using 3.8 GB of memory. The Phi-3.5-vision model at 4.2B parameters fits at Q5_K_M quantization using 3.6 GB of memory.

Text and audio models also fit within the local video memory. The Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, and Danube 3 4B models all fit at Q6_K quantization using 3.9 GB of memory. For audio tasks, Fish Speech 1.5 or OpenAudio S1 at 4B parameters fits at Q6_K quantization using 3.9 GB of memory. The Phi-4-mini-instruct and Phi-3.5 Mini models at 3.8B parameters fit at Q6_K quantization using 3.7 GB of memory. The SmolLM3 3B and Replit Code v1.5 3B models fit at Q8_0 quantization using 3.8 GB of memory.

When a model is too large for your 4 GB video memory, you can offload parts of it to your system RAM. This CPU offload process assumes you have 32 GB of system RAM. For example, Mistral 7B at Q4_K_M quantization needs 5.7 GB of video memory and 7.7 GB of system RAM. The Qwen2.5 7B, OLMo 2 7B, Falcon 3 7B, Command R7B, and OpenHermes 2.5 models all need 5.1 GB of video memory and 7.1 GB of system RAM at Q4_K_M quantization. This offload method allows you to run larger models but reduces generation speed.

Running these models at their maximum context window will increase memory usage. The memory figures listed are calculated using a standard 4k context window. If you increase the context length to process longer documents, the system will require more memory. This extra demand can exceed your 4 GB limit and force the system to slow down.