Best local AI models for AMD R9 M365X
4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Lumina-Next / Lumina-Image 2.0 | 5B | Q4_K_M | 3.7 GB |
| CogVideoX 2B / 5B | 5B | Q4_K_M | 3.7 GB |
| DeepSeek-VL2 | 4.5B | Q5_K_M | 3.8 GB |
| DeepFloyd IF | 4.3B | Q5_K_M | 3.7 GB |
| Phi-3.5-vision | 4.2B | Q5_K_M | 3.6 GB |
| Qwen3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 4 E4B | 4B | Q6_K | 3.9 GB |
| MiniCPM 3 4B | 4B | Q6_K | 3.9 GB |
| Danube 3 4B | 4B | Q6_K | 3.9 GB |
| Fish Speech 1.5 / OpenAudio S1 | 4B | Q6_K | 3.9 GB |
| Phi-4-mini-instruct | 3.8B | Q6_K | 3.7 GB |
| Phi-3.5 Mini | 3.8B | Q6_K | 3.7 GB |
| OmniGen / OmniGen2 | 3.8B | Q6_K | 3.7 GB |
| SD Cascade (Würstchen v3) | 3.6B | Q6_K | 3.5 GB |
| SDXL Turbo | 3.5B | Q6_K | 3.4 GB |
| SDXL Lightning | 3.5B | Q6_K | 3.4 GB |
| ACE-Step | 3.5B | Q6_K | 3.4 GB |
| MusicGen small/medium/large | 3.3B | Q6_K | 3.2 GB |
| SmolLM3 3B | 3B | Q8_0 | 3.8 GB |
| Replit Code v1.5 3B | 3B | Q8_0 | 3.8 GB |
| Kandinsky 3.1 | 3B | Q8_0 | 3.8 GB |
| Voxtral Mini / Small | 3B | Q8_0 | 3.8 GB |
| Orpheus TTS | 3B | Q8_0 | 3.8 GB |
| Higgs Audio v2 | 3B | Q8_0 | 3.8 GB |
| Allegro | 2.8B | Q8_0 | 3.6 GB |
| Open-Sora Plan | 2.7B | Q8_0 | 3.4 GB |
| LFM2 1.2B / 2.6B | 2.6B | Q8_0 | 3.3 GB |
| Playground v2.5 | 2.6B | Q8_0 | 3.3 GB |
| Stable Diffusion 3.5 Medium | 2.5B | Q8_0 | 3.2 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| Stable Diffusion XL | 3.417B | 4.1 GB needed | 6.1 GB |
| Phi-3 Mini | 3.8B | 4.4 GB needed | 6.4 GB |
| Phi-4-multimodal | 5.6B | 4.1 GB needed | 6.1 GB |
| Magicoder-S-DS 6.7B | 6.7B | 4.9 GB needed | 6.9 GB |
| Mistral 7B | 7B | 5.7 GB needed | 7.7 GB |
| Qwen2.5 0.5B / 1.5B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| OLMo 2 1B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Falcon 3 1B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Command R7B | 7B | 5.1 GB needed | 7.1 GB |
| OpenHermes 2.5 | 7B | 5.1 GB needed | 7.1 GB |
How to read this
The AMD Radeon R9 M365X graphics card features 4 GB of GDDR5 video memory. This dedicated memory size determines which artificial intelligence models can run entirely on your hardware. For local execution, the model files must fit within this 4 GB limit to maintain acceptable processing speeds. If a model exceeds this limit, your system must use slower system memory to process the remaining data.
The quantization column indicates the compression level used to shrink these models. Quantization formats like Q4_K_M, Q5_K_M, Q6_K, and Q8_0 reduce the precision of model weights to save space. A lower number like Q4_K_M uses less video memory but reduces output quality. A higher number like Q8_0 preserves more original quality but requires much more of your 4 GB video memory.
Several capable models fit completely within the 4 GB video memory limit of your AMD R9 M365X. The Lumina-Next or Lumina-Image 2.0 model at 5B parameters fits using a Q4_K_M quantization which consumes 3.7 GB of memory. The CogVideoX 5B model also fits at Q4_K_M quantization using 3.7 GB of memory. For vision tasks, DeepSeek-VL2 at 4.5B parameters fits with a Q5_K_M quantization using 3.8 GB of memory. The Phi-3.5-vision model at 4.2B parameters fits at Q5_K_M quantization using 3.6 GB of memory.
Text and audio models also fit within the local video memory. The Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, and Danube 3 4B models all fit at Q6_K quantization using 3.9 GB of memory. For audio tasks, Fish Speech 1.5 or OpenAudio S1 at 4B parameters fits at Q6_K quantization using 3.9 GB of memory. The Phi-4-mini-instruct and Phi-3.5 Mini models at 3.8B parameters fit at Q6_K quantization using 3.7 GB of memory. The SmolLM3 3B and Replit Code v1.5 3B models fit at Q8_0 quantization using 3.8 GB of memory.
When a model is too large for your 4 GB video memory, you can offload parts of it to your system RAM. This CPU offload process assumes you have 32 GB of system RAM. For example, Mistral 7B at Q4_K_M quantization needs 5.7 GB of video memory and 7.7 GB of system RAM. The Qwen2.5 7B, OLMo 2 7B, Falcon 3 7B, Command R7B, and OpenHermes 2.5 models all need 5.1 GB of video memory and 7.1 GB of system RAM at Q4_K_M quantization. This offload method allows you to run larger models but reduces generation speed.
Running these models at their maximum context window will increase memory usage. The memory figures listed are calculated using a standard 4k context window. If you increase the context length to process longer documents, the system will require more memory. This extra demand can exceed your 4 GB limit and force the system to slow down.