Best local AI models for AMD R9 295X2
4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Lumina-Next / Lumina-Image 2.0 | 5B | Q4_K_M | 3.7 GB |
| CogVideoX 2B / 5B | 5B | Q4_K_M | 3.7 GB |
| DeepSeek-VL2 | 4.5B | Q5_K_M | 3.8 GB |
| DeepFloyd IF | 4.3B | Q5_K_M | 3.7 GB |
| Phi-3.5-vision | 4.2B | Q5_K_M | 3.6 GB |
| Qwen3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 4 E4B | 4B | Q6_K | 3.9 GB |
| MiniCPM 3 4B | 4B | Q6_K | 3.9 GB |
| Danube 3 4B | 4B | Q6_K | 3.9 GB |
| Fish Speech 1.5 / OpenAudio S1 | 4B | Q6_K | 3.9 GB |
| Phi-4-mini-instruct | 3.8B | Q6_K | 3.7 GB |
| Phi-3.5 Mini | 3.8B | Q6_K | 3.7 GB |
| OmniGen / OmniGen2 | 3.8B | Q6_K | 3.7 GB |
| SD Cascade (Würstchen v3) | 3.6B | Q6_K | 3.5 GB |
| SDXL Turbo | 3.5B | Q6_K | 3.4 GB |
| SDXL Lightning | 3.5B | Q6_K | 3.4 GB |
| ACE-Step | 3.5B | Q6_K | 3.4 GB |
| MusicGen small/medium/large | 3.3B | Q6_K | 3.2 GB |
| SmolLM3 3B | 3B | Q8_0 | 3.8 GB |
| Replit Code v1.5 3B | 3B | Q8_0 | 3.8 GB |
| Kandinsky 3.1 | 3B | Q8_0 | 3.8 GB |
| Voxtral Mini / Small | 3B | Q8_0 | 3.8 GB |
| Orpheus TTS | 3B | Q8_0 | 3.8 GB |
| Higgs Audio v2 | 3B | Q8_0 | 3.8 GB |
| Allegro | 2.8B | Q8_0 | 3.6 GB |
| Open-Sora Plan | 2.7B | Q8_0 | 3.4 GB |
| LFM2 1.2B / 2.6B | 2.6B | Q8_0 | 3.3 GB |
| Playground v2.5 | 2.6B | Q8_0 | 3.3 GB |
| Stable Diffusion 3.5 Medium | 2.5B | Q8_0 | 3.2 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| Stable Diffusion XL | 3.417B | 4.1 GB needed | 6.1 GB |
| Phi-3 Mini | 3.8B | 4.4 GB needed | 6.4 GB |
| Phi-4-multimodal | 5.6B | 4.1 GB needed | 6.1 GB |
| Magicoder-S-DS 6.7B | 6.7B | 4.9 GB needed | 6.9 GB |
| Mistral 7B | 7B | 5.7 GB needed | 7.7 GB |
| Qwen2.5 0.5B / 1.5B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| OLMo 2 1B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Falcon 3 1B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Command R7B | 7B | 5.1 GB needed | 7.1 GB |
| OpenHermes 2.5 | 7B | 5.1 GB needed | 7.1 GB |
How to read this
The AMD Radeon R9 295X2 is a dual GPU graphics card. Each individual GPU has access to 4 GB of GDDR5 memory. When running local AI models, your model must fit within the memory of a single GPU to run at full speed. The memory size column shows how much VRAM the model requires. If a model exceeds 4 GB, it cannot run entirely in the fast graphics memory.
The quantization column shows the best compression format for each model. Quantization reduces the size of the model weights. A Q4_K_M quant uses four bits per weight and offers a good balance of speed and quality. A Q6_K quant uses six bits per weight for higher precision. A Q8_0 quant uses eight bits per weight to deliver maximum accuracy within the 4 GB limit.
Several high quality models fit entirely within the 4 GB VRAM limit. The Lumina-Next and Lumina-Image 2.0 5B models fit using a Q4_K_M quant which uses 3.7 GB of memory. CogVideoX 5B also fits at Q4_K_M using 3.7 GB. For vision tasks, DeepSeek-VL2 4.5B fits at Q5_K_M using 3.8 GB. Phi-3.5-vision 4.2B fits at Q5_K_M using 3.6 GB. DeepFloyd IF 4.3B fits at Q5_K_M using 3.7 GB.
Many 4B and 3.8B models run well at Q6_K quantization. Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, Danube 3 4B, and Fish Speech 1.5 use 3.9 GB of memory. Phi-4-mini-instruct, Phi-3.5 Mini, and OmniGen use 3.7 GB. SD Cascade uses 3.5 GB. SDXL Turbo, SDXL Lightning, and ACE-Step use 3.4 GB. MusicGen uses 3.2 GB. At Q8_0 quantization, SmolLM3 3B, Replit Code v1.5 3B, Kandinsky 3.1, Voxtral Mini, Orpheus TTS, and Higgs Audio v2 use 3.8 GB. Allegro uses 3.6 GB. Open-Sora Plan uses 3.4 GB. LFM2 2.6B and Playground v2.5 use 3.3 GB. Stable Diffusion 3.5 Medium uses 3.2 GB.
When a model is too large for the 4 GB VRAM, you can offload parts of it to your system RAM. This offload process requires a system with 32 GB of system RAM. Offloading allows you to run larger models, but it reduces generation speed. Stable Diffusion XL needs 4.1 GB at FP8 and uses 6.1 GB of system RAM. Phi-3 Mini needs 4.4 GB at Q4_K_M and uses 6.4 GB of system RAM. Phi-4-multimodal needs 4.1 GB at Q4_K_M and uses 6.1 GB of system RAM. Magicoder-S-DS 6.7B needs 4.9 GB at Q4_K_M and uses 6.9 GB of system RAM.
Larger 7B models can also run via CPU offloading. Mistral 7B needs 5.7 GB at Q4_K_M and uses 7.7 GB of system RAM. Qwen2.5 7B, OLMo 2 7B, Falcon 3 7B, Command R7B, and OpenHermes 2.5 need 5.1 GB at Q4_K_M and use 7.1 GB of system RAM. Be aware that these memory figures are calculated using a standard 4k context window. If you increase the context window, the model will require more memory and may exceed your available VRAM.