Best local AI models for AMD R9 FURY

4 GB HBM. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Lumina-Next / Lumina-Image 2.05BQ4_K_M3.7 GB
CogVideoX 2B / 5B5BQ4_K_M3.7 GB
DeepSeek-VL24.5BQ5_K_M3.8 GB
DeepFloyd IF4.3BQ5_K_M3.7 GB
Phi-3.5-vision4.2BQ5_K_M3.6 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-4-mini-instruct3.8BQ6_K3.7 GB
Phi-3.5 Mini3.8BQ6_K3.7 GB
OmniGen / OmniGen23.8BQ6_K3.7 GB
SD Cascade (Würstchen v3)3.6BQ6_K3.5 GB
SDXL Turbo3.5BQ6_K3.4 GB
SDXL Lightning3.5BQ6_K3.4 GB
ACE-Step3.5BQ6_K3.4 GB
MusicGen small/medium/large3.3BQ6_K3.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB
Open-Sora Plan2.7BQ8_03.4 GB
LFM2 1.2B / 2.6B2.6BQ8_03.3 GB
Playground v2.52.6BQ8_03.3 GB
Stable Diffusion 3.5 Medium2.5BQ8_03.2 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB

How to read this

The AMD R9 FURY graphics card features 4 GB of High Bandwidth Memory. This dedicated memory pool determines which local AI models can run entirely on the hardware. To run a model without slowdowns, the model files and the active memory must fit within this 4 GB limit.

The best quant column shows the optimal quantization level for each model. Quantization reduces the size of model weights to save space. For example, the 5B Lumina-Next model fits in 3.7 GB of memory using a Q4_K_M quantization. Models like Qwen3 4B and Phi-4-mini-instruct can run at a higher Q6_K quantization while using 3.9 GB and 3.7 GB of memory. Smaller models like SmolLM3 3B and Kandinsky 3.1 can run at Q8_0 quantization using 3.8 GB of memory.

Running models at these maximum quantization levels leaves very little free space. A major caveat is the 4k context window limit. Generating longer responses or processing large prompts increases memory usage. If your context exceeds this limit, the system will run out of memory and crash.

You can run larger models by offloading parts of the workload to your system RAM. This process requires a system with 32 GB of system RAM. Offloading allows you to run models that exceed the 4 GB physical limit of the card, but it reduces processing speed because system RAM is slower than graphics memory.

With CPU offloading, you can run Mistral 7B at Q4_K_M quantization. This setup needs 5.7 GB of graphics memory and 7.7 GB of system RAM. You can also run Qwen2.5 7B, OLMo 2 7B, Falcon 3 7B, Command R7B, and OpenHermes 2.5 at Q4_K_M quantization. These models require 5.1 GB of graphics memory and 7.1 GB of system RAM.

Other offload options include Magicoder-S-DS 6.7B which needs 4.9 GB of graphics memory and 6.9 GB of system RAM. Stable Diffusion XL needs 4.1 GB of graphics memory at FP8 or optimized settings along with 6.1 GB of system RAM. These offload configurations expand your options beyond the physical limits of the card.