Best local AI models for AMD R9 FURY X

4 GB HBM. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Lumina-Next / Lumina-Image 2.05BQ4_K_M3.7 GB
CogVideoX 2B / 5B5BQ4_K_M3.7 GB
DeepSeek-VL24.5BQ5_K_M3.8 GB
DeepFloyd IF4.3BQ5_K_M3.7 GB
Phi-3.5-vision4.2BQ5_K_M3.6 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-4-mini-instruct3.8BQ6_K3.7 GB
Phi-3.5 Mini3.8BQ6_K3.7 GB
OmniGen / OmniGen23.8BQ6_K3.7 GB
SD Cascade (Würstchen v3)3.6BQ6_K3.5 GB
SDXL Turbo3.5BQ6_K3.4 GB
SDXL Lightning3.5BQ6_K3.4 GB
ACE-Step3.5BQ6_K3.4 GB
MusicGen small/medium/large3.3BQ6_K3.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB
Open-Sora Plan2.7BQ8_03.4 GB
LFM2 1.2B / 2.6B2.6BQ8_03.3 GB
Playground v2.52.6BQ8_03.3 GB
Stable Diffusion 3.5 Medium2.5BQ8_03.2 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB

How to read this

The AMD Radeon R9 FURY X features 4 GB of High Bandwidth Memory. This memory capacity dictates the size of the artificial intelligence models you can run locally. To fit inside this 4 GB frame buffer, models must be compressed using quantization. Quantization reduces the precision of model weights to save space. The best quantization level column shows the highest quality format that fits entirely within your video memory.

For fully local execution, several model types fit within the 4 GB limit. Lumina-Next and Lumina-Image 2.0 at 5B parameters run at the Q4_K_M quantization using 3.7 GB of memory. CogVideoX 2B and 5B also fits at Q4_K_M using 3.7 GB. DeepSeek-VL2 at 4.5B parameters fits at Q5_K_M using 3.8 GB. DeepFloyd IF at 4.3B parameters fits at Q5_K_M using 3.7 GB. Phi-3.5-vision at 4.2B parameters fits at Q5_K_M using 3.6 GB.

Smaller models can run at higher precision levels. Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, Danube 3 4B, and Fish Speech 1.5 or OpenAudio S1 all fit at the Q6_K quantization using 3.9 GB of memory. Phi-4-mini-instruct, Phi-3.5 Mini, and OmniGen or OmniGen2 use 3.7 GB at Q6_K. SD Cascade uses 3.5 GB at Q6_K. SDXL Turbo, SDXL Lightning, and ACE-Step use 3.4 GB at Q6_K. MusicGen small, medium, or large uses 3.2 GB at Q6_K.

Models at 3B parameters or fewer can run at the Q8_0 quantization level. SmolLM3 3B, Replit Code v1.5 3B, Kandinsky 3.1, Voxtral Mini or Small, Orpheus TTS, and Higgs Audio v2 use 3.8 GB of memory. Allegro uses 3.6 GB at Q8_0. Open-Sora Plan uses 3.4 GB at Q8_0. LFM2 1.2B or 2.6B and Playground v2.5 use 3.3 GB at Q8_0. Stable Diffusion 3.5 Medium uses 3.2 GB at Q8_0.

When a model exceeds the 4 GB video memory limit, you must offload layers to your system RAM. This offloading process allows you to run larger models but reduces processing speed. For these setups, we assume a system with 32 GB of system RAM. Stable Diffusion XL requires 4.1 GB at FP8 or optimized settings and uses 6.1 GB of system RAM. Phi-3 Mini requires 4.4 GB at Q4_K_M and uses 6.4 GB of system RAM. Phi-4-multimodal requires 4.1 GB at Q4_K_M and uses 6.1 GB of system RAM.

Larger models require even more system RAM offloading. Magicoder-S-DS 6.7B requires 4.9 GB at Q4_K_M and uses 6.9 GB of system RAM. Mistral 7B requires 5.7 GB at Q4_K_M and uses 7.7 GB of system RAM. Qwen2.5 0.5B, 1.5B, 3B, or 7B, OLMo 2 1B or 7B, Falcon 3 1B, 3B, or 7B, Command R7B, and OpenHermes 2.5 all require 5.1 GB at Q4_K_M and use 7.1 GB of system RAM.

Be aware of the memory cost of context length. The listed memory usage figures are calculated using a standard 4k context window. If you increase the context length to process longer documents or conversations, the memory usage will rise. This extra memory demand can exceed the 4 GB limit of your card and force the system to slow down.