Best local AI models for AMD HD 7990
3 GB GDDR5. At a 4k context, 76 of the 233 models in our catalog with verified parameter counts fit fully, up to Qwen3 4B at 4B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 76 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Qwen3 4B | 4B | Q4_K_M | 2.9 GB |
| Gemma 3 4B | 4B | Q4_K_M | 2.9 GB |
| Gemma 4 E4B | 4B | Q4_K_M | 2.9 GB |
| MiniCPM 3 4B | 4B | Q4_K_M | 2.9 GB |
| Danube 3 4B | 4B | Q4_K_M | 2.9 GB |
| Fish Speech 1.5 / OpenAudio S1 | 4B | Q4_K_M | 2.9 GB |
| Phi-4-mini-instruct | 3.8B | Q4_K_M | 2.8 GB |
| Phi-3.5 Mini | 3.8B | Q4_K_M | 2.8 GB |
| OmniGen / OmniGen2 | 3.8B | Q4_K_M | 2.8 GB |
| SD Cascade (Würstchen v3) | 3.6B | Q4_K_M | 2.6 GB |
| SDXL Turbo | 3.5B | Q5_K_M | 3 GB |
| SDXL Lightning | 3.5B | Q5_K_M | 3 GB |
| ACE-Step | 3.5B | Q5_K_M | 3 GB |
| MusicGen small/medium/large | 3.3B | Q5_K_M | 2.8 GB |
| SmolLM3 3B | 3B | Q6_K | 3 GB |
| Replit Code v1.5 3B | 3B | Q6_K | 3 GB |
| Kandinsky 3.1 | 3B | Q6_K | 3 GB |
| Voxtral Mini / Small | 3B | Q6_K | 3 GB |
| Orpheus TTS | 3B | Q6_K | 3 GB |
| Higgs Audio v2 | 3B | Q6_K | 3 GB |
| Allegro | 2.8B | Q6_K | 2.8 GB |
| Open-Sora Plan | 2.7B | Q6_K | 2.7 GB |
| LFM2 1.2B / 2.6B | 2.6B | Q6_K | 2.6 GB |
| Playground v2.5 | 2.6B | Q6_K | 2.6 GB |
| Stable Diffusion 3.5 Medium | 2.5B | Q6_K | 2.5 GB |
| Canary 1B / Qwen-2.5B | 2.5B | Q6_K | 2.5 GB |
| SeamlessM4T v2 | 2.3B | Q8_0 | 2.9 GB |
| Parler-TTS | 2.2B | Q8_0 | 2.8 GB |
| Kimi K3 DSpark | 2.2B | Q8_0 | 2.9 GB |
| SmolVLM 256M / 500M / 2B | 2B | Q8_0 | 2.5 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| Stable Diffusion XL | 3.417B | 4.1 GB needed | 6.1 GB |
| Phi-3 Mini | 3.8B | 4.4 GB needed | 6.4 GB |
| Phi-3.5-vision | 4.2B | 3.1 GB needed | 5.1 GB |
| DeepFloyd IF | 4.3B | 3.1 GB needed | 5.1 GB |
| DeepSeek-VL2 | 4.5B | 3.3 GB needed | 5.3 GB |
| Lumina-Next / Lumina-Image 2.0 | 5B | 3.7 GB needed | 5.7 GB |
| CogVideoX 2B / 5B | 5B | 3.7 GB needed | 5.7 GB |
| Phi-4-multimodal | 5.6B | 4.1 GB needed | 6.1 GB |
| Magicoder-S-DS 6.7B | 6.7B | 4.9 GB needed | 6.9 GB |
| Mistral 7B | 7B | 5.7 GB needed | 7.7 GB |
How to read this
The AMD HD 7990 graphics card features 3 GB of GDDR5 memory. This memory size determines which local AI models can run entirely on the hardware. When a model fits completely within this 3 GB limit, it runs at maximum speed because the graphics processor can access the model weights directly without waiting for the system memory.
The quant column shows the quantization level used to compress each model. Quantization reduces the precision of model weights to save space. For example, Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, Danube 3 4B, and Fish Speech 1.5 / OpenAudio S1 use a Q4_K_M quant to fit within 2.9 GB. Phi-4-mini-instruct, Phi-3.5 Mini, and OmniGen / OmniGen2 use the same Q4_K_M quant to fit within 2.8 GB. SD Cascade (Würstchen v3) fits at 2.6 GB using Q4_K_M.
Models like SDXL Turbo, SDXL Lightning, and ACE-Step use a Q5_K_M quant to fit exactly 3 GB of memory. MusicGen small/medium/large fits within 2.8 GB using Q5_K_M. Other models use a higher quality Q6_K quant to fit. These include SmolLM3 3B, Replit Code v1.5 3B, Kandinsky 3.1, Voxtral Mini / Small, Orpheus TTS, and Higgs Audio v2 at 3 GB. Allegro fits at 2.8 GB, Open-Sora Plan fits at 2.7 GB, LFM2 1.2B / 2.6B and Playground v2.5 fit at 2.6 GB, while Stable Diffusion 3.5 Medium and Canary 1B / Qwen-2.5B fit at 2.5 GB using Q6_K.
The highest precision quants on this hardware use Q8_0. SeamlessM4T v2 uses Q8_0 to fit within 2.9 GB. Parler-TTS fits within 2.8 GB, Kimi K3 DSpark fits within 2.9 GB, and SmolVLM 256M / 500M / 2B fits within 2.5 GB using Q8_0. Running models at these sizes requires careful attention to context length. A standard 4k context window requires extra memory for active tokens, which can easily exceed the 3 GB limit if the model size is already close to the maximum.
When a model is too large for the 3 GB graphics memory, you can offload parts of it to your system RAM. This offload process allows you to run larger models but reduces processing speed. For these cases, we assume your computer has 32 GB of system RAM. Stable Diffusion XL needs 4.1 GB at FP8 or optimized settings, which requires 6.1 GB of system RAM. Phi-3 Mini needs 4.4 GB at Q4_K_M, requiring 6.4 GB of system RAM.
Other offload options include Phi-3.5-vision and DeepFloyd IF, which both need 3.1 GB at Q4_K_M and require 5.1 GB of system RAM. DeepSeek-VL2 needs 3.3 GB at Q4_K_M and requires 5.3 GB of system RAM. Lumina-Next / Lumina-Image 2.0 and CogVideoX 2B / 5B both need 3.7 GB at Q4_K_M, requiring 5.7 GB of system RAM. Phi-4-multimodal needs 4.1 GB at Q4_K_M and requires 6.1 GB of system RAM.
The largest offload models require the most system memory. Magicoder-S-DS 6.7B needs 4.9 GB at Q4_K_M, which requires 6.9 GB of system RAM. Mistral 7B needs 5.7 GB at Q4_K_M, requiring 7.7 GB of system RAM. Offloading these models allows the AMD HD 7990 to assist with computation even though the model files exceed the onboard physical memory.