Best local AI models for AMD HD 8970M

4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Lumina-Next / Lumina-Image 2.05BQ4_K_M3.7 GB
CogVideoX 2B / 5B5BQ4_K_M3.7 GB
DeepSeek-VL24.5BQ5_K_M3.8 GB
DeepFloyd IF4.3BQ5_K_M3.7 GB
Phi-3.5-vision4.2BQ5_K_M3.6 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-4-mini-instruct3.8BQ6_K3.7 GB
Phi-3.5 Mini3.8BQ6_K3.7 GB
OmniGen / OmniGen23.8BQ6_K3.7 GB
SD Cascade (Würstchen v3)3.6BQ6_K3.5 GB
SDXL Turbo3.5BQ6_K3.4 GB
SDXL Lightning3.5BQ6_K3.4 GB
ACE-Step3.5BQ6_K3.4 GB
MusicGen small/medium/large3.3BQ6_K3.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB
Open-Sora Plan2.7BQ8_03.4 GB
LFM2 1.2B / 2.6B2.6BQ8_03.3 GB
Playground v2.52.6BQ8_03.3 GB
Stable Diffusion 3.5 Medium2.5BQ8_03.2 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB

How to read this

The AMD HD 8970M is a legacy mobile graphics card equipped with 4 GB of GDDR5 memory. This onboard memory capacity determines the maximum size of the artificial intelligence models you can run entirely on the hardware. When a model runs fully inside the graphics memory, it achieves the fastest possible processing speeds. If a model exceeds this limit, it cannot load or must rely on system memory sharing.

To fit larger models into the 4 GB limit, developers use quantization. The quant column shows the compression level applied to the weights of each model. For example, a Q4_K_M quant uses approximately four bits per weight, while Q6_K and Q8_0 quants offer higher precision at the cost of larger file sizes. Choosing the best quant balances the intelligence of the model against the strict physical memory limit of your hardware.

Several highly capable models fit entirely within the graphics memory. The Lumina-Next and Lumina-Image 2.0 models at 5B parameters can run using the Q4_K_M quant, consuming 3.7 GB of memory. The DeepSeek-VL2 model at 4.5B parameters fits using the Q5_K_M quant at 3.8 GB. For text and reasoning, the Phi-4-mini-instruct and Phi-3.5 Mini models at 3.8B parameters run efficiently using the Q6_K quant, which uses 3.7 GB of memory.

You can also run larger models by offloading parts of the workload to your system RAM. Assuming your computer has 32 GB of system RAM, you can run models that exceed 4 GB. For instance, Mistral 7B requires 5.7 GB of memory at the Q4_K_M quant, which uses 7.7 GB of system RAM. Similarly, the Qwen2.5 7B, OLMo 2 7B, Falcon 3 7B, Command R7B, and OpenHermes 2.5 models require 5.1 GB of memory at the Q4_K_M quant, which uses 7.1 GB of system RAM.

CPU offloading allows you to run these larger models, but it comes with a performance cost. Transferring data between the system RAM and the graphics card is much slower than using the onboard GDDR5 memory. This transfer bottleneck significantly reduces the generation speed. Additionally, you must consider the 4k context caveat. Running models with long context windows increases memory consumption during active use, which can quickly exceed your limits.