Best local AI models for AMD FirePro W7170M

4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Lumina-Next / Lumina-Image 2.05BQ4_K_M3.7 GB
CogVideoX 2B / 5B5BQ4_K_M3.7 GB
DeepSeek-VL24.5BQ5_K_M3.8 GB
DeepFloyd IF4.3BQ5_K_M3.7 GB
Phi-3.5-vision4.2BQ5_K_M3.6 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-4-mini-instruct3.8BQ6_K3.7 GB
Phi-3.5 Mini3.8BQ6_K3.7 GB
OmniGen / OmniGen23.8BQ6_K3.7 GB
SD Cascade (Würstchen v3)3.6BQ6_K3.5 GB
SDXL Turbo3.5BQ6_K3.4 GB
SDXL Lightning3.5BQ6_K3.4 GB
ACE-Step3.5BQ6_K3.4 GB
MusicGen small/medium/large3.3BQ6_K3.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB
Open-Sora Plan2.7BQ8_03.4 GB
LFM2 1.2B / 2.6B2.6BQ8_03.3 GB
Playground v2.52.6BQ8_03.3 GB
Stable Diffusion 3.5 Medium2.5BQ8_03.2 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB

How to read this

The AMD FirePro W7170M is a mobile workstation graphics card equipped with 4 GB of GDDR5 memory. This dedicated video memory determines the maximum size of the artificial intelligence models you can run entirely on the hardware. When a model fits completely within this 4 GB frame, it runs at the maximum speed supported by your GPU.

To fit larger models into this memory space, developers use quantization. The quant column shows the compression level applied to each model. For example, the 5B Lumina-Next model fits using a Q4_K_M quantization which takes 3.7 GB of memory. Smaller models like the 4B Qwen3 can run at a higher quality Q6_K quantization while using 3.9 GB of memory. The 3B SmolLM3 can run at an even higher quality Q8_0 quantization using 3.8 GB of memory.

If you want to run models that exceed the 4 GB limit, you must use CPU offloading. This process splits the model between your graphics card and your system RAM. We assume your system has 32 GB of system RAM for these calculations. Offloading allows you to run larger models like Mistral 7B at Q4_K_M which needs 5.7 GB of memory and uses 7.7 GB of system RAM. This method makes larger models accessible but it reduces processing speed because system RAM is slower than GDDR5 memory.

Other models also benefit from CPU offloading on this hardware. You can run the 5.6B Phi-4-multimodal model at Q4_K_M using 4.1 GB of memory and 6.1 GB of system RAM. The Magicoder-S-DS 6.7B model runs at Q4_K_M using 4.9 GB of memory and 6.9 GB of system RAM. Popular 7B models like Falcon 3 and Command R7B run at Q4_K_M using 5.1 GB of memory and 7.1 GB of system RAM.

When running these models, you must consider the context window size. The memory calculations shown here are based on a standard 4k context window. If you increase the context window to process longer documents or chat histories, the model will require more memory. This extra memory demand might force you to use a lower quantization level or offload more layers to your system RAM.