Best local AI models for AMD FirePro W6150M

4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Lumina-Next / Lumina-Image 2.05BQ4_K_M3.7 GB
CogVideoX 2B / 5B5BQ4_K_M3.7 GB
DeepSeek-VL24.5BQ5_K_M3.8 GB
DeepFloyd IF4.3BQ5_K_M3.7 GB
Phi-3.5-vision4.2BQ5_K_M3.6 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-4-mini-instruct3.8BQ6_K3.7 GB
Phi-3.5 Mini3.8BQ6_K3.7 GB
OmniGen / OmniGen23.8BQ6_K3.7 GB
SD Cascade (Würstchen v3)3.6BQ6_K3.5 GB
SDXL Turbo3.5BQ6_K3.4 GB
SDXL Lightning3.5BQ6_K3.4 GB
ACE-Step3.5BQ6_K3.4 GB
MusicGen small/medium/large3.3BQ6_K3.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB
Open-Sora Plan2.7BQ8_03.4 GB
LFM2 1.2B / 2.6B2.6BQ8_03.3 GB
Playground v2.52.6BQ8_03.3 GB
Stable Diffusion 3.5 Medium2.5BQ8_03.2 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB

How to read this

The AMD FirePro W6150M is a mobile workstation graphics card equipped with 4 GB of GDDR5 memory. This dedicated video memory determines the maximum size of the artificial intelligence models you can run entirely on the hardware. To load and run a model successfully on this GPU, the model files and the active working memory must fit within this 4 GB limit.

The quantization level represents the compression format of the model weights. Choosing a lower quantization like Q4_K_M or Q5_K_M reduces the memory footprint of larger models so they fit on your hardware. For example, you can run the 5B Lumina-Next or CogVideoX 2B / 5B models at Q4_K_M quantization using 3.7 GB of video memory. Similarly, DeepSeek-VL2 at 4.5B fits at Q5_K_M quantization using 3.8 GB of memory.

For models with 4B parameters, you can often use a higher Q6_K quantization. Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, Danube 3 4B, and Fish Speech 1.5 / OpenAudio S1 all run at Q6_K quantization using 3.9 GB of video memory. The 3.8B models like Phi-4-mini-instruct, Phi-3.5 Mini, and OmniGen / OmniGen2 also run at Q6_K quantization using 3.7 GB of memory.

Models under 3B parameters can run at the high quality Q8_0 quantization. SmolLM3 3B, Replit Code v1.5 3B, Kandinsky 3.1, Voxtral Mini / Small, Orpheus TTS, and Higgs Audio v2 use 3.8 GB of video memory at Q8_0. Stable Diffusion 3.5 Medium at 2.5B parameters fits within 3.2 GB of video memory at Q8_0 quantization.

When a model exceeds the 4 GB video memory limit, you must use CPU offload. This process splits the model layers between your GPU and your system RAM. For instance, running Mistral 7B at Q4_K_M requires 5.7 GB of total memory, which uses your GPU and 7.7 GB of system RAM. Running Stable Diffusion XL at FP8 requires 4.1 GB of memory, which uses your GPU and 6.1 GB of system RAM.

Be aware that active context windows consume extra memory during inference. The listed memory usage figures assume a standard 4k context window. If you increase the context length to process longer documents, the memory requirements will rise. This can push a model past the 4 GB limit and trigger slow system RAM offloading.