Best local AI models for AMD Pro 460

4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Lumina-Next / Lumina-Image 2.05BQ4_K_M3.7 GB
CogVideoX 2B / 5B5BQ4_K_M3.7 GB
DeepSeek-VL24.5BQ5_K_M3.8 GB
DeepFloyd IF4.3BQ5_K_M3.7 GB
Phi-3.5-vision4.2BQ5_K_M3.6 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-4-mini-instruct3.8BQ6_K3.7 GB
Phi-3.5 Mini3.8BQ6_K3.7 GB
OmniGen / OmniGen23.8BQ6_K3.7 GB
SD Cascade (Würstchen v3)3.6BQ6_K3.5 GB
SDXL Turbo3.5BQ6_K3.4 GB
SDXL Lightning3.5BQ6_K3.4 GB
ACE-Step3.5BQ6_K3.4 GB
MusicGen small/medium/large3.3BQ6_K3.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB
Open-Sora Plan2.7BQ8_03.4 GB
LFM2 1.2B / 2.6B2.6BQ8_03.3 GB
Playground v2.52.6BQ8_03.3 GB
Stable Diffusion 3.5 Medium2.5BQ8_03.2 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB

How to read this

The AMD Radeon Pro 460 is a mobile graphics card equipped with 4 GB of GDDR5 dedicated memory. This hardware limit determines which local AI models you can run entirely on the GPU. When a model fits completely within this 4 GB frame, it benefits from the fastest processing speeds the hardware can offer. Keeping the model size under the available VRAM prevents severe performance drops.

To fit larger models into this memory space, we use quantized versions. Quantization reduces the precision of model weights to save space. The best quant column shows the highest quality quantization level that still fits within your VRAM. For example, a 5B model like Lumina-Next or CogVideoX 2B / 5B can run at Q4_K_M quantization using 3.7 GB of memory. Smaller models like SmolLM3 3B or Replit Code v1.5 3B can run at a higher quality Q8_0 quantization using 3.8 GB of memory.

Several vision and image models fit within the local limits of this card. DeepSeek-VL2 at 4.5B fits using a Q5_K_M quant with 3.8 GB used. Phi-3.5-vision at 4.2B fits using a Q5_K_M quant with 3.6 GB used. For image generation, SDXL Turbo and SDXL Lightning at 3.5B use 3.4 GB of memory at Q6_K quantization. Stable Diffusion 3.5 Medium at 2.5B fits using a Q8_0 quant with 3.2 GB used.

Text and audio models also run locally. Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, and Danube 3 4B all use 3.9 GB of memory at Q6_K quantization. Phi-4-mini-instruct and Phi-3.5 Mini at 3.8B use 3.7 GB of memory at Q6_K quantization. Audio models like Fish Speech 1.5 / OpenAudio S1 at 4B use 3.9 GB at Q6_K, while Orpheus TTS and Higgs Audio v2 at 3B use 3.8 GB at Q8_0 quantization.

When a model is too large for the 4 GB VRAM, you must offload parts of it to your system RAM. This CPU offload process allows you to run larger models but slows down generation speeds. For these cases, we assume a system with 32 GB of system RAM. Under this setup, Mistral 7B requires 5.7 GB of VRAM at Q4_K_M quantization and needs 7.7 GB of system RAM. Similarly, Qwen2.5 0.5B / 1.5B / 3B / 7B at the 7B size requires 5.1 GB of VRAM at Q4_K_M and needs 7.1 GB of system RAM.

Other offload options include Phi-4-multimodal at 5.6B, which needs 4.1 GB of VRAM at Q4_K_M and 6.1 GB of system RAM. Magicoder-S-DS 6.7B needs 4.9 GB of VRAM at Q4_K_M and 6.9 GB of system RAM. Popular 7B models like OLMo 2 1B / 7B, Falcon 3 1B / 3B / 7B, Command R7B, and OpenHermes 2.5 all require 5.1 GB of VRAM at Q4_K_M quantization and 7.1 GB of system RAM to run.

Be aware of the context window limit when running these models. The memory calculations shown here are based on a standard 4k context window. If you increase the context length to process longer documents or chat histories, the memory usage will rise. This extra memory demand can push a model past the 4 GB VRAM limit and trigger slow CPU offloading.