Best local AI models for AMD Pro 460
4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Lumina-Next / Lumina-Image 2.0 | 5B | Q4_K_M | 3.7 GB |
| CogVideoX 2B / 5B | 5B | Q4_K_M | 3.7 GB |
| DeepSeek-VL2 | 4.5B | Q5_K_M | 3.8 GB |
| DeepFloyd IF | 4.3B | Q5_K_M | 3.7 GB |
| Phi-3.5-vision | 4.2B | Q5_K_M | 3.6 GB |
| Qwen3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 4 E4B | 4B | Q6_K | 3.9 GB |
| MiniCPM 3 4B | 4B | Q6_K | 3.9 GB |
| Danube 3 4B | 4B | Q6_K | 3.9 GB |
| Fish Speech 1.5 / OpenAudio S1 | 4B | Q6_K | 3.9 GB |
| Phi-4-mini-instruct | 3.8B | Q6_K | 3.7 GB |
| Phi-3.5 Mini | 3.8B | Q6_K | 3.7 GB |
| OmniGen / OmniGen2 | 3.8B | Q6_K | 3.7 GB |
| SD Cascade (Würstchen v3) | 3.6B | Q6_K | 3.5 GB |
| SDXL Turbo | 3.5B | Q6_K | 3.4 GB |
| SDXL Lightning | 3.5B | Q6_K | 3.4 GB |
| ACE-Step | 3.5B | Q6_K | 3.4 GB |
| MusicGen small/medium/large | 3.3B | Q6_K | 3.2 GB |
| SmolLM3 3B | 3B | Q8_0 | 3.8 GB |
| Replit Code v1.5 3B | 3B | Q8_0 | 3.8 GB |
| Kandinsky 3.1 | 3B | Q8_0 | 3.8 GB |
| Voxtral Mini / Small | 3B | Q8_0 | 3.8 GB |
| Orpheus TTS | 3B | Q8_0 | 3.8 GB |
| Higgs Audio v2 | 3B | Q8_0 | 3.8 GB |
| Allegro | 2.8B | Q8_0 | 3.6 GB |
| Open-Sora Plan | 2.7B | Q8_0 | 3.4 GB |
| LFM2 1.2B / 2.6B | 2.6B | Q8_0 | 3.3 GB |
| Playground v2.5 | 2.6B | Q8_0 | 3.3 GB |
| Stable Diffusion 3.5 Medium | 2.5B | Q8_0 | 3.2 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| Stable Diffusion XL | 3.417B | 4.1 GB needed | 6.1 GB |
| Phi-3 Mini | 3.8B | 4.4 GB needed | 6.4 GB |
| Phi-4-multimodal | 5.6B | 4.1 GB needed | 6.1 GB |
| Magicoder-S-DS 6.7B | 6.7B | 4.9 GB needed | 6.9 GB |
| Mistral 7B | 7B | 5.7 GB needed | 7.7 GB |
| Qwen2.5 0.5B / 1.5B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| OLMo 2 1B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Falcon 3 1B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Command R7B | 7B | 5.1 GB needed | 7.1 GB |
| OpenHermes 2.5 | 7B | 5.1 GB needed | 7.1 GB |
How to read this
The AMD Radeon Pro 460 is a mobile graphics card equipped with 4 GB of GDDR5 dedicated memory. This hardware limit determines which local AI models you can run entirely on the GPU. When a model fits completely within this 4 GB frame, it benefits from the fastest processing speeds the hardware can offer. Keeping the model size under the available VRAM prevents severe performance drops.
To fit larger models into this memory space, we use quantized versions. Quantization reduces the precision of model weights to save space. The best quant column shows the highest quality quantization level that still fits within your VRAM. For example, a 5B model like Lumina-Next or CogVideoX 2B / 5B can run at Q4_K_M quantization using 3.7 GB of memory. Smaller models like SmolLM3 3B or Replit Code v1.5 3B can run at a higher quality Q8_0 quantization using 3.8 GB of memory.
Several vision and image models fit within the local limits of this card. DeepSeek-VL2 at 4.5B fits using a Q5_K_M quant with 3.8 GB used. Phi-3.5-vision at 4.2B fits using a Q5_K_M quant with 3.6 GB used. For image generation, SDXL Turbo and SDXL Lightning at 3.5B use 3.4 GB of memory at Q6_K quantization. Stable Diffusion 3.5 Medium at 2.5B fits using a Q8_0 quant with 3.2 GB used.
Text and audio models also run locally. Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, and Danube 3 4B all use 3.9 GB of memory at Q6_K quantization. Phi-4-mini-instruct and Phi-3.5 Mini at 3.8B use 3.7 GB of memory at Q6_K quantization. Audio models like Fish Speech 1.5 / OpenAudio S1 at 4B use 3.9 GB at Q6_K, while Orpheus TTS and Higgs Audio v2 at 3B use 3.8 GB at Q8_0 quantization.
When a model is too large for the 4 GB VRAM, you must offload parts of it to your system RAM. This CPU offload process allows you to run larger models but slows down generation speeds. For these cases, we assume a system with 32 GB of system RAM. Under this setup, Mistral 7B requires 5.7 GB of VRAM at Q4_K_M quantization and needs 7.7 GB of system RAM. Similarly, Qwen2.5 0.5B / 1.5B / 3B / 7B at the 7B size requires 5.1 GB of VRAM at Q4_K_M and needs 7.1 GB of system RAM.
Other offload options include Phi-4-multimodal at 5.6B, which needs 4.1 GB of VRAM at Q4_K_M and 6.1 GB of system RAM. Magicoder-S-DS 6.7B needs 4.9 GB of VRAM at Q4_K_M and 6.9 GB of system RAM. Popular 7B models like OLMo 2 1B / 7B, Falcon 3 1B / 3B / 7B, Command R7B, and OpenHermes 2.5 all require 5.1 GB of VRAM at Q4_K_M quantization and 7.1 GB of system RAM to run.
Be aware of the context window limit when running these models. The memory calculations shown here are based on a standard 4k context window. If you increase the context length to process longer documents or chat histories, the memory usage will rise. This extra memory demand can push a model past the 4 GB VRAM limit and trigger slow CPU offloading.