Best local AI models for NVIDIA GTX 1050 Ti
4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Lumina-Next / Lumina-Image 2.0 | 5B | Q4_K_M | 3.7 GB |
| CogVideoX 2B / 5B | 5B | Q4_K_M | 3.7 GB |
| DeepSeek-VL2 | 4.5B | Q5_K_M | 3.8 GB |
| DeepFloyd IF | 4.3B | Q5_K_M | 3.7 GB |
| Phi-3.5-vision | 4.2B | Q5_K_M | 3.6 GB |
| Qwen3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 4 E4B | 4B | Q6_K | 3.9 GB |
| MiniCPM 3 4B | 4B | Q6_K | 3.9 GB |
| Danube 3 4B | 4B | Q6_K | 3.9 GB |
| Fish Speech 1.5 / OpenAudio S1 | 4B | Q6_K | 3.9 GB |
| Phi-4-mini-instruct | 3.8B | Q6_K | 3.7 GB |
| Phi-3.5 Mini | 3.8B | Q6_K | 3.7 GB |
| OmniGen / OmniGen2 | 3.8B | Q6_K | 3.7 GB |
| SD Cascade (Würstchen v3) | 3.6B | Q6_K | 3.5 GB |
| SDXL Turbo | 3.5B | Q6_K | 3.4 GB |
| SDXL Lightning | 3.5B | Q6_K | 3.4 GB |
| ACE-Step | 3.5B | Q6_K | 3.4 GB |
| MusicGen small/medium/large | 3.3B | Q6_K | 3.2 GB |
| SmolLM3 3B | 3B | Q8_0 | 3.8 GB |
| Replit Code v1.5 3B | 3B | Q8_0 | 3.8 GB |
| Kandinsky 3.1 | 3B | Q8_0 | 3.8 GB |
| Voxtral Mini / Small | 3B | Q8_0 | 3.8 GB |
| Orpheus TTS | 3B | Q8_0 | 3.8 GB |
| Higgs Audio v2 | 3B | Q8_0 | 3.8 GB |
| Allegro | 2.8B | Q8_0 | 3.6 GB |
| Open-Sora Plan | 2.7B | Q8_0 | 3.4 GB |
| LFM2 1.2B / 2.6B | 2.6B | Q8_0 | 3.3 GB |
| Playground v2.5 | 2.6B | Q8_0 | 3.3 GB |
| Stable Diffusion 3.5 Medium | 2.5B | Q8_0 | 3.2 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| Stable Diffusion XL | 3.417B | 4.1 GB needed | 6.1 GB |
| Phi-3 Mini | 3.8B | 4.4 GB needed | 6.4 GB |
| Phi-4-multimodal | 5.6B | 4.1 GB needed | 6.1 GB |
| Magicoder-S-DS 6.7B | 6.7B | 4.9 GB needed | 6.9 GB |
| Mistral 7B | 7B | 5.7 GB needed | 7.7 GB |
| Qwen2.5 0.5B / 1.5B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| OLMo 2 1B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Falcon 3 1B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Command R7B | 7B | 5.1 GB needed | 7.1 GB |
| OpenHermes 2.5 | 7B | 5.1 GB needed | 7.1 GB |
How to read this
The NVIDIA GTX 1050 Ti features 4 GB of GDDR5 memory. This physical limit determines which local AI models can run entirely on your graphics card. To fit within this hardware boundary, models must use quantization. Quantization is a compression method that reduces the precision of model weights. This process lowers the memory footprint of a model while preserving most of its original capabilities.
The best quant column shows the highest quality quantization level that fits within the available VRAM. For example, the Lumina-Next or Lumina-Image 2.0 5B model fits at a Q4_K_M quantization using 3.7 GB of VRAM. Similarly, the DeepSeek-VL2 4.5B model fits at Q5_K_M using 3.8 GB. Smaller models like the Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, Danube 3 4B, and Fish Speech 1.5 or OpenAudio S1 4B models can run at a higher Q6_K quantization using 3.9 GB of VRAM.
Other models also fit comfortably within the 4 GB limit of your card. The Phi-4-mini-instruct 3.8B, Phi-3.5 Mini 3.8B, and OmniGen or OmniGen2 3.8B models use 3.7 GB of VRAM at Q6_K. Image and video generation models like SD Cascade (Würstchen v3) 3.6B use 3.5 GB, while SDXL Turbo 3.5B and SDXL Lightning 3.5B use 3.4 GB at Q6_K. Audio models such as MusicGen small/medium/large 3.3B fit at Q6_K using 3.2 GB of VRAM.
For maximum precision, you can run smaller models at Q8_0 quantization. The SmolLM3 3B, Replit Code v1.5 3B, Kandinsky 3.1 3B, Voxtral Mini or Small 3B, Orpheus TTS 3B, and Higgs Audio v2 3B models all use 3.8 GB of VRAM at Q8_0. Extremely compact models like Stable Diffusion 3.5 Medium 2.5B run at Q8_0 using only 3.2 GB of VRAM.
When a model exceeds 4 GB of VRAM, you must use CPU offload. This technique splits the model between your graphics card and system RAM. CPU offload allows you to run larger models like Mistral 7B, which needs 5.7 GB at Q4_K_M and 7.7 GB of system RAM. However, offloading introduces a severe performance cost. Shuffling data between system RAM and VRAM over the PCIe bus slows down generation speeds significantly.
Using CPU offload requires a system with sufficient RAM, such as 32 GB. Under this setup, you can run the Phi-4-multimodal 5.6B model, which needs 4.1 GB at Q4_K_M and 6.1 GB of system RAM. Popular 7B models like Qwen2.5, OLMo 2, Falcon 3, Command R7B, and OpenHermes 2.5 need 5.1 GB at Q4_K_M and 7.1 GB of system RAM. You must also monitor your context window. Running a model at its maximum 4k context limit increases VRAM usage, which can cause out of memory errors if your allocation is too close to the 4 GB limit.