Best local AI models for NVIDIA GTX 780 Ti
3 GB GDDR5. At a 4k context, 76 of the 233 models in our catalog with verified parameter counts fit fully, up to Qwen3 4B at 4B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 76 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Qwen3 4B | 4B | Q4_K_M | 2.9 GB |
| Gemma 3 4B | 4B | Q4_K_M | 2.9 GB |
| Gemma 4 E4B | 4B | Q4_K_M | 2.9 GB |
| MiniCPM 3 4B | 4B | Q4_K_M | 2.9 GB |
| Danube 3 4B | 4B | Q4_K_M | 2.9 GB |
| Fish Speech 1.5 / OpenAudio S1 | 4B | Q4_K_M | 2.9 GB |
| Phi-4-mini-instruct | 3.8B | Q4_K_M | 2.8 GB |
| Phi-3.5 Mini | 3.8B | Q4_K_M | 2.8 GB |
| OmniGen / OmniGen2 | 3.8B | Q4_K_M | 2.8 GB |
| SD Cascade (Würstchen v3) | 3.6B | Q4_K_M | 2.6 GB |
| SDXL Turbo | 3.5B | Q5_K_M | 3 GB |
| SDXL Lightning | 3.5B | Q5_K_M | 3 GB |
| ACE-Step | 3.5B | Q5_K_M | 3 GB |
| MusicGen small/medium/large | 3.3B | Q5_K_M | 2.8 GB |
| SmolLM3 3B | 3B | Q6_K | 3 GB |
| Replit Code v1.5 3B | 3B | Q6_K | 3 GB |
| Kandinsky 3.1 | 3B | Q6_K | 3 GB |
| Voxtral Mini / Small | 3B | Q6_K | 3 GB |
| Orpheus TTS | 3B | Q6_K | 3 GB |
| Higgs Audio v2 | 3B | Q6_K | 3 GB |
| Allegro | 2.8B | Q6_K | 2.8 GB |
| Open-Sora Plan | 2.7B | Q6_K | 2.7 GB |
| LFM2 1.2B / 2.6B | 2.6B | Q6_K | 2.6 GB |
| Playground v2.5 | 2.6B | Q6_K | 2.6 GB |
| Stable Diffusion 3.5 Medium | 2.5B | Q6_K | 2.5 GB |
| Canary 1B / Qwen-2.5B | 2.5B | Q6_K | 2.5 GB |
| SeamlessM4T v2 | 2.3B | Q8_0 | 2.9 GB |
| Parler-TTS | 2.2B | Q8_0 | 2.8 GB |
| Kimi K3 DSpark | 2.2B | Q8_0 | 2.9 GB |
| SmolVLM 256M / 500M / 2B | 2B | Q8_0 | 2.5 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| Stable Diffusion XL | 3.417B | 4.1 GB needed | 6.1 GB |
| Phi-3 Mini | 3.8B | 4.4 GB needed | 6.4 GB |
| Phi-3.5-vision | 4.2B | 3.1 GB needed | 5.1 GB |
| DeepFloyd IF | 4.3B | 3.1 GB needed | 5.1 GB |
| DeepSeek-VL2 | 4.5B | 3.3 GB needed | 5.3 GB |
| Lumina-Next / Lumina-Image 2.0 | 5B | 3.7 GB needed | 5.7 GB |
| CogVideoX 2B / 5B | 5B | 3.7 GB needed | 5.7 GB |
| Phi-4-multimodal | 5.6B | 4.1 GB needed | 6.1 GB |
| Magicoder-S-DS 6.7B | 6.7B | 4.9 GB needed | 6.9 GB |
| Mistral 7B | 7B | 5.7 GB needed | 7.7 GB |
How to read this
The NVIDIA GTX 780 Ti features 3 GB of GDDR5 graphics memory. This memory capacity determines the size of the artificial intelligence models you can run locally. To fit inside this 3 GB limit, models must undergo quantization. Quantization is a compression method that reduces model size while preserving most of the original reasoning capability.
The best quantization column shows the optimal format for each model on this hardware. For example, Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, Danube 3 4B, and Fish Speech 1.5 / OpenAudio S1 all run at the Q4_K_M quantization level. This compression allows these 4B models to fit into 2.9 GB of graphics memory. Models like Phi-4-mini-instruct, Phi-3.5 Mini, and OmniGen / OmniGen2 use the same Q4_K_M quantization to fit 3.8B parameters into 2.8 GB.
Other models use different quantization levels to balance size and quality. SD Cascade (Würstchen v3) fits 3.6B parameters into 2.6 GB using Q4_K_M. SDXL Turbo, SDXL Lightning, and ACE-Step use Q5_K_M to fit 3.5B parameters into 3 GB. MusicGen small/medium/large fits 3.3B parameters into 2.8 GB using Q5_K_M. SmolLM3 3B, Replit Code v1.5 3B, Kandinsky 3.1, Voxtral Mini / Small, Orpheus TTS, and Higgs Audio v2 use Q6_K to fit 3B parameters into 3 GB.
Smaller models also run efficiently on this card. Allegro fits 2.8B parameters into 2.8 GB using Q6_K. Open-Sora Plan fits 2.7B parameters into 2.7 GB using Q6_K. LFM2 1.2B / 2.6B and Playground v2.5 fit 2.6B parameters into 2.6 GB using Q6_K. Stable Diffusion 3.5 Medium and Canary 1B / Qwen-2.5B fit 2.5B parameters into 2.5 GB using Q6_K. SeamlessM4T v2 fits 2.3B parameters into 2.9 GB using Q8_0. Parler-TTS fits 2.2B parameters into 2.8 GB using Q8_0. Kimi K3 DSpark fits 2.2B parameters into 2.9 GB using Q8_0. SmolVLM 256M / 500M / 2B fits 2B parameters into 2.5 GB using Q8_0.
You can run larger models by offloading parts of the workload to your system RAM. This process requires a system with 32 GB of system RAM. Offloading allows you to run Stable Diffusion XL, which needs 4.1 GB at FP8 / optimized and 6.1 GB of system RAM. Phi-3 Mini needs 4.4 GB at Q4_K_M and 6.4 GB of system RAM. Phi-3.5-vision and DeepFloyd IF both need 3.1 GB at Q4_K_M and 5.1 GB of system RAM. DeepSeek-VL2 needs 3.3 GB at Q4_K_M and 5.3 GB of system RAM.
Larger offloaded models require even more system memory. Lumina-Next / Lumina-Image 2.0 and CogVideoX 2B / 5B need 3.7 GB at Q4_K_M and 5.7 GB of system RAM. Phi-4-multimodal needs 4.1 GB at Q4_K_M and 6.1 GB of system RAM. Magicoder-S-DS 6.7B needs 4.9 GB at Q4_K_M and 6.9 GB of system RAM. Mistral 7B needs 5.7 GB at Q4_K_M and 7.7 GB of system RAM. Offloading makes these models fit but it reduces processing speed because system RAM is slower than graphics memory.
When running text models, you must consider the context limit. The memory figures listed here assume a standard 4k context window. If you increase the context window to process longer documents, the model will require more graphics memory. This extra memory usage can exceed the 3 GB limit of your card and cause the system to slow down or fail.