Best local AI models for NVIDIA GTX 780

3 GB GDDR5. At a 4k context, 76 of the 233 models in our catalog with verified parameter counts fit fully, up to Qwen3 4B at 4B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 76 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Qwen3 4B4BQ4_K_M2.9 GB
Gemma 3 4B4BQ4_K_M2.9 GB
Gemma 4 E4B4BQ4_K_M2.9 GB
MiniCPM 3 4B4BQ4_K_M2.9 GB
Danube 3 4B4BQ4_K_M2.9 GB
Fish Speech 1.5 / OpenAudio S14BQ4_K_M2.9 GB
Phi-4-mini-instruct3.8BQ4_K_M2.8 GB
Phi-3.5 Mini3.8BQ4_K_M2.8 GB
OmniGen / OmniGen23.8BQ4_K_M2.8 GB
SD Cascade (Würstchen v3)3.6BQ4_K_M2.6 GB
SDXL Turbo3.5BQ5_K_M3 GB
SDXL Lightning3.5BQ5_K_M3 GB
ACE-Step3.5BQ5_K_M3 GB
MusicGen small/medium/large3.3BQ5_K_M2.8 GB
SmolLM3 3B3BQ6_K3 GB
Replit Code v1.5 3B3BQ6_K3 GB
Kandinsky 3.13BQ6_K3 GB
Voxtral Mini / Small3BQ6_K3 GB
Orpheus TTS3BQ6_K3 GB
Higgs Audio v23BQ6_K3 GB
Allegro2.8BQ6_K2.8 GB
Open-Sora Plan2.7BQ6_K2.7 GB
LFM2 1.2B / 2.6B2.6BQ6_K2.6 GB
Playground v2.52.6BQ6_K2.6 GB
Stable Diffusion 3.5 Medium2.5BQ6_K2.5 GB
Canary 1B / Qwen-2.5B2.5BQ6_K2.5 GB
SeamlessM4T v22.3BQ8_02.9 GB
Parler-TTS2.2BQ8_02.8 GB
Kimi K3 DSpark2.2BQ8_02.9 GB
SmolVLM 256M / 500M / 2B2BQ8_02.5 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-3.5-vision4.2B3.1 GB needed5.1 GB
DeepFloyd IF4.3B3.1 GB needed5.1 GB
DeepSeek-VL24.5B3.3 GB needed5.3 GB
Lumina-Next / Lumina-Image 2.05B3.7 GB needed5.7 GB
CogVideoX 2B / 5B5B3.7 GB needed5.7 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB

How to read this

The NVIDIA GTX 780 graphics card has 3 GB GDDR5 memory. This physical memory size determines which local AI models can run directly on your hardware. To fit inside this limit, you must look at the model size and the quantization level. Quantization reduces the precision of model weights to make the files smaller. The quant column shows the best possible quantization level that still fits within your video memory.

For models that fit entirely on the card, you can run several 4B parameter models at a Q4_K_M quantization. Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, Danube 3 4B, and Fish Speech 1.5 / OpenAudio S1 all use 2.9 GB of video memory. Phi-4-mini-instruct, Phi-3.5 Mini, and OmniGen / OmniGen2 are 3.8B models that use 2.8 GB of video memory at Q4_K_M. SD Cascade (Würstchen v3) is a 3.6B model that uses 2.6 GB of video memory at Q4_K_M.

Slightly smaller models can run at higher quantization levels for better quality. SDXL Turbo, SDXL Lightning, and ACE-Step are 3.5B models that use exactly 3 GB of video memory at Q5_K_M. MusicGen small/medium/large is a 3.3B model that uses 2.8 GB of video memory at Q5_K_M. You can also run 3B models like SmolLM3 3B, Replit Code v1.5 3B, Kandinsky 3.1, Voxtral Mini / Small, Orpheus TTS, and Higgs Audio v2 at Q6_K quantization using 3 GB of video memory.

Other models that fit completely include Allegro at 2.8B using 2.8 GB of video memory at Q6_K. Open-Sora Plan is a 2.7B model using 2.7 GB of video memory at Q6_K. LFM2 1.2B / 2.6B and Playground v2.5 are 2.6B models using 2.6 GB of video memory at Q6_K. Stable Diffusion 3.5 Medium and Canary 1B / Qwen-2.5B are 2.5B models using 2.5 GB of video memory at Q6_K. SeamlessM4T v2, Parler-TTS, Kimi K3 DSpark, and SmolVLM 256M / 500M / 2B use Q8_0 quantization to fit within 2.5 GB to 2.9 GB of video memory.

When a model is too large for the 3 GB video memory, you can offload parts of it to your system RAM. This offload process allows you to run larger models but it reduces your processing speed. For these cases, we assume your system has 32 GB system RAM. For example, Stable Diffusion XL needs 4.1 GB at FP8 / optimized and uses 6.1 GB system RAM. Phi-3 Mini needs 4.4 GB at Q4_K_M and uses 6.4 GB system RAM. Phi-3.5-vision and DeepFloyd IF need 3.1 GB at Q4_K_M and use 5.1 GB system RAM.

Other offload options include DeepSeek-VL2 which needs 3.3 GB at Q4_K_M and uses 5.3 GB system RAM. Lumina-Next / Lumina-Image 2.0 and CogVideoX 2B / 5B need 3.7 GB at Q4_K_M and use 5.7 GB system RAM. Phi-4-multimodal needs 4.1 GB at Q4_K_M and uses 6.1 GB system RAM. Magicoder-S-DS 6.7B needs 4.9 GB at Q4_K_M and uses 6.9 GB system RAM. Mistral 7B needs 5.7 GB at Q4_K_M and uses 7.7 GB system RAM.

You must consider the context window when running these models. The memory numbers listed here are calculated for a standard 4k context window. If you increase the context window size, the model will require more memory. This extra memory usage can exceed your 3 GB limit and force the system to slow down or fail to run.