Best local AI models for NVIDIA RTX 3080
10 GB GDDR6X. At a 4k context, 136 of the 233 models in our catalog with verified parameter counts fit fully, up to Vicuna 13B at 13B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 136 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Vicuna 13B | 13B | Q4_K_M | 9.5 GB |
| HunyuanVideo | 13B | Q4_K_M | 9.5 GB |
| HunyuanVideo-Avatar | 13B | Q4_K_M | 9.5 GB |
| LTX-Video / LTX-2 | 13B | Q4_K_M | 9.5 GB |
| FramePack | 13B | Q4_K_M | 9.5 GB |
| Gemma 3 12B | 12B | Q4_K_M | 8.8 GB |
| Gemma 4 12B | 12B | Q4_K_M | 8.8 GB |
| Mistral NeMo 12B | 12B | Q4_K_M | 8.8 GB |
| Pixtral 12B | 12B | Q4_K_M | 8.8 GB |
| FLUX.1 schnell | 12B | Q4_K_M | 8.8 GB |
| FLUX.1 Kontext dev | 12B | Q4_K_M | 8.8 GB |
| FLUX.1 Krea dev | 12B | Q4_K_M | 8.8 GB |
| Open-Sora 2.0 | 11B | Q5_K_M | 9.4 GB |
| Mochi 1 | 10B | Q6_K | 9.8 GB |
| Gemma 2 9B | 9B | Q5_K_M | 9.1 GB |
| Nemotron Nano 4B / 9B | 9B | Q6_K | 8.9 GB |
| GLM-4 9B / GLM-4.5-Air | 9B | Q6_K | 8.9 GB |
| Yi-Coder 1.5B / 9B | 9B | Q6_K | 8.9 GB |
| GLM-4-9B-Chat / CodeGeeX4 | 9B | Q6_K | 8.9 GB |
| GLM-4V-9B / GLM-4.1V-Thinking | 9B | Q6_K | 8.9 GB |
| Chroma | 8.9B | Q6_K | 8.8 GB |
| Llama 3.1 8B | 8B | Q6_K | 8.4 GB |
| Granite 3.3 2B / 8B | 8B | Q6_K | 7.9 GB |
| Ministral 3B / 8B | 8B | Q6_K | 7.9 GB |
| InternLM 3 8B | 8B | Q6_K | 7.9 GB |
| OpenCoder 1.5B / 8B | 8B | Q6_K | 7.9 GB |
| Seed-Coder 8B | 8B | Q6_K | 7.9 GB |
| MiniCPM-V 2.6 / MiniCPM-o 2.6 | 8B | Q6_K | 7.9 GB |
| Idefics 3 8B | 8B | Q6_K | 7.9 GB |
| Fuyu-8B | 8B | Q6_K | 7.9 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| FLUX.1 dev | 12B | 14.4 GB needed | 16.4 GB |
| Phi-3 Medium | 14B | 10.2 GB needed | 12.2 GB |
| Phi-4 | 14B | 10.2 GB needed | 12.2 GB |
| Phi-4-reasoning / -plus | 14B | 10.2 GB needed | 12.2 GB |
| Wan 2.2 T2I | 14B | 10.2 GB needed | 12.2 GB |
| Wan 2.1 (1.3B / 14B) | 14B | 10.2 GB needed | 12.2 GB |
| SkyReels V2 | 14B | 10.2 GB needed | 12.2 GB |
| Qwen2.5 14B | 14.7B | 11.6 GB needed | 13.6 GB |
| Apriel-1.5-15B-Thinker | 15B | 11 GB needed | 13 GB |
| StarCoder2 3B / 7B / 15B | 15B | 11 GB needed | 13 GB |
How to read this
The NVIDIA RTX 3080 graphics card has 10 GB of GDDR6X memory. This video memory determines which local AI models you can run directly on your hardware. To fit a model into this space, you must look at its size and quantization level. The memory size listed for each model shows how much video memory is needed to load and run that specific version.
Quantization is a method that compresses AI models to make them smaller. The best quant column shows the highest quality version that fits within your video memory. For example, Vicuna 13B, HunyuanVideo, HunyuanVideo-Avatar, LTX-Video / LTX-2, and FramePack are 13B models that fit using the Q4_K_M quantization, which uses 9.5 GB of video memory. Models like Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, FLUX.1 Kontext dev, and FLUX.1 Krea dev fit using the Q4_K_M quantization, which uses 8.8 GB of video memory.
Slightly smaller models can use higher quality quantizations. Open-Sora 2.0 is an 11B model using the Q5_K_M quantization at 9.4 GB of video memory. Mochi 1 is a 10B model using the Q6_K quantization at 9.8 GB of video memory. Gemma 2 9B uses the Q5_K_M quantization at 9.1 GB of video memory. Nemotron Nano 4B / 9B, GLM-4 9B / GLM-4.5-Air, Yi-Coder 1.5B / 9B, GLM-4-9B-Chat / CodeGeeX4, and GLM-4V-9B / GLM-4.1V-Thinking are 9B models using the Q6_K quantization at 8.9 GB of video memory. Chroma is an 8.9B model using the Q6_K quantization at 8.8 GB of video memory.
Many 8B models run with excellent quality using the Q6_K quantization. Llama 3.1 8B uses 8.4 GB of video memory. Granite 3.3 2B / 8B, Ministral 3B / 8B, InternLM 3 8B, OpenCoder 1.5B / 8B, Seed-Coder 8B, MiniCPM-V 2.6 / MiniCPM-o 2.6, Idefics 3 8B, and Fuyu-8B all use 7.9 GB of video memory. Running these models leaves some video memory free for your system display.
When a model is too large for your video memory, you can use CPU offload if you have 32 GB of system RAM. This process splits the model between your graphics card and system memory, which slows down generation speeds. FLUX.1 dev is a 12B model that needs 14.4 GB at FP8 / optimized and uses 16.4 GB of system RAM. Phi-3 Medium, Phi-4, Phi-4-reasoning / -plus, Wan 2.2 T2I, Wan 2.1 (1.3B / 14B), and SkyReels V2 are 14B models that need 10.2 GB at Q4_K_M and use 12.2 GB of system RAM. Qwen2.5 14B is a 14.7B model that needs 11.6 GB at Q4_K_M and uses 13.6 GB of system RAM. Apriel-1.5-15B-Thinker and StarCoder2 3B / 7B / 15B are 15B models that need 11 GB at Q4_K_M and use 13 GB of system RAM.
All memory calculations assume a standard 4k context window. If you increase the context window to process longer texts, the model will require more video memory. You may need to choose a smaller model or a lower quantization level to prevent running out of memory during long conversations.