Best local AI models for NVIDIA RTX 3080 Ti
12 GB GDDR6X. At a 4k context, 147 of the 233 models in our catalog with verified parameter counts fit fully, up to DeepSeek-Coder-V2 16B / 236B at 16B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 147 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| DeepSeek-Coder-V2 16B / 236B | 16B | Q4_K_M | 11.7 GB |
| Kimi-VL A3B | 16B | Q4_K_M | 11.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q4_K_M | 11 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q4_K_M | 11 GB |
| Qwen2.5 14B | 14.7B | Q4_K_M | 11.6 GB |
| Phi-3 Medium | 14B | Q5_K_M | 11.9 GB |
| Phi-4 | 14B | Q5_K_M | 11.9 GB |
| Phi-4-reasoning / -plus | 14B | Q5_K_M | 11.9 GB |
| Wan 2.2 T2I | 14B | Q5_K_M | 11.9 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q5_K_M | 11.9 GB |
| SkyReels V2 | 14B | Q5_K_M | 11.9 GB |
| Vicuna 13B | 13B | Q5_K_M | 11.1 GB |
| HunyuanVideo | 13B | Q5_K_M | 11.1 GB |
| HunyuanVideo-Avatar | 13B | Q5_K_M | 11.1 GB |
| LTX-Video / LTX-2 | 13B | Q5_K_M | 11.1 GB |
| FramePack | 13B | Q5_K_M | 11.1 GB |
| Gemma 3 12B | 12B | Q6_K | 11.8 GB |
| Gemma 4 12B | 12B | Q6_K | 11.8 GB |
| Mistral NeMo 12B | 12B | Q6_K | 11.8 GB |
| Pixtral 12B | 12B | Q6_K | 11.8 GB |
| FLUX.1 schnell | 12B | Q6_K | 11.8 GB |
| FLUX.1 Kontext dev | 12B | Q6_K | 11.8 GB |
| FLUX.1 Krea dev | 12B | Q6_K | 11.8 GB |
| Open-Sora 2.0 | 11B | Q6_K | 10.8 GB |
| Mochi 1 | 10B | Q6_K | 9.8 GB |
| Gemma 2 9B | 9B | Q6_K | 10.3 GB |
| Nemotron Nano 4B / 9B | 9B | Q8_0 | 11.4 GB |
| GLM-4 9B / GLM-4.5-Air | 9B | Q8_0 | 11.4 GB |
| Yi-Coder 1.5B / 9B | 9B | Q8_0 | 11.4 GB |
| GLM-4-9B-Chat / CodeGeeX4 | 9B | Q8_0 | 11.4 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| FLUX.1 dev | 12B | 14.4 GB needed | 16.4 GB |
| Ling-Coder-Lite | 16.8B | 12.3 GB needed | 14.3 GB |
| HunyuanImage 2.1 / 3.0 | 17B | 12.4 GB needed | 14.4 GB |
| CogVLM2 | 19B | 13.9 GB needed | 15.9 GB |
| Qwen-Image | 20B | 14.6 GB needed | 16.6 GB |
| Qwen-Image-Edit | 20B | 14.6 GB needed | 16.6 GB |
| gpt-oss-20b | 21B | 15.4 GB needed | 17.4 GB |
| Reka Flash 3 | 21B | 15.4 GB needed | 17.4 GB |
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
How to read this
The NVIDIA RTX 3080 Ti graphics card features 12 GB of GDDR6X memory. This onboard memory size dictates the maximum size of the artificial intelligence models you can run locally. To run a model entirely on the graphics card, the model files and the active memory space must fit within this 12 GB limit. Running models locally on your graphics hardware ensures the fastest processing speeds.
The quantization column indicates the compression level applied to each model. Quantization reduces the precision of model weights to make the files smaller. A Q4_K_M quant represents a four bit quantization level which offers a good balance of speed and accuracy. Higher quants like Q5_K_M or Q6_K require more memory but preserve more original model quality. Q8_0 quants provide the highest fidelity but consume the most space.
For models that fit completely within your 12 GB limit, you can run DeepSeek-Coder-V2 16B or Kimi-VL A3B at Q4_K_M quant using 11.7 GB of memory. Other options include Apriel-1.5-15B-Thinker and StarCoder2 15B at Q4_K_M quant using 11 GB. Qwen2.5 14.7B fits at Q4_K_M quant using 11.6 GB. You can also run Phi-3 Medium, Phi-4, Phi-4-reasoning / -plus, Wan 2.2 T2I, Wan 2.1 14B, or SkyReels V2 at Q5_K_M quant using 11.9 GB.
Additional fully fitting models include Vicuna 13B, HunyuanVideo, HunyuanVideo-Avatar, LTX-Video / LTX-2, and FramePack at Q5_K_M quant using 11.1 GB. Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, FLUX.1 Kontext dev, and FLUX.1 Krea dev fit at Q6_K quant using 11.8 GB. Open-Sora 2.0 fits at Q6_K quant using 10.8 GB. Mochi 1 fits at Q6_K quant using 9.8 GB. Gemma 2 9B fits at Q6_K quant using 10.3 GB. Nemotron Nano 9B, GLM-4 9B / GLM-4.5-Air, Yi-Coder 9B, and GLM-4-9B-Chat / CodeGeeX4 fit at Q8_0 quant using 11.4 GB.
When a model is too large for the 12 GB graphics memory, you must offload some layers to your system RAM. This offloading process allows you to run larger models but slows down the processing speed significantly. For these cases, we assume a system with 32 GB of system RAM. FLUX.1 dev requires 14.4 GB at FP8 / optimized and needs 16.4 GB of system RAM. Ling-Coder-Lite 16.8B requires 12.3 GB at Q4_K_M quant and needs 14.3 GB of system RAM. HunyuanImage 2.1 / 3.0 17B requires 12.4 GB at Q4_K_M quant and needs 14.4 GB of system RAM.
Other offload options include CogVLM2 19B which requires 13.9 GB at Q4_K_M quant and needs 15.9 GB of system RAM. Qwen-Image 20B and Qwen-Image-Edit 20B require 14.6 GB at Q4_K_M quant and need 16.6 GB of system RAM. The gpt-oss-20b model and Reka Flash 3 21B require 15.4 GB at Q4_K_M quant and need 17.4 GB of system RAM. Solar Pro 22B and Codestral 22B require 16.1 GB at Q4_K_M quant and need 18.1 GB of system RAM.
Be aware of the context limit when running these models. The memory numbers listed here are calculated using a basic 4k context window. If you increase the context window to process longer texts or larger prompts, the memory usage will rise. This extra memory demand can push a model past your 12 GB limit and trigger slow system RAM offloading.