Best local AI models for NVIDIA RTX 2080 Ti
11 GB GDDR6. At a 4k context, 144 of the 233 models in our catalog with verified parameter counts fit fully, up to Apriel-1.5-15B-Thinker at 15B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 144 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Apriel-1.5-15B-Thinker | 15B | Q4_K_M | 11 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q4_K_M | 11 GB |
| Phi-3 Medium | 14B | Q4_K_M | 10.2 GB |
| Phi-4 | 14B | Q4_K_M | 10.2 GB |
| Phi-4-reasoning / -plus | 14B | Q4_K_M | 10.2 GB |
| Wan 2.2 T2I | 14B | Q4_K_M | 10.2 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q4_K_M | 10.2 GB |
| SkyReels V2 | 14B | Q4_K_M | 10.2 GB |
| Vicuna 13B | 13B | Q4_K_M | 9.5 GB |
| HunyuanVideo | 13B | Q4_K_M | 9.5 GB |
| HunyuanVideo-Avatar | 13B | Q4_K_M | 9.5 GB |
| LTX-Video / LTX-2 | 13B | Q4_K_M | 9.5 GB |
| FramePack | 13B | Q4_K_M | 9.5 GB |
| Gemma 3 12B | 12B | Q5_K_M | 10.2 GB |
| Gemma 4 12B | 12B | Q5_K_M | 10.2 GB |
| Mistral NeMo 12B | 12B | Q5_K_M | 10.2 GB |
| Pixtral 12B | 12B | Q5_K_M | 10.2 GB |
| FLUX.1 schnell | 12B | Q5_K_M | 10.2 GB |
| FLUX.1 Kontext dev | 12B | Q5_K_M | 10.2 GB |
| FLUX.1 Krea dev | 12B | Q5_K_M | 10.2 GB |
| Open-Sora 2.0 | 11B | Q6_K | 10.8 GB |
| Mochi 1 | 10B | Q6_K | 9.8 GB |
| Gemma 2 9B | 9B | Q6_K | 10.3 GB |
| Nemotron Nano 4B / 9B | 9B | Q6_K | 8.9 GB |
| GLM-4 9B / GLM-4.5-Air | 9B | Q6_K | 8.9 GB |
| Yi-Coder 1.5B / 9B | 9B | Q6_K | 8.9 GB |
| GLM-4-9B-Chat / CodeGeeX4 | 9B | Q6_K | 8.9 GB |
| GLM-4V-9B / GLM-4.1V-Thinking | 9B | Q6_K | 8.9 GB |
| Chroma | 8.9B | Q6_K | 8.8 GB |
| Llama 3.1 8B | 8B | Q8_0 | 10.7 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| FLUX.1 dev | 12B | 14.4 GB needed | 16.4 GB |
| Qwen2.5 14B | 14.7B | 11.6 GB needed | 13.6 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | 11.7 GB needed | 13.7 GB |
| Kimi-VL A3B | 16B | 11.7 GB needed | 13.7 GB |
| Ling-Coder-Lite | 16.8B | 12.3 GB needed | 14.3 GB |
| HunyuanImage 2.1 / 3.0 | 17B | 12.4 GB needed | 14.4 GB |
| CogVLM2 | 19B | 13.9 GB needed | 15.9 GB |
| Qwen-Image | 20B | 14.6 GB needed | 16.6 GB |
| Qwen-Image-Edit | 20B | 14.6 GB needed | 16.6 GB |
| gpt-oss-20b | 21B | 15.4 GB needed | 17.4 GB |
How to read this
The NVIDIA RTX 2080 Ti features 11 GB GDDR6 memory. This onboard memory determines the maximum size of the artificial intelligence models you can run locally. To fit a model entirely on this graphics card, the model files and the active memory must stay under this 11 GB limit. Running models completely inside your video memory ensures the fastest possible processing speeds.
Quantization is a method that compresses model files so they require less memory. The best quantization column shows the optimal balance of size and quality for this hardware. For example, the 15B models Apriel-1.5-15B-Thinker and StarCoder2 15B fit within 11 GB used when using the Q4_K_M quantization. Other models like Phi-3 Medium, Phi-4, Phi-4-reasoning / -plus, Wan 2.2 T2I, Wan 2.1 14B, and SkyReels V2 use 10.2 GB of memory at this same Q4_K_M level.
Smaller models can run with higher precision quantizations because they have lower memory requirements. The 12B models Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, FLUX.1 Kontext dev, and FLUX.1 Krea dev use 10.2 GB of memory with the Q5_K_M quantization. Models like Open-Sora 2.0 use 10.8 GB at Q6_K, while Llama 3.1 8B can run at the high quality Q8_0 quantization using 10.7 GB of video memory.
When a model exceeds the 11 GB video memory limit, you must use CPU offloading. This process splits the model workload between your graphics card and your system memory. Offloading allows you to run larger models, but it costs significant processing speed because system RAM is much slower than GDDR6 video memory. For these cases, we assume your computer has 32 GB of system RAM available.
Several larger models require this CPU offload setup to function. FLUX.1 dev needs 14.4 GB at FP8 or optimized settings, which requires 16.4 GB of system RAM. The Qwen2.5 14B model needs 11.6 GB at Q4_K_M and 13.6 GB of system RAM. Other offload options include DeepSeek-Coder-V2 16B, Kimi-VL A3B, Ling-Coder-Lite, HunyuanImage 2.1 / 3.0, CogVLM2, Qwen-Image, Qwen-Image-Edit, and gpt-oss-20b.
You must also consider the memory cost of text history. The memory figures listed here are measured using a standard 4k context window. If you increase the context window to remember longer conversations, the system will require more video memory. This extra memory demand can force a model that normally fits on the card to require CPU offloading instead.