Best local AI models for NVIDIA RTX 2080 Ti

11 GB GDDR6. At a 4k context, 144 of the 233 models in our catalog with verified parameter counts fit fully, up to Apriel-1.5-15B-Thinker at 15B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 144 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Apriel-1.5-15B-Thinker15BQ4_K_M11 GB
StarCoder2 3B / 7B / 15B15BQ4_K_M11 GB
Phi-3 Medium14BQ4_K_M10.2 GB
Phi-414BQ4_K_M10.2 GB
Phi-4-reasoning / -plus14BQ4_K_M10.2 GB
Wan 2.2 T2I14BQ4_K_M10.2 GB
Wan 2.1 (1.3B / 14B)14BQ4_K_M10.2 GB
SkyReels V214BQ4_K_M10.2 GB
Vicuna 13B13BQ4_K_M9.5 GB
HunyuanVideo13BQ4_K_M9.5 GB
HunyuanVideo-Avatar13BQ4_K_M9.5 GB
LTX-Video / LTX-213BQ4_K_M9.5 GB
FramePack13BQ4_K_M9.5 GB
Gemma 3 12B12BQ5_K_M10.2 GB
Gemma 4 12B12BQ5_K_M10.2 GB
Mistral NeMo 12B12BQ5_K_M10.2 GB
Pixtral 12B12BQ5_K_M10.2 GB
FLUX.1 schnell12BQ5_K_M10.2 GB
FLUX.1 Kontext dev12BQ5_K_M10.2 GB
FLUX.1 Krea dev12BQ5_K_M10.2 GB
Open-Sora 2.011BQ6_K10.8 GB
Mochi 110BQ6_K9.8 GB
Gemma 2 9B9BQ6_K10.3 GB
Nemotron Nano 4B / 9B9BQ6_K8.9 GB
GLM-4 9B / GLM-4.5-Air9BQ6_K8.9 GB
Yi-Coder 1.5B / 9B9BQ6_K8.9 GB
GLM-4-9B-Chat / CodeGeeX49BQ6_K8.9 GB
GLM-4V-9B / GLM-4.1V-Thinking9BQ6_K8.9 GB
Chroma8.9BQ6_K8.8 GB
Llama 3.1 8B8BQ8_010.7 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
FLUX.1 dev12B14.4 GB needed16.4 GB
Qwen2.5 14B14.7B11.6 GB needed13.6 GB
DeepSeek-Coder-V2 16B / 236B16B11.7 GB needed13.7 GB
Kimi-VL A3B16B11.7 GB needed13.7 GB
Ling-Coder-Lite16.8B12.3 GB needed14.3 GB
HunyuanImage 2.1 / 3.017B12.4 GB needed14.4 GB
CogVLM219B13.9 GB needed15.9 GB
Qwen-Image20B14.6 GB needed16.6 GB
Qwen-Image-Edit20B14.6 GB needed16.6 GB
gpt-oss-20b21B15.4 GB needed17.4 GB

How to read this

The NVIDIA RTX 2080 Ti features 11 GB GDDR6 memory. This onboard memory determines the maximum size of the artificial intelligence models you can run locally. To fit a model entirely on this graphics card, the model files and the active memory must stay under this 11 GB limit. Running models completely inside your video memory ensures the fastest possible processing speeds.

Quantization is a method that compresses model files so they require less memory. The best quantization column shows the optimal balance of size and quality for this hardware. For example, the 15B models Apriel-1.5-15B-Thinker and StarCoder2 15B fit within 11 GB used when using the Q4_K_M quantization. Other models like Phi-3 Medium, Phi-4, Phi-4-reasoning / -plus, Wan 2.2 T2I, Wan 2.1 14B, and SkyReels V2 use 10.2 GB of memory at this same Q4_K_M level.

Smaller models can run with higher precision quantizations because they have lower memory requirements. The 12B models Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, FLUX.1 Kontext dev, and FLUX.1 Krea dev use 10.2 GB of memory with the Q5_K_M quantization. Models like Open-Sora 2.0 use 10.8 GB at Q6_K, while Llama 3.1 8B can run at the high quality Q8_0 quantization using 10.7 GB of video memory.

When a model exceeds the 11 GB video memory limit, you must use CPU offloading. This process splits the model workload between your graphics card and your system memory. Offloading allows you to run larger models, but it costs significant processing speed because system RAM is much slower than GDDR6 video memory. For these cases, we assume your computer has 32 GB of system RAM available.

Several larger models require this CPU offload setup to function. FLUX.1 dev needs 14.4 GB at FP8 or optimized settings, which requires 16.4 GB of system RAM. The Qwen2.5 14B model needs 11.6 GB at Q4_K_M and 13.6 GB of system RAM. Other offload options include DeepSeek-Coder-V2 16B, Kimi-VL A3B, Ling-Coder-Lite, HunyuanImage 2.1 / 3.0, CogVLM2, Qwen-Image, Qwen-Image-Edit, and gpt-oss-20b.

You must also consider the memory cost of text history. The memory figures listed here are measured using a standard 4k context window. If you increase the context window to remember longer conversations, the system will require more video memory. This extra memory demand can force a model that normally fits on the card to require CPU offloading instead.