Best local AI models for NVIDIA GTX 1080 Ti

11 GB GDDR5X. At a 4k context, 144 of the 233 models in our catalog with verified parameter counts fit fully, up to Apriel-1.5-15B-Thinker at 15B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 144 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Apriel-1.5-15B-Thinker15BQ4_K_M11 GB
StarCoder2 3B / 7B / 15B15BQ4_K_M11 GB
Phi-3 Medium14BQ4_K_M10.2 GB
Phi-414BQ4_K_M10.2 GB
Phi-4-reasoning / -plus14BQ4_K_M10.2 GB
Wan 2.2 T2I14BQ4_K_M10.2 GB
Wan 2.1 (1.3B / 14B)14BQ4_K_M10.2 GB
SkyReels V214BQ4_K_M10.2 GB
Vicuna 13B13BQ4_K_M9.5 GB
HunyuanVideo13BQ4_K_M9.5 GB
HunyuanVideo-Avatar13BQ4_K_M9.5 GB
LTX-Video / LTX-213BQ4_K_M9.5 GB
FramePack13BQ4_K_M9.5 GB
Gemma 3 12B12BQ5_K_M10.2 GB
Gemma 4 12B12BQ5_K_M10.2 GB
Mistral NeMo 12B12BQ5_K_M10.2 GB
Pixtral 12B12BQ5_K_M10.2 GB
FLUX.1 schnell12BQ5_K_M10.2 GB
FLUX.1 Kontext dev12BQ5_K_M10.2 GB
FLUX.1 Krea dev12BQ5_K_M10.2 GB
Open-Sora 2.011BQ6_K10.8 GB
Mochi 110BQ6_K9.8 GB
Gemma 2 9B9BQ6_K10.3 GB
Nemotron Nano 4B / 9B9BQ6_K8.9 GB
GLM-4 9B / GLM-4.5-Air9BQ6_K8.9 GB
Yi-Coder 1.5B / 9B9BQ6_K8.9 GB
GLM-4-9B-Chat / CodeGeeX49BQ6_K8.9 GB
GLM-4V-9B / GLM-4.1V-Thinking9BQ6_K8.9 GB
Chroma8.9BQ6_K8.8 GB
Llama 3.1 8B8BQ8_010.7 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
FLUX.1 dev12B14.4 GB needed16.4 GB
Qwen2.5 14B14.7B11.6 GB needed13.6 GB
DeepSeek-Coder-V2 16B / 236B16B11.7 GB needed13.7 GB
Kimi-VL A3B16B11.7 GB needed13.7 GB
Ling-Coder-Lite16.8B12.3 GB needed14.3 GB
HunyuanImage 2.1 / 3.017B12.4 GB needed14.4 GB
CogVLM219B13.9 GB needed15.9 GB
Qwen-Image20B14.6 GB needed16.6 GB
Qwen-Image-Edit20B14.6 GB needed16.6 GB
gpt-oss-20b21B15.4 GB needed17.4 GB

How to read this

The NVIDIA GTX 1080 Ti features 11 GB of GDDR5X memory. This video memory size determines which local AI models can run entirely on your graphics card. When a model fits completely within this memory space, you get the fastest possible processing speeds. If a model exceeds this limit, you must use system memory offloading which slows down performance.

The quantization column shows the compression level used to fit these models into your video memory. A quant like Q4_K_M or Q5_K_M reduces the size of the model weights. This compression allows larger models to run on your hardware with very little loss in output quality. Lower quants like Q4_K_M use less memory while higher quants like Q8_0 or Q6_K offer better precision but require more space.

For maximum performance, you can run models up to 15B parameters entirely on the card. The Apriel-1.5-15B-Thinker and StarCoder2 15B models fit within the limit at Q4_K_M using exactly 11 GB. You can also run the 14B models like Phi-3 Medium, Phi-4, Phi-4-reasoning, Phi-4-plus, Wan 2.2 T2I, Wan 2.1 14B, and SkyReels V2 at Q4_K_M which use 10.2 GB. The 13B models like Vicuna 13B, HunyuanVideo, HunyuanVideo-Avatar, LTX-Video, LTX-2, and FramePack use 9.5 GB at Q4_K_M.

Slightly smaller models can run at higher quantization levels for better accuracy. The 12B models like Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, FLUX.1 Kontext dev, and FLUX.1 Krea dev use 10.2 GB at Q5_K_M. The Open-Sora 2.0 model uses 10.8 GB at Q6_K. The Mochi 1 model uses 9.8 GB at Q6_K. The Gemma 2 9B model uses 10.3 GB at Q6_K. Other 9B models like Nemotron Nano 9B, GLM-4 9B, GLM-4.5-Air, Yi-Coder 9B, GLM-4-9B-Chat, CodeGeeX4, GLM-4V-9B, and GLM-4.1V-Thinking use 8.9 GB at Q6_K. The Chroma model uses 8.8 GB at Q6_K. The Llama 3.1 8B model uses 10.7 GB at Q8_0.

When a model is too large for the 11 GB video memory, you can offload the extra data to your system RAM. This process requires a system with 32 GB of system RAM. Offloading allows you to run FLUX.1 dev which needs 14.4 GB at FP8 and uses 16.4 GB of system RAM. You can also run Qwen2.5 14B which needs 11.6 GB at Q4_K_M and uses 13.6 GB of system RAM. DeepSeek-Coder-V2 16B and Kimi-VL A3B need 11.7 GB at Q4_K_M and use 13.7 GB of system RAM. Ling-Coder-Lite needs 12.3 GB at Q4_K_M and uses 14.3 GB of system RAM. HunyuanImage 2.1 and HunyuanImage 3.0 need 12.4 GB at Q4_K_M and use 14.4 GB of system RAM. CogVLM2 needs 13.9 GB at Q4_K_M and uses 15.9 GB of system RAM. Qwen-Image and Qwen-Image-Edit need 14.6 GB at Q4_K_M and use 16.6 GB of system RAM. The gpt-oss-20b model needs 15.4 GB at Q4_K_M and uses 17.4 GB of system RAM.

You must consider the context limit when planning your memory usage. The memory figures listed here are calculated using a standard 4k context window. If you increase the context length to process longer documents, the memory usage will grow. Running close to the 11 GB limit of your card with a large context window can cause out of memory errors or force your system into slow CPU offloading.