Best local AI models for NVIDIA GTX 1660

6 GB GDDR5. At a 4k context, 114 of the 233 models in our catalog with verified parameter counts fit fully, up to Granite 3.3 2B / 8B at 8B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 114 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Granite 3.3 2B / 8B8BQ4_K_M5.9 GB
Ministral 3B / 8B8BQ4_K_M5.9 GB
InternLM 3 8B8BQ4_K_M5.9 GB
OpenCoder 1.5B / 8B8BQ4_K_M5.9 GB
Seed-Coder 8B8BQ4_K_M5.9 GB
MiniCPM-V 2.6 / MiniCPM-o 2.68BQ4_K_M5.9 GB
Idefics 3 8B8BQ4_K_M5.9 GB
Fuyu-8B8BQ4_K_M5.9 GB
Emu38BQ4_K_M5.9 GB
Stable Diffusion 3.5 Large / Turbo8BQ4_K_M5.9 GB
EXAONE 3.5 2.4B / 7.8B7.8BQ4_K_M5.7 GB
Mistral 7B7BQ4_K_M5.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7BQ5_K_M6 GB
OLMo 2 1B / 7B7BQ5_K_M6 GB
Falcon 3 1B / 3B / 7B7BQ5_K_M6 GB
Command R7B7BQ5_K_M6 GB
OpenHermes 2.57BQ5_K_M6 GB
Zephyr 7B Beta7BQ5_K_M6 GB
OpenChat 3.57BQ5_K_M6 GB
Starling LM 7B7BQ5_K_M6 GB
Codestral Mamba 7B7BQ5_K_M6 GB
CodeGemma 2B / 7B7BQ5_K_M6 GB
aiXcoder-7B7BQ5_K_M6 GB
Nxcode / CodeQwen 1.5 7B7BQ5_K_M6 GB
Janus-Pro 1B / 7B7BQ5_K_M6 GB
Ruyi-Mini-7B7BQ5_K_M6 GB
Qwen2-Audio 7B7BQ5_K_M6 GB
Qwen2.5-Omni 3B / 7B7BQ5_K_M6 GB
YuE7BQ5_K_M6 GB
Magicoder-S-DS 6.7B6.7BQ5_K_M5.7 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at Q4_K_MSystem RAM at 4k
Llama 3.1 8B8B6.4 GB needed8.4 GB
Chroma8.9B6.5 GB needed8.5 GB
Gemma 2 9B9B8 GB needed10 GB
Nemotron Nano 4B / 9B9B6.6 GB needed8.6 GB
GLM-4 9B / GLM-4.5-Air9B6.6 GB needed8.6 GB
Yi-Coder 1.5B / 9B9B6.6 GB needed8.6 GB
GLM-4-9B-Chat / CodeGeeX49B6.6 GB needed8.6 GB
GLM-4V-9B / GLM-4.1V-Thinking9B6.6 GB needed8.6 GB
Mochi 110B7.3 GB needed9.3 GB
Open-Sora 2.011B8.1 GB needed10.1 GB

How to read this

The NVIDIA GTX 1660 graphics card features 6 GB of GDDR5 video memory. This memory limit determines which local AI models can run directly on your hardware. To fit models within this space, we use quantized versions. The quantization column shows the compression level used to reduce model size while keeping accuracy high. A Q4_K_M quant represents a four bit quantization level. A Q5_K_M quant represents a five bit quantization level.

For complete on board execution, models must fit within the 6 GB limit. The largest fully fitting models include Granite 8B, Ministral 8B, InternLM 3 8B, OpenCoder 8B, Seed-Coder 8B, MiniCPM-V 2.6, Idefics 3 8B, Fuyu-8B, Emu3, and Stable Diffusion 3.5 Large. These models run at the Q4_K_M quant and use 5.9 GB of video memory. EXAONE 3.5 7.8B and Mistral 7B also fit completely using 5.7 GB of video memory at the Q4_K_M quant.

Several 7B models can run entirely in video memory at a higher Q5_K_M quant using exactly 6 GB of space. This group includes Qwen2.5 7B, OLMo 2 7B, Falcon 3 7B, Command R7B, OpenHermes 2.5, Zephyr 7B Beta, OpenChat 3.5, Starling LM 7B, Codestral Mamba 7B, CodeGemma 7B, aiXcoder-7B, Nxcode, Janus-Pro 7B, Ruyi-Mini-7B, Qwen2-Audio 7B, Qwen2.5-Omni 7B, and YuE. Magicoder-S-DS 6.7B also fits at Q5_K_M using 5.7 GB of video memory.

When a model exceeds the 6 GB video memory limit, you can use CPU offload. This technique splits the model between your graphics card and system RAM. CPU offloading requires a system with 32 GB of system RAM. Offloading allows you to run larger models, but it costs performance. Your processing speed will decrease because system RAM is much slower than the GDDR5 video memory on your graphics card.

Several popular models require CPU offloading on this hardware. Llama 3.1 8B needs 6.4 GB at Q4_K_M and uses 8.4 GB of system RAM. Chroma needs 6.5 GB at Q4_K_M and uses 8.5 GB of system RAM. Gemma 2 9B needs 8 GB at Q4_K_M and uses 10 GB of system RAM. Nemotron Nano 9B, GLM-4 9B, Yi-Coder 9B, GLM-4-9B-Chat, and GLM-4V-9B all need 6.6 GB at Q4_K_M and use 8.6 GB of system RAM. Mochi 1 needs 7.3 GB at Q4_K_M and uses 9.3 GB of system RAM. Open-Sora 2.0 needs 8.1 GB at Q4_K_M and uses 10.1 GB of system RAM.

All memory calculations are based on a standard 4k context window. Running a model with a larger context window requires more memory for the active conversation history. If you increase the context length beyond 4k tokens, the model will exceed the 6 GB video memory limit much faster. This will force the system to use CPU offload even for models that normally fit entirely on the graphics card.