Best local AI models for NVIDIA GTX 1660 Ti Laptop
6 GB GDDR6. At a 4k context, 114 of the 233 models in our catalog with verified parameter counts fit fully, up to Granite 3.3 2B / 8B at 8B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 114 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Granite 3.3 2B / 8B | 8B | Q4_K_M | 5.9 GB |
| Ministral 3B / 8B | 8B | Q4_K_M | 5.9 GB |
| InternLM 3 8B | 8B | Q4_K_M | 5.9 GB |
| OpenCoder 1.5B / 8B | 8B | Q4_K_M | 5.9 GB |
| Seed-Coder 8B | 8B | Q4_K_M | 5.9 GB |
| MiniCPM-V 2.6 / MiniCPM-o 2.6 | 8B | Q4_K_M | 5.9 GB |
| Idefics 3 8B | 8B | Q4_K_M | 5.9 GB |
| Fuyu-8B | 8B | Q4_K_M | 5.9 GB |
| Emu3 | 8B | Q4_K_M | 5.9 GB |
| Stable Diffusion 3.5 Large / Turbo | 8B | Q4_K_M | 5.9 GB |
| EXAONE 3.5 2.4B / 7.8B | 7.8B | Q4_K_M | 5.7 GB |
| Mistral 7B | 7B | Q4_K_M | 5.7 GB |
| Qwen2.5 0.5B / 1.5B / 3B / 7B | 7B | Q5_K_M | 6 GB |
| OLMo 2 1B / 7B | 7B | Q5_K_M | 6 GB |
| Falcon 3 1B / 3B / 7B | 7B | Q5_K_M | 6 GB |
| Command R7B | 7B | Q5_K_M | 6 GB |
| OpenHermes 2.5 | 7B | Q5_K_M | 6 GB |
| Zephyr 7B Beta | 7B | Q5_K_M | 6 GB |
| OpenChat 3.5 | 7B | Q5_K_M | 6 GB |
| Starling LM 7B | 7B | Q5_K_M | 6 GB |
| Codestral Mamba 7B | 7B | Q5_K_M | 6 GB |
| CodeGemma 2B / 7B | 7B | Q5_K_M | 6 GB |
| aiXcoder-7B | 7B | Q5_K_M | 6 GB |
| Nxcode / CodeQwen 1.5 7B | 7B | Q5_K_M | 6 GB |
| Janus-Pro 1B / 7B | 7B | Q5_K_M | 6 GB |
| Ruyi-Mini-7B | 7B | Q5_K_M | 6 GB |
| Qwen2-Audio 7B | 7B | Q5_K_M | 6 GB |
| Qwen2.5-Omni 3B / 7B | 7B | Q5_K_M | 6 GB |
| YuE | 7B | Q5_K_M | 6 GB |
| Magicoder-S-DS 6.7B | 6.7B | Q5_K_M | 5.7 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Llama 3.1 8B | 8B | 6.4 GB needed | 8.4 GB |
| Chroma | 8.9B | 6.5 GB needed | 8.5 GB |
| Gemma 2 9B | 9B | 8 GB needed | 10 GB |
| Nemotron Nano 4B / 9B | 9B | 6.6 GB needed | 8.6 GB |
| GLM-4 9B / GLM-4.5-Air | 9B | 6.6 GB needed | 8.6 GB |
| Yi-Coder 1.5B / 9B | 9B | 6.6 GB needed | 8.6 GB |
| GLM-4-9B-Chat / CodeGeeX4 | 9B | 6.6 GB needed | 8.6 GB |
| GLM-4V-9B / GLM-4.1V-Thinking | 9B | 6.6 GB needed | 8.6 GB |
| Mochi 1 | 10B | 7.3 GB needed | 9.3 GB |
| Open-Sora 2.0 | 11B | 8.1 GB needed | 10.1 GB |
How to read this
The NVIDIA GeForce GTX 1660 Ti Laptop graphics card features 6 GB of GDDR6 video memory. This dedicated memory size is the most important factor when running local AI models. To get fast processing speeds, the entire active model must fit inside this 6 GB limit. If a model fits completely in your video memory, your graphics processor can run it at full speed.
The quant column shows the level of model compression used to save space. Quantization reduces the precision of model weights to make the file smaller. For example, the Q4_K_M quant represents a medium four bit compression level. The Q5_K_M quant represents a five bit compression level. These compression levels let you run larger models like the 8B or 7B variants while keeping high response quality.
Many excellent models fit completely within your 6 GB limit. You can run Granite 3.3 8B, Ministral 8B, InternLM 3 8B, OpenCoder 8B, Seed-Coder 8B, MiniCPM-V 2.6, Idefics 3 8B, Fuyu-8B, Emu3, and Stable Diffusion 3.5 Large at the Q4_K_M quant using 5.9 GB of memory. EXAONE 3.5 7.8B and Mistral 7B also fit at the Q4_K_M quant using 5.7 GB of memory.
You can run other models at the slightly higher Q5_K_M quant using exactly 6 GB of memory. This list includes Qwen2.5 7B, OLMo 2 7B, Falcon 3 7B, Command R7B, OpenHermes 2.5, Zephyr 7B Beta, OpenChat 3.5, Starling LM 7B, Codestral Mamba 7B, CodeGemma 7B, aiXcoder-7B, Nxcode, Janus-Pro 7B, Ruyi-Mini-7B, Qwen2-Audio 7B, Qwen2.5-Omni 7B, and YuE. Magicoder-S-DS 6.7B fits at the Q5_K_M quant using 5.7 GB of memory.
When a model is too large for 6 GB of video memory, you must use CPU offload. This method splits the model between your graphics card and your system RAM. CPU offload allows you to run Llama 3.1 8B, Chroma, Gemma 2 9B, Nemotron Nano 9B, GLM-4 9B, Yi-Coder 9B, GLM-4-9B-Chat, GLM-4V-9B, Mochi 1, and Open-Sora 2.0. Offloading prevents out of memory errors but it makes generation speeds much slower because system RAM is slower than GDDR6 memory.
You must also consider the 4k context limit when planning your memory use. The memory figures listed here are calculated using a standard 4000 token context window. If you increase the context window to process longer documents, the model will require more memory. This extra memory demand can push a fitting model over your 6 GB limit and force slow CPU offload.