Best local AI models for NVIDIA GTX 1060 MAX-Q

6 GB GDDR5. At a 4k context, 114 of the 233 models in our catalog with verified parameter counts fit fully, up to Granite 3.3 2B / 8B at 8B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 114 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Granite 3.3 2B / 8B8BQ4_K_M5.9 GB
Ministral 3B / 8B8BQ4_K_M5.9 GB
InternLM 3 8B8BQ4_K_M5.9 GB
OpenCoder 1.5B / 8B8BQ4_K_M5.9 GB
Seed-Coder 8B8BQ4_K_M5.9 GB
MiniCPM-V 2.6 / MiniCPM-o 2.68BQ4_K_M5.9 GB
Idefics 3 8B8BQ4_K_M5.9 GB
Fuyu-8B8BQ4_K_M5.9 GB
Emu38BQ4_K_M5.9 GB
Stable Diffusion 3.5 Large / Turbo8BQ4_K_M5.9 GB
EXAONE 3.5 2.4B / 7.8B7.8BQ4_K_M5.7 GB
Mistral 7B7BQ4_K_M5.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7BQ5_K_M6 GB
OLMo 2 1B / 7B7BQ5_K_M6 GB
Falcon 3 1B / 3B / 7B7BQ5_K_M6 GB
Command R7B7BQ5_K_M6 GB
OpenHermes 2.57BQ5_K_M6 GB
Zephyr 7B Beta7BQ5_K_M6 GB
OpenChat 3.57BQ5_K_M6 GB
Starling LM 7B7BQ5_K_M6 GB
Codestral Mamba 7B7BQ5_K_M6 GB
CodeGemma 2B / 7B7BQ5_K_M6 GB
aiXcoder-7B7BQ5_K_M6 GB
Nxcode / CodeQwen 1.5 7B7BQ5_K_M6 GB
Janus-Pro 1B / 7B7BQ5_K_M6 GB
Ruyi-Mini-7B7BQ5_K_M6 GB
Qwen2-Audio 7B7BQ5_K_M6 GB
Qwen2.5-Omni 3B / 7B7BQ5_K_M6 GB
YuE7BQ5_K_M6 GB
Magicoder-S-DS 6.7B6.7BQ5_K_M5.7 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at Q4_K_MSystem RAM at 4k
Llama 3.1 8B8B6.4 GB needed8.4 GB
Chroma8.9B6.5 GB needed8.5 GB
Gemma 2 9B9B8 GB needed10 GB
Nemotron Nano 4B / 9B9B6.6 GB needed8.6 GB
GLM-4 9B / GLM-4.5-Air9B6.6 GB needed8.6 GB
Yi-Coder 1.5B / 9B9B6.6 GB needed8.6 GB
GLM-4-9B-Chat / CodeGeeX49B6.6 GB needed8.6 GB
GLM-4V-9B / GLM-4.1V-Thinking9B6.6 GB needed8.6 GB
Mochi 110B7.3 GB needed9.3 GB
Open-Sora 2.011B8.1 GB needed10.1 GB

How to read this

The NVIDIA GTX 1060 MAX-Q is a laptop graphics card equipped with 6 GB GDDR5 memory. This dedicated video memory determines the maximum size of the artificial intelligence models you can run entirely on your graphics hardware. Running models completely within this video memory ensures the fastest possible processing speeds for your local tasks.

To fit larger models into the 6 GB limit you must use quantized versions. The quant column indicates the compression level applied to the model weights. For example a Q4_K_M quant represents a medium four bit quantization while a Q5_K_M quant represents a five bit quantization. These compressed formats significantly reduce memory usage while preserving most of the original model intelligence.

Several high quality models fit directly into your video memory. The Granite 3.3 8B Ministral 8B InternLM 3 8B OpenCoder 8B Seed-Coder 8B MiniCPM-V 2.6 MiniCPM-o 2.6 Idefics 3 8B Fuyu-8B Emu3 and Stable Diffusion 3.5 Large or Turbo can all run at the Q4_K_M quant using 5.9 GB of memory. The EXAONE 3.5 7.8B and Mistral 7B also fit well at Q4_K_M using 5.7 GB of memory.

Other models utilize a Q5_K_M quant to fit exactly within the 6 GB limit. These include Qwen2.5 7B OLMo 2 7B Falcon 3 7B Command R7B OpenHermes 2.5 Zephyr 7B Beta OpenChat 3.5 Starling LM 7B Codestral Mamba 7B CodeGemma 7B aiXcoder-7B Nxcode or CodeQwen 1.5 7B Janus-Pro 7B Ruyi-Mini-7B Qwen2-Audio 7B Qwen2.5-Omni 7B and YuE. The Magicoder-S-DS 6.7B also fits at Q5_K_M using 5.7 GB of memory.

If you want to run larger models you must offload some layers to your system RAM. Assuming you have 32 GB of system RAM you can run Llama 3.1 8B which needs 6.4 GB at Q4_K_M and 8.4 GB of system RAM. You can also run Chroma 8.9B Nemotron Nano 9B GLM-4 9B Yi-Coder 9B GLM-4-9B-Chat or CodeGeeX4 and GLM-4V or GLM-4.1V-Thinking which all need 6.6 GB at Q4_K_M and 8.6 GB of system RAM. Gemma 2 9B needs 8 GB at Q4_K_M and 10 GB of system RAM. Mochi 1 needs 7.3 GB at Q4_K_M and 9.3 GB of system RAM while Open-Sora 2.0 needs 8.1 GB at Q4_K_M and 10.1 GB of system RAM.

Offloading layers to system RAM comes with a performance cost. The transfer of data between your system memory and your graphics card is much slower than using dedicated video memory alone. Additionally these memory calculations are based on a standard 4k context window. If you increase the context window to process longer texts the memory usage will rise and may exceed your limits.