Best local AI models for NVIDIA GTX 960M
4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Lumina-Next / Lumina-Image 2.0 | 5B | Q4_K_M | 3.7 GB |
| CogVideoX 2B / 5B | 5B | Q4_K_M | 3.7 GB |
| DeepSeek-VL2 | 4.5B | Q5_K_M | 3.8 GB |
| DeepFloyd IF | 4.3B | Q5_K_M | 3.7 GB |
| Phi-3.5-vision | 4.2B | Q5_K_M | 3.6 GB |
| Qwen3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 4 E4B | 4B | Q6_K | 3.9 GB |
| MiniCPM 3 4B | 4B | Q6_K | 3.9 GB |
| Danube 3 4B | 4B | Q6_K | 3.9 GB |
| Fish Speech 1.5 / OpenAudio S1 | 4B | Q6_K | 3.9 GB |
| Phi-4-mini-instruct | 3.8B | Q6_K | 3.7 GB |
| Phi-3.5 Mini | 3.8B | Q6_K | 3.7 GB |
| OmniGen / OmniGen2 | 3.8B | Q6_K | 3.7 GB |
| SD Cascade (Würstchen v3) | 3.6B | Q6_K | 3.5 GB |
| SDXL Turbo | 3.5B | Q6_K | 3.4 GB |
| SDXL Lightning | 3.5B | Q6_K | 3.4 GB |
| ACE-Step | 3.5B | Q6_K | 3.4 GB |
| MusicGen small/medium/large | 3.3B | Q6_K | 3.2 GB |
| SmolLM3 3B | 3B | Q8_0 | 3.8 GB |
| Replit Code v1.5 3B | 3B | Q8_0 | 3.8 GB |
| Kandinsky 3.1 | 3B | Q8_0 | 3.8 GB |
| Voxtral Mini / Small | 3B | Q8_0 | 3.8 GB |
| Orpheus TTS | 3B | Q8_0 | 3.8 GB |
| Higgs Audio v2 | 3B | Q8_0 | 3.8 GB |
| Allegro | 2.8B | Q8_0 | 3.6 GB |
| Open-Sora Plan | 2.7B | Q8_0 | 3.4 GB |
| LFM2 1.2B / 2.6B | 2.6B | Q8_0 | 3.3 GB |
| Playground v2.5 | 2.6B | Q8_0 | 3.3 GB |
| Stable Diffusion 3.5 Medium | 2.5B | Q8_0 | 3.2 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| Stable Diffusion XL | 3.417B | 4.1 GB needed | 6.1 GB |
| Phi-3 Mini | 3.8B | 4.4 GB needed | 6.4 GB |
| Phi-4-multimodal | 5.6B | 4.1 GB needed | 6.1 GB |
| Magicoder-S-DS 6.7B | 6.7B | 4.9 GB needed | 6.9 GB |
| Mistral 7B | 7B | 5.7 GB needed | 7.7 GB |
| Qwen2.5 0.5B / 1.5B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| OLMo 2 1B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Falcon 3 1B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Command R7B | 7B | 5.1 GB needed | 7.1 GB |
| OpenHermes 2.5 | 7B | 5.1 GB needed | 7.1 GB |
How to read this
The NVIDIA GTX 960M is a mobile graphics card equipped with 4 GB of GDDR5 memory. This physical memory size determines the maximum size of the artificial intelligence models you can run locally. To fit a model entirely within this 4 GB limit, you must select the correct model size and quantization level. Running models locally on your hardware ensures complete privacy and removes any dependence on external cloud servers.
Quantization is a compression method that reduces the precision of model weights. The quant column indicates the optimal format for each model to maximize quality within your hardware limits. For example, a Q4_K_M quant uses approximately four bits per weight, while a Q8_0 quant uses eight bits per weight. Higher quantization levels like Q8_0 provide better output quality but require more memory, while lower levels like Q4_K_M allow larger models to fit into your video memory.
With 4 GB of video memory, the largest fully fitting models include Lumina-Next or Lumina-Image 2.0 at 5B using a Q4_K_M quant which consumes 3.7 GB of memory. CogVideoX 2B or 5B also fits at 5B using a Q4_K_M quant with 3.7 GB used. For vision tasks, DeepSeek-VL2 at 4.5B fits using a Q5_K_M quant with 3.8 GB used, and Phi-3.5-vision at 4.2B fits using a Q5_K_M quant with 3.6 GB used. Text models like Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, and Danube 3 4B all run at 4B using a Q6_K quant which uses 3.9 GB of memory.
Other models that fit entirely in your video memory include Phi-4-mini-instruct and Phi-3.5 Mini at 3.8B using a Q6_K quant with 3.7 GB used. Image generation models like SDXL Turbo and SDXL Lightning fit at 3.5B using a Q6_K quant with 3.4 GB used. For audio tasks, MusicGen small/medium/large fits at 3.3B using a Q6_K quant with 3.2 GB used. Smaller models like SmolLM3 3B and Replit Code v1.5 3B can run at a higher Q8_0 quant using 3.8 GB of memory.
When a model exceeds 4 GB of video memory, you must use CPU offload. This technique splits the model weights between your video memory and your system RAM. For these cases, we assume you have 32 GB of system RAM. For example, Mistral 7B requires 5.7 GB of video memory at a Q4_K_M quant and 7.7 GB of system RAM. Qwen2.5 0.5B / 1.5B / 3B / 7B at the 7B size requires 5.1 GB of video memory at a Q4_K_M quant and 7.1 GB of system RAM. Offloading allows you to run larger models, but it costs processing speed because transferring data between system RAM and video memory is slow.
You must also consider the context window limit when running these models. The memory figures listed here are calculated using a basic 4k context window. If you increase the context window to process longer documents or chat histories, the memory usage will rise significantly. To prevent out of memory errors on your 4 GB card, you should keep your context window at 4k or lower.