Best local AI models for NVIDIA GTX 950M
4 GB DDR3. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Lumina-Next / Lumina-Image 2.0 | 5B | Q4_K_M | 3.7 GB |
| CogVideoX 2B / 5B | 5B | Q4_K_M | 3.7 GB |
| DeepSeek-VL2 | 4.5B | Q5_K_M | 3.8 GB |
| DeepFloyd IF | 4.3B | Q5_K_M | 3.7 GB |
| Phi-3.5-vision | 4.2B | Q5_K_M | 3.6 GB |
| Qwen3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 4 E4B | 4B | Q6_K | 3.9 GB |
| MiniCPM 3 4B | 4B | Q6_K | 3.9 GB |
| Danube 3 4B | 4B | Q6_K | 3.9 GB |
| Fish Speech 1.5 / OpenAudio S1 | 4B | Q6_K | 3.9 GB |
| Phi-4-mini-instruct | 3.8B | Q6_K | 3.7 GB |
| Phi-3.5 Mini | 3.8B | Q6_K | 3.7 GB |
| OmniGen / OmniGen2 | 3.8B | Q6_K | 3.7 GB |
| SD Cascade (Würstchen v3) | 3.6B | Q6_K | 3.5 GB |
| SDXL Turbo | 3.5B | Q6_K | 3.4 GB |
| SDXL Lightning | 3.5B | Q6_K | 3.4 GB |
| ACE-Step | 3.5B | Q6_K | 3.4 GB |
| MusicGen small/medium/large | 3.3B | Q6_K | 3.2 GB |
| SmolLM3 3B | 3B | Q8_0 | 3.8 GB |
| Replit Code v1.5 3B | 3B | Q8_0 | 3.8 GB |
| Kandinsky 3.1 | 3B | Q8_0 | 3.8 GB |
| Voxtral Mini / Small | 3B | Q8_0 | 3.8 GB |
| Orpheus TTS | 3B | Q8_0 | 3.8 GB |
| Higgs Audio v2 | 3B | Q8_0 | 3.8 GB |
| Allegro | 2.8B | Q8_0 | 3.6 GB |
| Open-Sora Plan | 2.7B | Q8_0 | 3.4 GB |
| LFM2 1.2B / 2.6B | 2.6B | Q8_0 | 3.3 GB |
| Playground v2.5 | 2.6B | Q8_0 | 3.3 GB |
| Stable Diffusion 3.5 Medium | 2.5B | Q8_0 | 3.2 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| Stable Diffusion XL | 3.417B | 4.1 GB needed | 6.1 GB |
| Phi-3 Mini | 3.8B | 4.4 GB needed | 6.4 GB |
| Phi-4-multimodal | 5.6B | 4.1 GB needed | 6.1 GB |
| Magicoder-S-DS 6.7B | 6.7B | 4.9 GB needed | 6.9 GB |
| Mistral 7B | 7B | 5.7 GB needed | 7.7 GB |
| Qwen2.5 0.5B / 1.5B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| OLMo 2 1B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Falcon 3 1B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Command R7B | 7B | 5.1 GB needed | 7.1 GB |
| OpenHermes 2.5 | 7B | 5.1 GB needed | 7.1 GB |
How to read this
The NVIDIA GTX 950M is an entry level mobile graphics card equipped with 4 GB of DDR3 video memory. This dedicated memory size determines which artificial intelligence models can run entirely on your graphics hardware. To run a model smoothly without system slowdowns, the model files and the active memory workspace must fit within this 4 GB limit.
The best quant column shows the optimal quantization level for each model. Quantization is a compression method that reduces the size of model weights. Using a Q4_K_M or Q5_K_M quant allows larger models to fit into the 4 GB video memory. Using a Q6_K or Q8_0 quant provides higher precision for smaller models because they have more free memory space.
For models that fit entirely in the graphics memory, Lumina-Next and Lumina-Image 2.0 at 5B parameters can run using the Q4_K_M quant which uses 3.7 GB. CogVideoX 2B or 5B also fits at 5B parameters using the Q4_K_M quant with 3.7 GB used. DeepSeek-VL2 at 4.5B parameters fits using the Q5_K_M quant with 3.8 GB used. Other options include Phi-3.5-vision at 4.2B parameters using the Q5_K_M quant with 3.6 GB used.
Several 4B parameter models like Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, Danube 3 4B, and Fish Speech 1.5 or OpenAudio S1 fit using the Q6_K quant with 3.9 GB used. Phi-4-mini-instruct, Phi-3.5 Mini, and OmniGen or OmniGen2 at 3.8B parameters use the Q6_K quant with 3.7 GB used. You can also run SD Cascade, SDXL Turbo, SDXL Lightning, and ACE-Step using the Q6_K quant. Smaller models like SmolLM3 3B, Replit Code v1.5 3B, and Kandinsky 3.1 run at the Q8_0 quant using 3.8 GB.
When a model exceeds the 4 GB video memory, you must use CPU offload. This process splits the workload between your graphics card and your system RAM. CPU offload allows you to run larger models like Mistral 7B which needs 5.7 GB at Q4_K_M and 7.7 GB system RAM. Other offload options include Qwen2.5 0.5B or 1.5B or 3B or 7B, OLMo 2 1B or 7B, Falcon 3 1B or 3B or 7B, Command R7B, and OpenHermes 2.5 which all need 5.1 GB at Q4_K_M and 7.1 GB system RAM.
Using CPU offload comes with a performance cost. Transferring data between the DDR3 video memory and the system RAM slows down the generation speed significantly. Additionally, running models close to the memory limit restricts your context window. Keeping your text prompts short is necessary because a long 4k context window requires extra memory space that may cause out of memory errors.