Best local AI models for NVIDIA GTX 680M

4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Lumina-Next / Lumina-Image 2.05BQ4_K_M3.7 GB
CogVideoX 2B / 5B5BQ4_K_M3.7 GB
DeepSeek-VL24.5BQ5_K_M3.8 GB
DeepFloyd IF4.3BQ5_K_M3.7 GB
Phi-3.5-vision4.2BQ5_K_M3.6 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-4-mini-instruct3.8BQ6_K3.7 GB
Phi-3.5 Mini3.8BQ6_K3.7 GB
OmniGen / OmniGen23.8BQ6_K3.7 GB
SD Cascade (Würstchen v3)3.6BQ6_K3.5 GB
SDXL Turbo3.5BQ6_K3.4 GB
SDXL Lightning3.5BQ6_K3.4 GB
ACE-Step3.5BQ6_K3.4 GB
MusicGen small/medium/large3.3BQ6_K3.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB
Open-Sora Plan2.7BQ8_03.4 GB
LFM2 1.2B / 2.6B2.6BQ8_03.3 GB
Playground v2.52.6BQ8_03.3 GB
Stable Diffusion 3.5 Medium2.5BQ8_03.2 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB

How to read this

The NVIDIA GTX 680M is a mobile graphics card equipped with 4 GB of GDDR5 memory. This physical memory size determines the maximum size of the artificial intelligence models you can run locally. To fit inside this hardware limit, models must be compressed using quantization. The best quant column shows the highest quality quantization level that fits within your video memory without causing out of memory errors.

For models running entirely on the graphics card, you can use options up to 5B parameters. Lumina-Next or Lumina-Image 2.0 and CogVideoX 2B or 5B fit at the Q4_K_M quantization level using 3.7 GB of memory. DeepSeek-VL2 fits at the Q5_K_M quantization level using 3.8 GB of memory. DeepFloyd IF fits at Q5_K_M using 3.7 GB of memory. Phi-3.5-vision fits at Q5_K_M using 3.6 GB of memory.

Several 4B parameter models run on this hardware at the Q6_K quantization level using 3.9 GB of memory. These models include Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, Danube 3 4B, and Fish Speech 1.5 or OpenAudio S1. You can also run Phi-4-mini-instruct, Phi-3.5 Mini, and OmniGen or OmniGen2 at Q6_K using 3.7 GB of memory. SD Cascade (Würstchen v3) runs at Q6_K using 3.5 GB of memory. SDXL Turbo, SDXL Lightning, and ACE-Step run at Q6_K using 3.4 GB of memory. MusicGen small/medium/large runs at Q6_K using 3.2 GB of memory.

Smaller models can run at the higher Q8_0 quantization level. SmolLM3 3B, Replit Code v1.5 3B, Kandinsky 3.1, Voxtral Mini or Small, Orpheus TTS, and Higgs Audio v2 use 3.8 GB of memory. Allegro runs at Q8_0 using 3.6 GB of memory. Open-Sora Plan uses 3.4 GB of memory. LFM2 1.2B or 2.6B and Playground v2.5 use 3.3 GB of memory. Stable Diffusion 3.5 Medium uses 3.2 GB of memory.

When a model is too large for the 4 GB video memory, you can offload parts of it to your system RAM. This requires a system with 32 GB of system RAM. Offloading allows you to run larger models but it costs processing speed because data must travel between the system RAM and the graphics card. For example, Stable Diffusion XL needs 4.1 GB at FP8 or optimized settings and uses 6.1 GB of system RAM. Phi-3 Mini needs 4.4 GB at Q4_K_M and uses 6.4 GB of system RAM. Phi-4-multimodal needs 4.1 GB at Q4_K_M and uses 6.1 GB of system RAM. Magicoder-S-DS 6.7B needs 4.9 GB at Q4_K_M and uses 6.9 GB of system RAM.

Other offload options include Mistral 7B which needs 5.7 GB at Q4_K_M and uses 7.7 GB of system RAM. Qwen2.5 0.5B or 1.5B or 3B or 7B, OLMo 2 1B or 7B, Falcon 3 1B or 3B or 7B, Command R7B, and OpenHermes 2.5 all need 5.1 GB at Q4_K_M and use 7.1 GB of system RAM. Keep in mind that active memory usage increases during processing. Running text models at a standard 4k context window requires additional memory for the context history which can exceed your limits if your base model already uses almost all the video memory.