Best local AI models for NVIDIA GTX 1650 Ti Laptop

4 GB GDDR6. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Lumina-Next / Lumina-Image 2.05BQ4_K_M3.7 GB
CogVideoX 2B / 5B5BQ4_K_M3.7 GB
DeepSeek-VL24.5BQ5_K_M3.8 GB
DeepFloyd IF4.3BQ5_K_M3.7 GB
Phi-3.5-vision4.2BQ5_K_M3.6 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-4-mini-instruct3.8BQ6_K3.7 GB
Phi-3.5 Mini3.8BQ6_K3.7 GB
OmniGen / OmniGen23.8BQ6_K3.7 GB
SD Cascade (Würstchen v3)3.6BQ6_K3.5 GB
SDXL Turbo3.5BQ6_K3.4 GB
SDXL Lightning3.5BQ6_K3.4 GB
ACE-Step3.5BQ6_K3.4 GB
MusicGen small/medium/large3.3BQ6_K3.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB
Open-Sora Plan2.7BQ8_03.4 GB
LFM2 1.2B / 2.6B2.6BQ8_03.3 GB
Playground v2.52.6BQ8_03.3 GB
Stable Diffusion 3.5 Medium2.5BQ8_03.2 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB

How to read this

The NVIDIA GeForce GTX 1650 Ti Laptop GPU features 4 GB of GDDR6 video memory. This dedicated memory size determines which artificial intelligence models can run entirely on your graphics hardware. When a model fits completely within this 4 GB limit, it processes tokens and generates outputs at the fastest speeds your hardware can support.

To fit models into this memory limit, developers use quantization. The quant column shows the specific compression level applied to each model. For example, a Q6_K quant represents a high quality six bit quantization that preserves most of the original model accuracy. A Q4_K_M quant uses four bit quantization to compress larger models like the 5B Lumina-Next or CogVideoX 2B / 5B down to 3.7 GB of memory usage.

For models that exceed your 4 GB of video memory, you can use CPU offload. This technique splits the model layers between your GPU and your system RAM. If your laptop has 32 GB of system RAM, you can run larger models like Mistral 7B. This model requires 5.7 GB of memory at Q4_K_M, which uses your 4 GB of video memory and 7.7 GB of system RAM. Offloading allows you to run these larger models, but it reduces processing speed because system RAM is slower than GDDR6 memory.

Several high quality models fit directly into your video memory without offloading. The Phi-4-mini-instruct and Phi-3.5 Mini models both use 3.7 GB of memory at the Q6_K quant. Image generation models like SDXL Turbo and SDXL Lightning also fit well, requiring 3.4 GB of video memory at the Q6_K quant. For audio tasks, the Higgs Audio v2 and Orpheus TTS models fit within 3.8 GB of memory using the Q8_0 quant.

When running these models, you must monitor your context window size. The memory figures listed are for the base model. Running a model with a long context window of 4k tokens or more will require additional video memory for the context cache. If your context cache exceeds your remaining video memory, your system will automatically offload data to system RAM and slow down performance.