Best local AI models for NVIDIA Quadro K2000M
2 GB DDR3. At a 4k context, 56 of the 233 models in our catalog with verified parameter counts fit fully, up to Allegro at 2.8B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 56 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Allegro | 2.8B | Q4_K_M | 2 GB |
| Open-Sora Plan | 2.7B | Q4_K_M | 2 GB |
| LFM2 1.2B / 2.6B | 2.6B | Q4_K_M | 1.9 GB |
| Playground v2.5 | 2.6B | Q4_K_M | 1.9 GB |
| Stable Diffusion 3.5 Medium | 2.5B | Q4_K_M | 1.8 GB |
| Canary 1B / Qwen-2.5B | 2.5B | Q4_K_M | 1.8 GB |
| SeamlessM4T v2 | 2.3B | Q5_K_M | 2 GB |
| Parler-TTS | 2.2B | Q5_K_M | 1.9 GB |
| Kimi K3 DSpark | 2.2B | Q5_K_M | 2 GB |
| SmolVLM 256M / 500M / 2B | 2B | Q6_K | 2 GB |
| Stable Diffusion 3 Medium | 2B | Q6_K | 2 GB |
| Pyramid Flow | 2B | Q6_K | 2 GB |
| Wav2Vec2 / XLS-R | 2B | Q6_K | 2 GB |
| Moondream 2 | 1.9B | Q6_K | 1.9 GB |
| Qwen3 1.7B | 1.7B | Q6_K | 1.7 GB |
| SmolLM2 135M / 360M / 1.7B | 1.7B | Q6_K | 1.7 GB |
| StableLM 2 1.6B | 1.6B | Q8_0 | 2 GB |
| Sana 0.6B / 1.6B | 1.6B | Q8_0 | 2 GB |
| Zonos 0.1 | 1.6B | Q8_0 | 2 GB |
| Dia 1.6B | 1.6B | Q8_0 | 2 GB |
| Whisper Large v3 | 1.55B | Q8_0 | 2 GB |
| ControlNet / T2I-Adapter / IP-Adapter | 1.5B | Q8_0 | 1.9 GB |
| Hunyuan-DiT | 1.5B | Q8_0 | 1.9 GB |
| Stable Video Diffusion | 1.5B | Q8_0 | 1.9 GB |
| Whisper Large v2 / turbo | 1.5B | Q8_0 | 1.9 GB |
| AudioGen | 1.5B | Q8_0 | 1.9 GB |
| AudioLDM 2 | 1.5B | Q8_0 | 1.9 GB |
| Tango 2 | 1.4B | Q8_0 | 1.8 GB |
| TinyLlama 1.1B | 1.1B | Q8_0 | 1.4 GB |
| SantaCoder 1.1B | 1.1B | Q8_0 | 1.4 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| SmolLM3 3B | 3B | 2.2 GB needed | 4.2 GB |
| Replit Code v1.5 3B | 3B | 2.2 GB needed | 4.2 GB |
| Kandinsky 3.1 | 3B | 2.2 GB needed | 4.2 GB |
| Voxtral Mini / Small | 3B | 2.2 GB needed | 4.2 GB |
| Orpheus TTS | 3B | 2.2 GB needed | 4.2 GB |
| Higgs Audio v2 | 3B | 2.2 GB needed | 4.2 GB |
| MusicGen small/medium/large | 3.3B | 2.4 GB needed | 4.4 GB |
| Stable Diffusion XL | 3.417B | 4.1 GB needed | 6.1 GB |
| SDXL Turbo | 3.5B | 2.6 GB needed | 4.6 GB |
| SDXL Lightning | 3.5B | 2.6 GB needed | 4.6 GB |
How to read this
The NVIDIA Quadro K2000M is an older mobile workstation graphics card equipped with 2 GB of DDR3 memory. This dedicated video memory capacity determines which artificial intelligence models can run directly on the hardware. Because the onboard memory is limited to 2 GB, selecting the correct model size and quantization level is essential to prevent out of memory errors during inference.
Quantization is a compression method that reduces the size of neural networks. The quant column shows the best format for each model to fit within your hardware limits. Formats like Q4_K_M and Q5_K_M compress the model weights to four or five bits. Formats like Q6_K and Q8_0 preserve more precision but require more memory. Using these optimized quants allows you to run larger architectures on limited hardware.
Several compact models can run entirely within the 2 GB video memory limit. The Allegro 2.8B model fits using the Q4_K_M quant which uses exactly 2 GB of memory. The Open-Sora Plan 2.7B model also fits using the Q4_K_M quant with 2 GB used. For slightly smaller footprints, the LFM2 2.6B and Playground v2.5 2.6B models require 1.9 GB of memory at the Q4_K_M quantization level. Stable Diffusion 3.5 Medium 2.5B and Canary 2.5B both use 1.8 GB of memory with the Q4_K_M quant.
Highly compressed smaller models can utilize higher precision quants. The SmolVLM 2B, Stable Diffusion 3 Medium 2B, Pyramid Flow 2B, and Wav2Vec2 2B models all run at the Q6_K quantization level using 2 GB of memory. If you require maximum precision, the TinyLlama 1.1B and SantaCoder 1.1B models can run at the Q8_0 quantization level. These models use 1.4 GB of video memory which leaves a small buffer.
When a model exceeds the 2 GB video memory limit, you must use CPU offloading. This process splits the model weights between your graphics card and your system memory. For example, the SmolLM3 3B model needs 2.2 GB of memory at Q4_K_M which requires 4.2 GB of system RAM to offload the overflow. Stable Diffusion XL requires 4.1 GB at FP8 which needs 6.1 GB of system RAM. Offloading allows you to run larger models but it reduces processing speed significantly.
Running language models with a standard 4k context window increases memory consumption during active use. The listed memory figures represent the base model size. Generating long responses or processing large text prompts will require additional memory beyond these base figures. For the best stability on the Quadro K2000M, you should monitor your memory usage and keep your context windows short.