Best local AI models for NVIDIA Quadro K2000M

2 GB DDR3. At a 4k context, 56 of the 233 models in our catalog with verified parameter counts fit fully, up to Allegro at 2.8B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 56 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Allegro2.8BQ4_K_M2 GB
Open-Sora Plan2.7BQ4_K_M2 GB
LFM2 1.2B / 2.6B2.6BQ4_K_M1.9 GB
Playground v2.52.6BQ4_K_M1.9 GB
Stable Diffusion 3.5 Medium2.5BQ4_K_M1.8 GB
Canary 1B / Qwen-2.5B2.5BQ4_K_M1.8 GB
SeamlessM4T v22.3BQ5_K_M2 GB
Parler-TTS2.2BQ5_K_M1.9 GB
Kimi K3 DSpark2.2BQ5_K_M2 GB
SmolVLM 256M / 500M / 2B2BQ6_K2 GB
Stable Diffusion 3 Medium2BQ6_K2 GB
Pyramid Flow2BQ6_K2 GB
Wav2Vec2 / XLS-R2BQ6_K2 GB
Moondream 21.9BQ6_K1.9 GB
Qwen3 1.7B1.7BQ6_K1.7 GB
SmolLM2 135M / 360M / 1.7B1.7BQ6_K1.7 GB
StableLM 2 1.6B1.6BQ8_02 GB
Sana 0.6B / 1.6B1.6BQ8_02 GB
Zonos 0.11.6BQ8_02 GB
Dia 1.6B1.6BQ8_02 GB
Whisper Large v31.55BQ8_02 GB
ControlNet / T2I-Adapter / IP-Adapter1.5BQ8_01.9 GB
Hunyuan-DiT1.5BQ8_01.9 GB
Stable Video Diffusion1.5BQ8_01.9 GB
Whisper Large v2 / turbo1.5BQ8_01.9 GB
AudioGen1.5BQ8_01.9 GB
AudioLDM 21.5BQ8_01.9 GB
Tango 21.4BQ8_01.8 GB
TinyLlama 1.1B1.1BQ8_01.4 GB
SantaCoder 1.1B1.1BQ8_01.4 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at Q4_K_MSystem RAM at 4k
SmolLM3 3B3B2.2 GB needed4.2 GB
Replit Code v1.5 3B3B2.2 GB needed4.2 GB
Kandinsky 3.13B2.2 GB needed4.2 GB
Voxtral Mini / Small3B2.2 GB needed4.2 GB
Orpheus TTS3B2.2 GB needed4.2 GB
Higgs Audio v23B2.2 GB needed4.2 GB
MusicGen small/medium/large3.3B2.4 GB needed4.4 GB
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
SDXL Turbo3.5B2.6 GB needed4.6 GB
SDXL Lightning3.5B2.6 GB needed4.6 GB

How to read this

The NVIDIA Quadro K2000M is an older mobile workstation graphics card equipped with 2 GB of DDR3 memory. This dedicated video memory capacity determines which artificial intelligence models can run directly on the hardware. Because the onboard memory is limited to 2 GB, selecting the correct model size and quantization level is essential to prevent out of memory errors during inference.

Quantization is a compression method that reduces the size of neural networks. The quant column shows the best format for each model to fit within your hardware limits. Formats like Q4_K_M and Q5_K_M compress the model weights to four or five bits. Formats like Q6_K and Q8_0 preserve more precision but require more memory. Using these optimized quants allows you to run larger architectures on limited hardware.

Several compact models can run entirely within the 2 GB video memory limit. The Allegro 2.8B model fits using the Q4_K_M quant which uses exactly 2 GB of memory. The Open-Sora Plan 2.7B model also fits using the Q4_K_M quant with 2 GB used. For slightly smaller footprints, the LFM2 2.6B and Playground v2.5 2.6B models require 1.9 GB of memory at the Q4_K_M quantization level. Stable Diffusion 3.5 Medium 2.5B and Canary 2.5B both use 1.8 GB of memory with the Q4_K_M quant.

Highly compressed smaller models can utilize higher precision quants. The SmolVLM 2B, Stable Diffusion 3 Medium 2B, Pyramid Flow 2B, and Wav2Vec2 2B models all run at the Q6_K quantization level using 2 GB of memory. If you require maximum precision, the TinyLlama 1.1B and SantaCoder 1.1B models can run at the Q8_0 quantization level. These models use 1.4 GB of video memory which leaves a small buffer.

When a model exceeds the 2 GB video memory limit, you must use CPU offloading. This process splits the model weights between your graphics card and your system memory. For example, the SmolLM3 3B model needs 2.2 GB of memory at Q4_K_M which requires 4.2 GB of system RAM to offload the overflow. Stable Diffusion XL requires 4.1 GB at FP8 which needs 6.1 GB of system RAM. Offloading allows you to run larger models but it reduces processing speed significantly.

Running language models with a standard 4k context window increases memory consumption during active use. The listed memory figures represent the base model size. Generating long responses or processing large text prompts will require additional memory beyond these base figures. For the best stability on the Quadro K2000M, you should monitor your memory usage and keep your context windows short.