Best local AI models for NVIDIA Quadro K2200

4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Lumina-Next / Lumina-Image 2.05BQ4_K_M3.7 GB
CogVideoX 2B / 5B5BQ4_K_M3.7 GB
DeepSeek-VL24.5BQ5_K_M3.8 GB
DeepFloyd IF4.3BQ5_K_M3.7 GB
Phi-3.5-vision4.2BQ5_K_M3.6 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-4-mini-instruct3.8BQ6_K3.7 GB
Phi-3.5 Mini3.8BQ6_K3.7 GB
OmniGen / OmniGen23.8BQ6_K3.7 GB
SD Cascade (Würstchen v3)3.6BQ6_K3.5 GB
SDXL Turbo3.5BQ6_K3.4 GB
SDXL Lightning3.5BQ6_K3.4 GB
ACE-Step3.5BQ6_K3.4 GB
MusicGen small/medium/large3.3BQ6_K3.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB
Open-Sora Plan2.7BQ8_03.4 GB
LFM2 1.2B / 2.6B2.6BQ8_03.3 GB
Playground v2.52.6BQ8_03.3 GB
Stable Diffusion 3.5 Medium2.5BQ8_03.2 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB

How to read this

The NVIDIA Quadro K2200 is equipped with 4 GB GDDR5 memory. This physical memory size limits the size of the artificial intelligence models you can run entirely on the graphics card. To run a model without system slowdowns, the model weights and the working memory must fit within this 4 GB boundary.

The quantization column shows the compression level used to shrink these models. Quantization formats like Q4_K_M, Q6_K, and Q8_0 represent different levels of precision. A Q4_K_M quant uses fewer bits per parameter to fit larger models like the 5B Lumina-Next or CogVideoX into 3.7 GB of graphics memory. A Q8_0 quant keeps higher precision but requires more memory per parameter, which is why the smaller 3B SmolLM3 uses 3.8 GB of memory.

When a model exceeds the 4 GB physical limit of your graphics card, you must use CPU offload. This technique splits the workload between your graphics card and your system RAM. For example, running Mistral 7B at Q4_K_M requires 5.7 GB of memory, which uses your graphics card and 7.7 GB of system RAM on a 32 GB system. Offloading allows you to run larger models like Falcon 3 7B or Command R7B, but it costs processing speed because data transfers slowly between the system RAM and the GPU.

You can run several vision and audio models locally within the hardware limits. The DeepSeek-VL2 model fits at Q5_K_M quant using 3.8 GB of memory. Image generators like SDXL Turbo and SDXL Lightning fit at Q6_K quant using 3.4 GB of memory. For audio tasks, MusicGen fits at Q6_K quant using 3.2 GB of memory.

Context window size also affects your memory usage. The memory figures listed here assume a standard 4k context window. If you increase the context window to process longer documents, the system requires more working memory. This extra demand can push a model that fits at 4k context over the 4 GB limit, which forces the system to offload data to your system RAM.