Best local AI models for NVIDIA Quadro P2200

5 GB GDDR5X. At a 4k context, 85 of the 233 models in our catalog with verified parameter counts fit fully, up to Magicoder-S-DS 6.7B at 6.7B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 85 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Magicoder-S-DS 6.7B6.7BQ4_K_M4.9 GB
Phi-4-multimodal5.6BQ5_K_M4.8 GB
Lumina-Next / Lumina-Image 2.05BQ6_K4.9 GB
CogVideoX 2B / 5B5BQ6_K4.9 GB
DeepSeek-VL24.5BQ6_K4.4 GB
DeepFloyd IF4.3BQ6_K4.2 GB
Phi-3.5-vision4.2BQ6_K4.1 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-3 Mini3.8BQ5_K_M4.8 GB
Phi-4-mini-instruct3.8BQ8_04.8 GB
Phi-3.5 Mini3.8BQ8_04.8 GB
OmniGen / OmniGen23.8BQ8_04.8 GB
SD Cascade (Würstchen v3)3.6BQ8_04.6 GB
SDXL Turbo3.5BQ8_04.5 GB
SDXL Lightning3.5BQ8_04.5 GB
ACE-Step3.5BQ8_04.5 GB
Stable Diffusion XL3.417BFP8 / optimized4.1 GB
MusicGen small/medium/large3.3BQ8_04.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at Q4_K_MSystem RAM at 4k
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB
Zephyr 7B Beta7B5.1 GB needed7.1 GB
OpenChat 3.57B5.1 GB needed7.1 GB
Starling LM 7B7B5.1 GB needed7.1 GB
Codestral Mamba 7B7B5.1 GB needed7.1 GB

How to read this

The NVIDIA Quadro P2200 is equipped with 5 GB of GDDR5X frame buffer memory. This dedicated video memory determines the size of the artificial intelligence models you can run locally. To fit inside this hardware limit, models must be compressed using quantization. Quantization reduces the precision of model weights to save space. The best quantization column shows the highest quality format that fits entirely within your graphics card memory.

For local execution without slowdowns, the model and its working memory must fit under the 5 GB limit. The Magicoder-S-DS 6.7B model fits at a Q4_K_M quantization which uses 4.9 GB of memory. The Phi-4-multimodal model fits at Q5_K_M quantization using 4.8 GB of memory. You can also run Lumina-Next or Lumina-Image 2.0 at Q6_K quantization using 4.9 GB of memory. CogVideoX 2B or 5B also fits at Q6_K quantization using 4.9 GB of memory.

Several highly capable models fit within the 4 GB range. DeepSeek-VL2 fits at Q6_K quantization using 4.4 GB of memory. DeepFloyd IF fits at Q6_K quantization using 4.2 GB of memory. Phi-3.5-vision fits at Q6_K quantization using 4.1 GB of memory. Models like Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, Danube 3 4B, and Fish Speech 1.5 or OpenAudio S1 all fit at Q6_K quantization using 3.9 GB of memory.

Highly optimized 3.8B models can run at Q8_0 quantization which offers excellent precision. These include Phi-4-mini-instruct, Phi-3.5 Mini, and OmniGen or OmniGen2 which all use 4.8 GB of memory. Phi-3 Mini also uses 4.8 GB of memory but at Q5_K_M quantization. Image and audio models like SD Cascade (Würstchen v3) use 4.6 GB of memory at Q8_0 quantization. SDXL Turbo and SDXL Lightning use 4.5 GB of memory at Q8_0 quantization. Stable Diffusion XL uses 4.1 GB of memory in FP8 or optimized format.

When a model is too large for the 5 GB video memory, you can use CPU offload if you have 32 GB of system RAM. Offloading splits the model between your graphics card and system memory. This allows you to run larger 7B models but it reduces processing speed. For example, Mistral 7B needs 5.7 GB at Q4_K_M quantization which requires 7.7 GB of system RAM. Other 7B models like Qwen2.5, OLMo 2, Falcon 3, Command R7B, OpenHermes 2.5, Zephyr 7B Beta, OpenChat 3.5, Starling LM 7B, and Codestral Mamba 7B need 5.1 GB at Q4_K_M quantization which requires 7.1 GB of system RAM.

You must consider the 4k context window caveat when running these models. The memory numbers listed only cover loading the model weights. As you type longer prompts and the model generates longer answers, the active context memory grows. Running models close to the 5 GB limit of your NVIDIA Quadro P2200 will leave very little room for long conversations or large image generation tasks.