Best local AI models for NVIDIA Quadro P2200
5 GB GDDR5X. At a 4k context, 85 of the 233 models in our catalog with verified parameter counts fit fully, up to Magicoder-S-DS 6.7B at 6.7B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 85 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Magicoder-S-DS 6.7B | 6.7B | Q4_K_M | 4.9 GB |
| Phi-4-multimodal | 5.6B | Q5_K_M | 4.8 GB |
| Lumina-Next / Lumina-Image 2.0 | 5B | Q6_K | 4.9 GB |
| CogVideoX 2B / 5B | 5B | Q6_K | 4.9 GB |
| DeepSeek-VL2 | 4.5B | Q6_K | 4.4 GB |
| DeepFloyd IF | 4.3B | Q6_K | 4.2 GB |
| Phi-3.5-vision | 4.2B | Q6_K | 4.1 GB |
| Qwen3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 4 E4B | 4B | Q6_K | 3.9 GB |
| MiniCPM 3 4B | 4B | Q6_K | 3.9 GB |
| Danube 3 4B | 4B | Q6_K | 3.9 GB |
| Fish Speech 1.5 / OpenAudio S1 | 4B | Q6_K | 3.9 GB |
| Phi-3 Mini | 3.8B | Q5_K_M | 4.8 GB |
| Phi-4-mini-instruct | 3.8B | Q8_0 | 4.8 GB |
| Phi-3.5 Mini | 3.8B | Q8_0 | 4.8 GB |
| OmniGen / OmniGen2 | 3.8B | Q8_0 | 4.8 GB |
| SD Cascade (Würstchen v3) | 3.6B | Q8_0 | 4.6 GB |
| SDXL Turbo | 3.5B | Q8_0 | 4.5 GB |
| SDXL Lightning | 3.5B | Q8_0 | 4.5 GB |
| ACE-Step | 3.5B | Q8_0 | 4.5 GB |
| Stable Diffusion XL | 3.417B | FP8 / optimized | 4.1 GB |
| MusicGen small/medium/large | 3.3B | Q8_0 | 4.2 GB |
| SmolLM3 3B | 3B | Q8_0 | 3.8 GB |
| Replit Code v1.5 3B | 3B | Q8_0 | 3.8 GB |
| Kandinsky 3.1 | 3B | Q8_0 | 3.8 GB |
| Voxtral Mini / Small | 3B | Q8_0 | 3.8 GB |
| Orpheus TTS | 3B | Q8_0 | 3.8 GB |
| Higgs Audio v2 | 3B | Q8_0 | 3.8 GB |
| Allegro | 2.8B | Q8_0 | 3.6 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Mistral 7B | 7B | 5.7 GB needed | 7.7 GB |
| Qwen2.5 0.5B / 1.5B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| OLMo 2 1B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Falcon 3 1B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Command R7B | 7B | 5.1 GB needed | 7.1 GB |
| OpenHermes 2.5 | 7B | 5.1 GB needed | 7.1 GB |
| Zephyr 7B Beta | 7B | 5.1 GB needed | 7.1 GB |
| OpenChat 3.5 | 7B | 5.1 GB needed | 7.1 GB |
| Starling LM 7B | 7B | 5.1 GB needed | 7.1 GB |
| Codestral Mamba 7B | 7B | 5.1 GB needed | 7.1 GB |
How to read this
The NVIDIA Quadro P2200 is equipped with 5 GB of GDDR5X frame buffer memory. This dedicated video memory determines the size of the artificial intelligence models you can run locally. To fit inside this hardware limit, models must be compressed using quantization. Quantization reduces the precision of model weights to save space. The best quantization column shows the highest quality format that fits entirely within your graphics card memory.
For local execution without slowdowns, the model and its working memory must fit under the 5 GB limit. The Magicoder-S-DS 6.7B model fits at a Q4_K_M quantization which uses 4.9 GB of memory. The Phi-4-multimodal model fits at Q5_K_M quantization using 4.8 GB of memory. You can also run Lumina-Next or Lumina-Image 2.0 at Q6_K quantization using 4.9 GB of memory. CogVideoX 2B or 5B also fits at Q6_K quantization using 4.9 GB of memory.
Several highly capable models fit within the 4 GB range. DeepSeek-VL2 fits at Q6_K quantization using 4.4 GB of memory. DeepFloyd IF fits at Q6_K quantization using 4.2 GB of memory. Phi-3.5-vision fits at Q6_K quantization using 4.1 GB of memory. Models like Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, Danube 3 4B, and Fish Speech 1.5 or OpenAudio S1 all fit at Q6_K quantization using 3.9 GB of memory.
Highly optimized 3.8B models can run at Q8_0 quantization which offers excellent precision. These include Phi-4-mini-instruct, Phi-3.5 Mini, and OmniGen or OmniGen2 which all use 4.8 GB of memory. Phi-3 Mini also uses 4.8 GB of memory but at Q5_K_M quantization. Image and audio models like SD Cascade (Würstchen v3) use 4.6 GB of memory at Q8_0 quantization. SDXL Turbo and SDXL Lightning use 4.5 GB of memory at Q8_0 quantization. Stable Diffusion XL uses 4.1 GB of memory in FP8 or optimized format.
When a model is too large for the 5 GB video memory, you can use CPU offload if you have 32 GB of system RAM. Offloading splits the model between your graphics card and system memory. This allows you to run larger 7B models but it reduces processing speed. For example, Mistral 7B needs 5.7 GB at Q4_K_M quantization which requires 7.7 GB of system RAM. Other 7B models like Qwen2.5, OLMo 2, Falcon 3, Command R7B, OpenHermes 2.5, Zephyr 7B Beta, OpenChat 3.5, Starling LM 7B, and Codestral Mamba 7B need 5.1 GB at Q4_K_M quantization which requires 7.1 GB of system RAM.
You must consider the 4k context window caveat when running these models. The memory numbers listed only cover loading the model weights. As you type longer prompts and the model generates longer answers, the active context memory grows. Running models close to the 5 GB limit of your NVIDIA Quadro P2200 will leave very little room for long conversations or large image generation tasks.