Best local AI models for NVIDIA Quadro K4200
4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Lumina-Next / Lumina-Image 2.0 | 5B | Q4_K_M | 3.7 GB |
| CogVideoX 2B / 5B | 5B | Q4_K_M | 3.7 GB |
| DeepSeek-VL2 | 4.5B | Q5_K_M | 3.8 GB |
| DeepFloyd IF | 4.3B | Q5_K_M | 3.7 GB |
| Phi-3.5-vision | 4.2B | Q5_K_M | 3.6 GB |
| Qwen3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 4 E4B | 4B | Q6_K | 3.9 GB |
| MiniCPM 3 4B | 4B | Q6_K | 3.9 GB |
| Danube 3 4B | 4B | Q6_K | 3.9 GB |
| Fish Speech 1.5 / OpenAudio S1 | 4B | Q6_K | 3.9 GB |
| Phi-4-mini-instruct | 3.8B | Q6_K | 3.7 GB |
| Phi-3.5 Mini | 3.8B | Q6_K | 3.7 GB |
| OmniGen / OmniGen2 | 3.8B | Q6_K | 3.7 GB |
| SD Cascade (Würstchen v3) | 3.6B | Q6_K | 3.5 GB |
| SDXL Turbo | 3.5B | Q6_K | 3.4 GB |
| SDXL Lightning | 3.5B | Q6_K | 3.4 GB |
| ACE-Step | 3.5B | Q6_K | 3.4 GB |
| MusicGen small/medium/large | 3.3B | Q6_K | 3.2 GB |
| SmolLM3 3B | 3B | Q8_0 | 3.8 GB |
| Replit Code v1.5 3B | 3B | Q8_0 | 3.8 GB |
| Kandinsky 3.1 | 3B | Q8_0 | 3.8 GB |
| Voxtral Mini / Small | 3B | Q8_0 | 3.8 GB |
| Orpheus TTS | 3B | Q8_0 | 3.8 GB |
| Higgs Audio v2 | 3B | Q8_0 | 3.8 GB |
| Allegro | 2.8B | Q8_0 | 3.6 GB |
| Open-Sora Plan | 2.7B | Q8_0 | 3.4 GB |
| LFM2 1.2B / 2.6B | 2.6B | Q8_0 | 3.3 GB |
| Playground v2.5 | 2.6B | Q8_0 | 3.3 GB |
| Stable Diffusion 3.5 Medium | 2.5B | Q8_0 | 3.2 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| Stable Diffusion XL | 3.417B | 4.1 GB needed | 6.1 GB |
| Phi-3 Mini | 3.8B | 4.4 GB needed | 6.4 GB |
| Phi-4-multimodal | 5.6B | 4.1 GB needed | 6.1 GB |
| Magicoder-S-DS 6.7B | 6.7B | 4.9 GB needed | 6.9 GB |
| Mistral 7B | 7B | 5.7 GB needed | 7.7 GB |
| Qwen2.5 0.5B / 1.5B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| OLMo 2 1B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Falcon 3 1B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Command R7B | 7B | 5.1 GB needed | 7.1 GB |
| OpenHermes 2.5 | 7B | 5.1 GB needed | 7.1 GB |
How to read this
The NVIDIA Quadro K4200 is an older workstation graphics card equipped with 4 GB of GDDR5 dedicated memory. When running artificial intelligence models locally, this hardware memory capacity is the main limiting factor. To fit within the physical limits of your card, models must be compressed to prevent out of memory errors. This page guides you through selecting the right model sizes and quantization levels for your hardware.
The quantization column shows the compression level used to shrink the model files. Quantization reduces the precision of model weights to save space. For example, the Lumina-Next or Lumina-Image 2.0 5B model fits on your card using a Q4_K_M quantization which consumes 3.7 GB of memory. Similarly, the DeepSeek-VL2 4.5B model fits using a Q5_K_M quantization and uses 3.8 GB of memory. Smaller models like the Qwen3 4B or Gemma 3 4B can run at a higher Q6_K quantization while using 3.9 GB of memory.
If you want maximum output quality, you can use smaller models with minimal compression. The SmolLM3 3B and Orpheus TTS 3B models run at a high Q8_0 quantization, using 3.8 GB of memory. Image generation models also fit within these limits. SDXL Turbo and SDXL Lightning 3.5B models use a Q6_K quantization and require 3.4 GB of memory. Stable Diffusion 3.5 Medium 2.5B fits at Q8_0 quantization using 3.2 GB of memory.
When a model is slightly too large for your 4 GB of video memory, you can use CPU offloading. This technique splits the model between your graphics card and your system RAM, but it will slow down processing speeds. For offloading, we assume your computer has 32 GB of system RAM. Under this setup, Mistral 7B requires 5.7 GB of memory at Q4_K_M quantization and uses 7.7 GB of system RAM. The Falcon 3 7B model needs 5.1 GB at Q4_K_M quantization and uses 7.1 GB of system RAM.
Other offloading options include the Phi-4-multimodal 5.6B model, which needs 4.1 GB at Q4_K_M quantization and uses 6.1 GB of system RAM. Stable Diffusion XL 3.417B needs 4.1 GB at FP8 or optimized settings and uses 6.1 GB of system RAM. Remember that these offloaded models will run much slower than models that fit entirely inside your 4 GB video memory.
You must also consider the memory cost of context length. Running text models with long conversations or large documents increases memory usage. The memory figures listed here are calculated using a standard 4k context window. If you increase the context window beyond 4000 tokens, the model will require more memory and might exceed your 4 GB limit.