Best local AI models for NVIDIA Quadro K1200

4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Lumina-Next / Lumina-Image 2.05BQ4_K_M3.7 GB
CogVideoX 2B / 5B5BQ4_K_M3.7 GB
DeepSeek-VL24.5BQ5_K_M3.8 GB
DeepFloyd IF4.3BQ5_K_M3.7 GB
Phi-3.5-vision4.2BQ5_K_M3.6 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-4-mini-instruct3.8BQ6_K3.7 GB
Phi-3.5 Mini3.8BQ6_K3.7 GB
OmniGen / OmniGen23.8BQ6_K3.7 GB
SD Cascade (Würstchen v3)3.6BQ6_K3.5 GB
SDXL Turbo3.5BQ6_K3.4 GB
SDXL Lightning3.5BQ6_K3.4 GB
ACE-Step3.5BQ6_K3.4 GB
MusicGen small/medium/large3.3BQ6_K3.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB
Open-Sora Plan2.7BQ8_03.4 GB
LFM2 1.2B / 2.6B2.6BQ8_03.3 GB
Playground v2.52.6BQ8_03.3 GB
Stable Diffusion 3.5 Medium2.5BQ8_03.2 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB

How to read this

The NVIDIA Quadro K1200 is equipped with 4 GB of GDDR5 memory. This hardware limit determines which local AI models can run entirely on your graphics card. When a model fits inside this 4 GB boundary, it executes quickly because the GPU can access the weights directly. If a model exceeds this limit, you must use CPU offloading to split the workload.

To fit larger models into the 4 GB memory space, you must use quantized versions. Quantization reduces the precision of model weights to save space. The quant column shows the best option for each model. For example, a Q4_K_M quantization allows the 5B Lumina-Next or CogVideoX 2B to run using 3.7 GB of memory. Models like Qwen3 4B and Gemma 3 4B run at Q6_K quantization using 3.9 GB of memory. Smaller models like SmolLM3 3B can run at a higher Q8_0 quantization using 3.8 GB of memory.

Running models with a 4k context window requires extra memory. The context window stores the history of your conversation. As your chat grows longer, the memory usage increases. If you use the maximum 4k context, you must ensure there is enough free space left within your 4 GB limit. Running too close to the limit can cause out of memory errors.

When a model is too large for the 4 GB memory, you can offload parts of it to your system RAM. This process requires a system with 32 GB of system RAM. Offloading allows you to run larger models, but it reduces processing speed because data must travel between the CPU and GPU. For example, Mistral 7B needs 5.7 GB of memory at Q4_K_M quantization, which requires 7.7 GB of system RAM to function.

Other models also benefit from CPU offloading. The Phi-4-multimodal model at 5.6B needs 4.1 GB at Q4_K_M quantization and 6.1 GB of system RAM. The Qwen2.5 7B, Falcon 3 7B, and Command R7B models each need 5.1 GB at Q4_K_M quantization and 7.1 GB of system RAM. Stable Diffusion XL needs 4.1 GB at FP8 or optimized settings, which requires 6.1 GB of system RAM to run.