Best local AI models for NVIDIA Quadro K600

1 GB DDR3. At a 4k context, 28 of the 233 models in our catalog with verified parameter counts fit fully, up to Tango 2 at 1.4B parameters.

Check your own machine against every model →

The largest models that fit fully

The 28 largest of the 28 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Tango 21.4BQ4_K_M1 GB
TinyLlama 1.1B1.1BQ5_K_M0.9 GB
SantaCoder 1.1B1.1BQ5_K_M0.9 GB
Stable Audio Open 1.0 / small1.1BQ5_K_M0.9 GB
Gemma 3 1B1BQ6_K1 GB
Llama 3.2 1B / 3B1BQ6_K1 GB
MMS (1100+ languages)1BQ6_K1 GB
CSM-1B1BQ6_K1 GB
IndexTTS 21BQ6_K1 GB
DiffRhythm1BQ6_K1 GB
Stable Diffusion 2.10.9BQ6_K0.9 GB
Bark0.9BQ6_K0.9 GB
Tortoise TTS0.9BQ6_K0.9 GB
Riffusion (SD-based)0.9BQ6_K0.9 GB
Magenta RT0.8BQ8_01 GB
Florence-2 base/large0.77BQ8_01 GB
Qwen3 0.6B0.6BQ8_00.8 GB
PixArt-α / PixArt-Σ0.6BQ8_00.8 GB
Parakeet TDT 0.6B v20.6BQ8_00.8 GB
XTTS v20.5BQ8_00.6 GB
Spark-TTS0.5BQ8_00.6 GB
CosyVoice 20.5BQ8_00.6 GB
VALL-E X (unofficial)0.4BFP161 GB
ERNIE 4.5 open weights0.3BFP160.7 GB
F5-TTS0.3BFP160.7 GB
E2-TTS0.3BFP160.7 GB
ChatTTS0.3BFP160.7 GB
StyleTTS 20.15BFP160.4 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion 1.51.07B1.3 GB needed3.3 GB
ControlNet / T2I-Adapter / IP-Adapter1.5B1.1 GB needed3.1 GB
Hunyuan-DiT1.5B1.1 GB needed3.1 GB
Stable Video Diffusion1.5B1.1 GB needed3.1 GB
Whisper Large v2 / turbo1.5B1.1 GB needed3.1 GB
AudioGen1.5B1.1 GB needed3.1 GB
AudioLDM 21.5B1.1 GB needed3.1 GB
Whisper Large v31.55B1.3 GB needed3.3 GB
StableLM 2 1.6B1.6B1.2 GB needed3.2 GB
Sana 0.6B / 1.6B1.6B1.2 GB needed3.2 GB

How to read this

The NVIDIA Quadro K600 is an entry level professional graphics card equipped with 1 GB DDR3 of onboard video memory. This hardware memory size represents the absolute limit for storing model weights directly on the graphics card. To run artificial intelligence models locally on this hardware, you must select small architectures or use aggressive quantization to reduce the memory footprint of the weights.

The quantization column indicates the specific numerical format used to compress the model. For example, a Q4_K_M quant uses approximately four bits per weight to compress the Tango 2 1.4B model down to 1 GB of video memory. Higher precision formats like FP16 preserve more original quality but require much smaller models like the StyleTTS 2 0.15B model which uses 0.4 GB of video memory.

When a model size exceeds the physical 1 GB limit of the graphics card, you must use CPU offload. This technique splits the workload between your graphics card and your system memory. To use CPU offload, your computer should have 32 GB system RAM. For example, running Stable Diffusion 1.5 requires 1.3 GB of video memory at FP8 or optimized settings, which forces 3.3 GB of data into your system RAM.

Offloading weights to system RAM carries a significant performance cost. System RAM and DDR3 video memory transfer data much slower than modern graphics memory. While offloading allows you to run larger architectures like the StableLM 2 1.6B model or Sana 1.6B model, the generation speed will be much slower because of the data transfer bottleneck between the processor and system memory.

You must also consider the context window when running text models. The listed memory usage figures only account for the static model weights. Running a model with a 4k context window requires additional video memory to store the active conversation history. This extra memory overhead can easily exceed the 1 GB limit of the card, so you may need to reduce the context length to prevent out of memory errors.