Best local AI models for NVIDIA T600

4 GB GDDR6. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Lumina-Next / Lumina-Image 2.05BQ4_K_M3.7 GB
CogVideoX 2B / 5B5BQ4_K_M3.7 GB
DeepSeek-VL24.5BQ5_K_M3.8 GB
DeepFloyd IF4.3BQ5_K_M3.7 GB
Phi-3.5-vision4.2BQ5_K_M3.6 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-4-mini-instruct3.8BQ6_K3.7 GB
Phi-3.5 Mini3.8BQ6_K3.7 GB
OmniGen / OmniGen23.8BQ6_K3.7 GB
SD Cascade (Würstchen v3)3.6BQ6_K3.5 GB
SDXL Turbo3.5BQ6_K3.4 GB
SDXL Lightning3.5BQ6_K3.4 GB
ACE-Step3.5BQ6_K3.4 GB
MusicGen small/medium/large3.3BQ6_K3.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB
Open-Sora Plan2.7BQ8_03.4 GB
LFM2 1.2B / 2.6B2.6BQ8_03.3 GB
Playground v2.52.6BQ8_03.3 GB
Stable Diffusion 3.5 Medium2.5BQ8_03.2 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB

How to read this

The NVIDIA T600 graphics card features 4 GB of GDDR6 memory. This dedicated memory determines the maximum size of the artificial intelligence models you can run directly on the hardware. To fit models within this limit, you must use quantized versions. Quantization reduces the precision of model weights to save space. The best quant column indicates the highest quality quantization level that fits entirely inside the 4 GB limit.

For running models completely on the graphics card, the largest options span from 2.5B to 5B parameters. Lumina-Next or Lumina-Image 2.0 and CogVideoX 2B or 5B both fit at the Q4_K_M quantization level using 3.7 GB of memory. DeepSeek-VL2 fits at Q5_K_M using 3.8 GB. DeepFloyd IF fits at Q5_K_M using 3.7 GB. Phi-3.5-vision fits at Q5_K_M using 3.6 GB.

Several 4B parameter models run efficiently on this hardware. Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, Danube 3 4B, and Fish Speech 1.5 or OpenAudio S1 all fit at the Q6_K quantization level using 3.9 GB of memory. Phi-4-mini-instruct, Phi-3.5 Mini, and OmniGen or OmniGen2 fit at Q6_K using 3.7 GB. SD Cascade (Würstchen v3) fits at Q6_K using 3.5 GB. SDXL Turbo, SDXL Lightning, and ACE-Step fit at Q6_K using 3.4 GB. MusicGen small/medium/large fits at Q6_K using 3.2 GB.

Models with 3B parameters or fewer can run at the higher Q8_0 quantization level. SmolLM3 3B, Replit Code v1.5 3B, Kandinsky 3.1, Voxtral Mini or Small, Orpheus TTS, and Higgs Audio v2 all use 3.8 GB of memory. Allegro fits at Q8_0 using 3.6 GB. Open-Sora Plan fits at Q8_0 using 3.4 GB. LFM2 1.2B or 2.6B and Playground v2.5 fit at Q8_0 using 3.3 GB. Stable Diffusion 3.5 Medium fits at Q8_0 using 3.2 GB.

When a model exceeds the 4 GB graphics memory, you must use CPU offloading. This process splits the model weights between your graphics card and your system RAM. Offloading allows you to run larger models but it reduces processing speed because system RAM is slower than graphics memory. These calculations assume your system has 32 GB of system RAM available.

For CPU offload cases, Stable Diffusion XL requires 4.1 GB at FP8 or optimized settings and uses 6.1 GB of system RAM. Phi-3 Mini requires 4.4 GB at Q4_K_M and uses 6.4 GB of system RAM. Phi-4-multimodal requires 4.1 GB at Q4_K_M and uses 6.1 GB of system RAM. Magicoder-S-DS 6.7B requires 4.9 GB at Q4_K_M and uses 6.9 GB of system RAM. Mistral 7B requires 5.7 GB at Q4_K_M and uses 7.7 GB of system RAM.

Other 7B models also run via offloading. Qwen2.5 0.5B or 1.5B or 3B or 7B, OLMo 2 1B or 7B, Falcon 3 1B or 3B or 7B, Command R7B, and OpenHermes 2.5 all require 5.1 GB at Q4_K_M and use 7.1 GB of system RAM. Note that memory usage calculations assume a standard 4k context window. Increasing the context window size will require more memory and may prevent these models from fitting.