Best local AI models for NVIDIA T600
4 GB GDDR6. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Lumina-Next / Lumina-Image 2.0 | 5B | Q4_K_M | 3.7 GB |
| CogVideoX 2B / 5B | 5B | Q4_K_M | 3.7 GB |
| DeepSeek-VL2 | 4.5B | Q5_K_M | 3.8 GB |
| DeepFloyd IF | 4.3B | Q5_K_M | 3.7 GB |
| Phi-3.5-vision | 4.2B | Q5_K_M | 3.6 GB |
| Qwen3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 4 E4B | 4B | Q6_K | 3.9 GB |
| MiniCPM 3 4B | 4B | Q6_K | 3.9 GB |
| Danube 3 4B | 4B | Q6_K | 3.9 GB |
| Fish Speech 1.5 / OpenAudio S1 | 4B | Q6_K | 3.9 GB |
| Phi-4-mini-instruct | 3.8B | Q6_K | 3.7 GB |
| Phi-3.5 Mini | 3.8B | Q6_K | 3.7 GB |
| OmniGen / OmniGen2 | 3.8B | Q6_K | 3.7 GB |
| SD Cascade (Würstchen v3) | 3.6B | Q6_K | 3.5 GB |
| SDXL Turbo | 3.5B | Q6_K | 3.4 GB |
| SDXL Lightning | 3.5B | Q6_K | 3.4 GB |
| ACE-Step | 3.5B | Q6_K | 3.4 GB |
| MusicGen small/medium/large | 3.3B | Q6_K | 3.2 GB |
| SmolLM3 3B | 3B | Q8_0 | 3.8 GB |
| Replit Code v1.5 3B | 3B | Q8_0 | 3.8 GB |
| Kandinsky 3.1 | 3B | Q8_0 | 3.8 GB |
| Voxtral Mini / Small | 3B | Q8_0 | 3.8 GB |
| Orpheus TTS | 3B | Q8_0 | 3.8 GB |
| Higgs Audio v2 | 3B | Q8_0 | 3.8 GB |
| Allegro | 2.8B | Q8_0 | 3.6 GB |
| Open-Sora Plan | 2.7B | Q8_0 | 3.4 GB |
| LFM2 1.2B / 2.6B | 2.6B | Q8_0 | 3.3 GB |
| Playground v2.5 | 2.6B | Q8_0 | 3.3 GB |
| Stable Diffusion 3.5 Medium | 2.5B | Q8_0 | 3.2 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| Stable Diffusion XL | 3.417B | 4.1 GB needed | 6.1 GB |
| Phi-3 Mini | 3.8B | 4.4 GB needed | 6.4 GB |
| Phi-4-multimodal | 5.6B | 4.1 GB needed | 6.1 GB |
| Magicoder-S-DS 6.7B | 6.7B | 4.9 GB needed | 6.9 GB |
| Mistral 7B | 7B | 5.7 GB needed | 7.7 GB |
| Qwen2.5 0.5B / 1.5B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| OLMo 2 1B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Falcon 3 1B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Command R7B | 7B | 5.1 GB needed | 7.1 GB |
| OpenHermes 2.5 | 7B | 5.1 GB needed | 7.1 GB |
How to read this
The NVIDIA T600 graphics card features 4 GB of GDDR6 memory. This dedicated memory determines the maximum size of the artificial intelligence models you can run directly on the hardware. To fit models within this limit, you must use quantized versions. Quantization reduces the precision of model weights to save space. The best quant column indicates the highest quality quantization level that fits entirely inside the 4 GB limit.
For running models completely on the graphics card, the largest options span from 2.5B to 5B parameters. Lumina-Next or Lumina-Image 2.0 and CogVideoX 2B or 5B both fit at the Q4_K_M quantization level using 3.7 GB of memory. DeepSeek-VL2 fits at Q5_K_M using 3.8 GB. DeepFloyd IF fits at Q5_K_M using 3.7 GB. Phi-3.5-vision fits at Q5_K_M using 3.6 GB.
Several 4B parameter models run efficiently on this hardware. Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, Danube 3 4B, and Fish Speech 1.5 or OpenAudio S1 all fit at the Q6_K quantization level using 3.9 GB of memory. Phi-4-mini-instruct, Phi-3.5 Mini, and OmniGen or OmniGen2 fit at Q6_K using 3.7 GB. SD Cascade (Würstchen v3) fits at Q6_K using 3.5 GB. SDXL Turbo, SDXL Lightning, and ACE-Step fit at Q6_K using 3.4 GB. MusicGen small/medium/large fits at Q6_K using 3.2 GB.
Models with 3B parameters or fewer can run at the higher Q8_0 quantization level. SmolLM3 3B, Replit Code v1.5 3B, Kandinsky 3.1, Voxtral Mini or Small, Orpheus TTS, and Higgs Audio v2 all use 3.8 GB of memory. Allegro fits at Q8_0 using 3.6 GB. Open-Sora Plan fits at Q8_0 using 3.4 GB. LFM2 1.2B or 2.6B and Playground v2.5 fit at Q8_0 using 3.3 GB. Stable Diffusion 3.5 Medium fits at Q8_0 using 3.2 GB.
When a model exceeds the 4 GB graphics memory, you must use CPU offloading. This process splits the model weights between your graphics card and your system RAM. Offloading allows you to run larger models but it reduces processing speed because system RAM is slower than graphics memory. These calculations assume your system has 32 GB of system RAM available.
For CPU offload cases, Stable Diffusion XL requires 4.1 GB at FP8 or optimized settings and uses 6.1 GB of system RAM. Phi-3 Mini requires 4.4 GB at Q4_K_M and uses 6.4 GB of system RAM. Phi-4-multimodal requires 4.1 GB at Q4_K_M and uses 6.1 GB of system RAM. Magicoder-S-DS 6.7B requires 4.9 GB at Q4_K_M and uses 6.9 GB of system RAM. Mistral 7B requires 5.7 GB at Q4_K_M and uses 7.7 GB of system RAM.
Other 7B models also run via offloading. Qwen2.5 0.5B or 1.5B or 3B or 7B, OLMo 2 1B or 7B, Falcon 3 1B or 3B or 7B, Command R7B, and OpenHermes 2.5 all require 5.1 GB at Q4_K_M and use 7.1 GB of system RAM. Note that memory usage calculations assume a standard 4k context window. Increasing the context window size will require more memory and may prevent these models from fitting.