Best local AI models for NVIDIA T400 4GB
4 GB GDDR6. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Lumina-Next / Lumina-Image 2.0 | 5B | Q4_K_M | 3.7 GB |
| CogVideoX 2B / 5B | 5B | Q4_K_M | 3.7 GB |
| DeepSeek-VL2 | 4.5B | Q5_K_M | 3.8 GB |
| DeepFloyd IF | 4.3B | Q5_K_M | 3.7 GB |
| Phi-3.5-vision | 4.2B | Q5_K_M | 3.6 GB |
| Qwen3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 3 4B | 4B | Q6_K | 3.9 GB |
| Gemma 4 E4B | 4B | Q6_K | 3.9 GB |
| MiniCPM 3 4B | 4B | Q6_K | 3.9 GB |
| Danube 3 4B | 4B | Q6_K | 3.9 GB |
| Fish Speech 1.5 / OpenAudio S1 | 4B | Q6_K | 3.9 GB |
| Phi-4-mini-instruct | 3.8B | Q6_K | 3.7 GB |
| Phi-3.5 Mini | 3.8B | Q6_K | 3.7 GB |
| OmniGen / OmniGen2 | 3.8B | Q6_K | 3.7 GB |
| SD Cascade (Würstchen v3) | 3.6B | Q6_K | 3.5 GB |
| SDXL Turbo | 3.5B | Q6_K | 3.4 GB |
| SDXL Lightning | 3.5B | Q6_K | 3.4 GB |
| ACE-Step | 3.5B | Q6_K | 3.4 GB |
| MusicGen small/medium/large | 3.3B | Q6_K | 3.2 GB |
| SmolLM3 3B | 3B | Q8_0 | 3.8 GB |
| Replit Code v1.5 3B | 3B | Q8_0 | 3.8 GB |
| Kandinsky 3.1 | 3B | Q8_0 | 3.8 GB |
| Voxtral Mini / Small | 3B | Q8_0 | 3.8 GB |
| Orpheus TTS | 3B | Q8_0 | 3.8 GB |
| Higgs Audio v2 | 3B | Q8_0 | 3.8 GB |
| Allegro | 2.8B | Q8_0 | 3.6 GB |
| Open-Sora Plan | 2.7B | Q8_0 | 3.4 GB |
| LFM2 1.2B / 2.6B | 2.6B | Q8_0 | 3.3 GB |
| Playground v2.5 | 2.6B | Q8_0 | 3.3 GB |
| Stable Diffusion 3.5 Medium | 2.5B | Q8_0 | 3.2 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| Stable Diffusion XL | 3.417B | 4.1 GB needed | 6.1 GB |
| Phi-3 Mini | 3.8B | 4.4 GB needed | 6.4 GB |
| Phi-4-multimodal | 5.6B | 4.1 GB needed | 6.1 GB |
| Magicoder-S-DS 6.7B | 6.7B | 4.9 GB needed | 6.9 GB |
| Mistral 7B | 7B | 5.7 GB needed | 7.7 GB |
| Qwen2.5 0.5B / 1.5B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| OLMo 2 1B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Falcon 3 1B / 3B / 7B | 7B | 5.1 GB needed | 7.1 GB |
| Command R7B | 7B | 5.1 GB needed | 7.1 GB |
| OpenHermes 2.5 | 7B | 5.1 GB needed | 7.1 GB |
How to read this
The NVIDIA T400 features 4 GB of GDDR6 memory. This dedicated VRAM determines the size of the neural networks you can run entirely on the graphics card. To run a model without slowdowns, the model files and the active memory must fit inside this 4 GB limit. If a model exceeds this capacity, your system must transfer data between the card and your system RAM, which reduces processing speeds.
Quantization is a method that compresses model files to save space. The quant column shows the best compression level for each model on this hardware. For example, a Q4_K_M quant uses four bit quantization to reduce file size while keeping high accuracy. A Q6_K or Q8_0 quant offers better output quality but requires more VRAM. Choosing the right quant allows you to run larger models like the 5B Lumina-Next or Lumina-Image 2.0 at Q4_K_M using 3.7 GB of VRAM.
Many capable models fit within the 4 GB VRAM limit of the NVIDIA T400. You can run the 5B CogVideoX at Q4_K_M using 3.7 GB of VRAM. The 4.5B DeepSeek-VL2 fits at Q5_K_M using 3.8 GB of VRAM. Other options include DeepFloyd IF at Q5_K_M using 3.7 GB of VRAM and Phi-3.5-vision at Q5_K_M using 3.6 GB of VRAM. For text and audio, you can run Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, Danube 3 4B, and Fish Speech 1.5 at Q6_K using 3.9 GB of VRAM.
Other models fit well within the hardware limits. Phi-4-mini-instruct, Phi-3.5 Mini, and OmniGen use 3.7 GB of VRAM at Q6_K. Image generators like SD Cascade use 3.5 GB of VRAM, while SDXL Turbo and SDXL Lightning use 3.4 GB of VRAM at Q6_K. You can also run ACE-Step using 3.4 GB of VRAM and MusicGen using 3.2 GB of VRAM. Models like SmolLM3 3B, Replit Code v1.5 3B, Kandinsky 3.1, Voxtral Mini, Orpheus TTS, and Higgs Audio v2 run at Q8_0 using 3.8 GB of VRAM. Allegro uses 3.6 GB of VRAM, Open-Sora Plan uses 3.4 GB of VRAM, LFM2 uses 3.3 GB of VRAM, Playground v2.5 uses 3.3 GB of VRAM, and Stable Diffusion 3.5 Medium uses 3.2 GB of VRAM.
When a model is too large for the 4 GB VRAM, you can use CPU offload. This technique splits the model between your graphics card and your system RAM. For these cases, we assume you have 32 GB of system RAM. Stable Diffusion XL needs 4.1 GB at FP8 and 6.1 GB of system RAM. Phi-3 Mini needs 4.4 GB at Q4_K_M and 6.4 GB of system RAM. Phi-4-multimodal needs 4.1 GB at Q4_K_M and 6.1 GB of system RAM. Magicoder-S-DS 6.7B needs 4.9 GB at Q4_K_M and 6.9 GB of system RAM. Mistral 7B needs 5.7 GB at Q4_K_M and 7.7 GB of system RAM.
Other larger models can also run using CPU offload. Qwen2.5 7B, OLMo 2 7B, Falcon 3 7B, Command R7B, and OpenHermes 2.5 need 5.1 GB at Q4_K_M and 7.1 GB of system RAM. While CPU offload allows you to run these larger models, it comes with a cost. Sharing data between the graphics card and system RAM slows down the generation speed compared to running entirely in VRAM.
You must also consider the context window when running local models. The VRAM usage numbers listed here are measured at a standard 4k context window. If you increase the context window to process longer documents or larger chat histories, the memory usage will rise. This extra memory demand can push a model over the 4 GB limit and trigger slow system RAM offloading.