Best local AI models for NVIDIA TITAN Xp
12 GB GDDR5X. At a 4k context, 147 of the 233 models in our catalog with verified parameter counts fit fully, up to DeepSeek-Coder-V2 16B / 236B at 16B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 147 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| DeepSeek-Coder-V2 16B / 236B | 16B | Q4_K_M | 11.7 GB |
| Kimi-VL A3B | 16B | Q4_K_M | 11.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q4_K_M | 11 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q4_K_M | 11 GB |
| Qwen2.5 14B | 14.7B | Q4_K_M | 11.6 GB |
| Phi-3 Medium | 14B | Q5_K_M | 11.9 GB |
| Phi-4 | 14B | Q5_K_M | 11.9 GB |
| Phi-4-reasoning / -plus | 14B | Q5_K_M | 11.9 GB |
| Wan 2.2 T2I | 14B | Q5_K_M | 11.9 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q5_K_M | 11.9 GB |
| SkyReels V2 | 14B | Q5_K_M | 11.9 GB |
| Vicuna 13B | 13B | Q5_K_M | 11.1 GB |
| HunyuanVideo | 13B | Q5_K_M | 11.1 GB |
| HunyuanVideo-Avatar | 13B | Q5_K_M | 11.1 GB |
| LTX-Video / LTX-2 | 13B | Q5_K_M | 11.1 GB |
| FramePack | 13B | Q5_K_M | 11.1 GB |
| Gemma 3 12B | 12B | Q6_K | 11.8 GB |
| Gemma 4 12B | 12B | Q6_K | 11.8 GB |
| Mistral NeMo 12B | 12B | Q6_K | 11.8 GB |
| Pixtral 12B | 12B | Q6_K | 11.8 GB |
| FLUX.1 schnell | 12B | Q6_K | 11.8 GB |
| FLUX.1 Kontext dev | 12B | Q6_K | 11.8 GB |
| FLUX.1 Krea dev | 12B | Q6_K | 11.8 GB |
| Open-Sora 2.0 | 11B | Q6_K | 10.8 GB |
| Mochi 1 | 10B | Q6_K | 9.8 GB |
| Gemma 2 9B | 9B | Q6_K | 10.3 GB |
| Nemotron Nano 4B / 9B | 9B | Q8_0 | 11.4 GB |
| GLM-4 9B / GLM-4.5-Air | 9B | Q8_0 | 11.4 GB |
| Yi-Coder 1.5B / 9B | 9B | Q8_0 | 11.4 GB |
| GLM-4-9B-Chat / CodeGeeX4 | 9B | Q8_0 | 11.4 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| FLUX.1 dev | 12B | 14.4 GB needed | 16.4 GB |
| Ling-Coder-Lite | 16.8B | 12.3 GB needed | 14.3 GB |
| HunyuanImage 2.1 / 3.0 | 17B | 12.4 GB needed | 14.4 GB |
| CogVLM2 | 19B | 13.9 GB needed | 15.9 GB |
| Qwen-Image | 20B | 14.6 GB needed | 16.6 GB |
| Qwen-Image-Edit | 20B | 14.6 GB needed | 16.6 GB |
| gpt-oss-20b | 21B | 15.4 GB needed | 17.4 GB |
| Reka Flash 3 | 21B | 15.4 GB needed | 17.4 GB |
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
How to read this
The NVIDIA TITAN Xp graphics card features 12 GB of GDDR5X onboard memory. This physical memory size determines the maximum size of the artificial intelligence models you can run locally. To fit a model entirely within this frame buffer, the total size of the model weights and the active context window must not exceed 12 GB. Keeping the model fully inside the graphics memory ensures the fastest possible processing speeds.
Models are compressed using quantization to fit into smaller memory footprints. The quant column shows the best available quantization level that fits your hardware. For example, DeepSeek-Coder-V2 16B and Kimi-VL A3B can run at the Q4_K_M quantization level using 11.7 GB of memory. Similarly, Apriel-1.5-15B-Thinker and StarCoder2 15B fit at Q4_K_M using 11 GB of memory. Qwen2.5 14B fits at Q4_K_M using 11.6 GB of memory.
Other models can run at higher precision levels because of their smaller base sizes. Phi-3 Medium, Phi-4, Phi-4-reasoning / -plus, Wan 2.2 T2I, Wan 2.1 (1.3B / 14B), and SkyReels V2 are 14B models that run at the Q5_K_M quantization level using 11.9 GB of memory. Vicuna 13B, HunyuanVideo, HunyuanVideo-Avatar, LTX-Video / LTX-2, and FramePack are 13B models that run at Q5_K_M using 11.1 GB of memory. Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, FLUX.1 Kontext dev, and FLUX.1 Krea dev are 12B models that run at the Q6_K level using 11.8 GB of memory.
Smaller models can run at even higher quality levels. Open-Sora 2.0 is an 11B model that runs at Q6_K using 10.8 GB of memory. Mochi 1 is a 10B model that runs at Q6_K using 9.8 GB of memory. Gemma 2 9B runs at Q6_K using 10.3 GB of memory. Nemotron Nano 4B / 9B, GLM-4 9B / GLM-4.5-Air, Yi-Coder 1.5B / 9B, and GLM-4-9B-Chat / CodeGeeX4 are 9B models that run at the Q8_0 quantization level using 11.4 GB of memory.
If you want to run larger models, you must use CPU offloading. This process splits the model weights between your graphics card and your system memory. Offloading allows you to run larger models but it reduces your generation speed. For a system with 32 GB of system RAM, FLUX.1 dev needs 14.4 GB at FP8 / optimized and 16.4 GB of system RAM. Ling-Coder-Lite needs 12.3 GB at Q4_K_M and 14.3 GB of system RAM. HunyuanImage 2.1 / 3.0 needs 12.4 GB at Q4_K_M and 14.4 GB of system RAM. CogVLM2 needs 13.9 GB at Q4_K_M and 15.9 GB of system RAM.
Other offload options include Qwen-Image and Qwen-Image-Edit, which need 14.6 GB at Q4_K_M and 16.6 GB of system RAM. The gpt-oss-20b and Reka Flash 3 models need 15.4 GB at Q4_K_M and 17.4 GB of system RAM. Solar Pro and Codestral 22B need 16.1 GB at Q4_K_M and 18.1 GB of system RAM. When planning your setup, remember that memory calculations assume a standard 4k context window. Running longer conversations or larger prompts will require more memory and might require a lower quantization level.