Best local AI models for NVIDIA RTX 4070 Ti SUPER
16 GB GDDR6X. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| gpt-oss-20b | 21B | Q4_K_M | 15.4 GB |
| Reka Flash 3 | 21B | Q4_K_M | 15.4 GB |
| Qwen-Image | 20B | Q4_K_M | 14.6 GB |
| Qwen-Image-Edit | 20B | Q4_K_M | 14.6 GB |
| CogVLM2 | 19B | Q4_K_M | 13.9 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q5_K_M | 14.5 GB |
| Ling-Coder-Lite | 16.8B | Q5_K_M | 14.3 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q6_K | 15.7 GB |
| Kimi-VL A3B | 16B | Q6_K | 15.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q6_K | 14.8 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q6_K | 14.8 GB |
| Qwen2.5 14B | 14.7B | Q6_K | 15.3 GB |
| Phi-3 Medium | 14B | Q6_K | 13.8 GB |
| Phi-4 | 14B | Q6_K | 13.8 GB |
| Phi-4-reasoning / -plus | 14B | Q6_K | 13.8 GB |
| Wan 2.2 T2I | 14B | Q6_K | 13.8 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q6_K | 13.8 GB |
| SkyReels V2 | 14B | Q6_K | 13.8 GB |
| Vicuna 13B | 13B | Q6_K | 12.8 GB |
| HunyuanVideo | 13B | Q6_K | 12.8 GB |
| HunyuanVideo-Avatar | 13B | Q6_K | 12.8 GB |
| LTX-Video / LTX-2 | 13B | Q6_K | 12.8 GB |
| FramePack | 13B | Q6_K | 12.8 GB |
| FLUX.1 dev | 12B | FP8 / optimized | 14.4 GB |
| Gemma 3 12B | 12B | Q8_0 | 15.3 GB |
| Gemma 4 12B | 12B | Q8_0 | 15.3 GB |
| Mistral NeMo 12B | 12B | Q8_0 | 15.3 GB |
| Pixtral 12B | 12B | Q8_0 | 15.3 GB |
| FLUX.1 schnell | 12B | Q8_0 | 15.3 GB |
| FLUX.1 Kontext dev | 12B | Q8_0 | 15.3 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
| Mistral Small 3.2 | 24B | 17.6 GB needed | 19.6 GB |
| Magistral Small | 24B | 17.6 GB needed | 19.6 GB |
| Devstral Small 1.1 | 24B | 17.6 GB needed | 19.6 GB |
| Aria | 25B | 18.3 GB needed | 20.3 GB |
| Gemma 4 26B-A4B | 26B | 19 GB needed | 21 GB |
| Gemma 4 (all sizes) | 26B | 19 GB needed | 21 GB |
| Gemma 3 27B | 27B | 19.8 GB needed | 21.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | 19.8 GB needed | 21.8 GB |
How to read this
The NVIDIA RTX 4070 Ti SUPER features 16 GB of GDDR6X memory. This dedicated video memory determines the size of the artificial intelligence models you can run locally. To run a model entirely on your graphics card, the model files and the active memory space must fit within this 16 GB limit. Running models locally on your graphics hardware ensures the fastest generation speeds.
The quantization column indicates the compression level applied to each model. Raw models are often too large for consumer hardware. Quantization reduces the precision of the model weights to make the files smaller. For example, the Q4_K_M quant represents a four bit medium quantization. The Q6_K and Q8_0 quants offer higher precision but require more memory. A higher quant level improves output quality but limits the maximum model size you can load.
Several large models fit completely within the 16 GB memory limit of your card. The gpt-oss-20b and Reka Flash 3 models use 15.4 GB of memory at the Q4_K_M quantization. You can also run Qwen-Image and Qwen-Image-Edit at Q4_K_M using 14.6 GB of memory. For higher precision, models like DeepSeek-Coder-V2 16B / 236B and Kimi-VL A3B run at Q6_K using 15.7 GB of memory. The FLUX.1 dev model runs at FP8 / optimized using 14.4 GB of memory.
When a model exceeds 16 GB of memory, you must offload some layers to your system RAM. This process requires at least 32 GB of system RAM to function. For example, Solar Pro and Codestral 22B require 16.1 GB of memory at Q4_K_M, which needs 18.1 GB of system RAM. Gemma 3 27B requires 19.8 GB of memory at Q4_K_M, which needs 21.8 GB of system RAM. Offloading allows you to run larger models like Mistral Small 3.2 or Gemma 4 26B-A4B, but it significantly reduces processing speed.
Memory calculations assume a standard 4k context window. The context window is the amount of text the model can remember during a conversation. If you increase the context window beyond 4k tokens, the model will require more memory. This extra memory usage might force you to use a lower quantization level or offload layers to system RAM to prevent out of memory errors.