Best local AI models for NVIDIA RTX 5070 Ti

16 GB GDDR7. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
gpt-oss-20b21BQ4_K_M15.4 GB
Reka Flash 321BQ4_K_M15.4 GB
Qwen-Image20BQ4_K_M14.6 GB
Qwen-Image-Edit20BQ4_K_M14.6 GB
CogVLM219BQ4_K_M13.9 GB
HunyuanImage 2.1 / 3.017BQ5_K_M14.5 GB
Ling-Coder-Lite16.8BQ5_K_M14.3 GB
DeepSeek-Coder-V2 16B / 236B16BQ6_K15.7 GB
Kimi-VL A3B16BQ6_K15.7 GB
Apriel-1.5-15B-Thinker15BQ6_K14.8 GB
StarCoder2 3B / 7B / 15B15BQ6_K14.8 GB
Qwen2.5 14B14.7BQ6_K15.3 GB
Phi-3 Medium14BQ6_K13.8 GB
Phi-414BQ6_K13.8 GB
Phi-4-reasoning / -plus14BQ6_K13.8 GB
Wan 2.2 T2I14BQ6_K13.8 GB
Wan 2.1 (1.3B / 14B)14BQ6_K13.8 GB
SkyReels V214BQ6_K13.8 GB
Vicuna 13B13BQ6_K12.8 GB
HunyuanVideo13BQ6_K12.8 GB
HunyuanVideo-Avatar13BQ6_K12.8 GB
LTX-Video / LTX-213BQ6_K12.8 GB
FramePack13BQ6_K12.8 GB
FLUX.1 dev12BFP8 / optimized14.4 GB
Gemma 3 12B12BQ8_015.3 GB
Gemma 4 12B12BQ8_015.3 GB
Mistral NeMo 12B12BQ8_015.3 GB
Pixtral 12B12BQ8_015.3 GB
FLUX.1 schnell12BQ8_015.3 GB
FLUX.1 Kontext dev12BQ8_015.3 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at Q4_K_MSystem RAM at 4k
Solar Pro22B16.1 GB needed18.1 GB
Codestral 22B22B16.1 GB needed18.1 GB
Mistral Small 3.224B17.6 GB needed19.6 GB
Magistral Small24B17.6 GB needed19.6 GB
Devstral Small 1.124B17.6 GB needed19.6 GB
Aria25B18.3 GB needed20.3 GB
Gemma 4 26B-A4B26B19 GB needed21 GB
Gemma 4 (all sizes)26B19 GB needed21 GB
Gemma 3 27B27B19.8 GB needed21.8 GB
Gemma 3 4B/12B/27B (vision)27B19.8 GB needed21.8 GB

How to read this

The NVIDIA RTX 5070 Ti features 16 GB of GDDR7 memory. This memory size determines which local artificial intelligence models can run entirely on your graphics hardware. When a model fits completely within this video memory, it runs at maximum speed. If a model exceeds this limit, it cannot run without modifications or system memory assistance.

To fit larger models into the 16 GB limit, we use quantization. The quant column shows the compression level applied to each model. For example, the gpt-oss-20b and Reka Flash 3 models use the Q4_K_M quantization to fit within 15.4 GB of memory. Smaller models like Gemma 3 12B and Mistral NeMo 12B can use the higher quality Q8_0 quantization because they require 15.3 GB of memory.

When you run models close to the memory limit, you must consider the context window. Running a model with a 4k context window requires extra memory to store the active conversation history. If you use the maximum context length, the system may run out of video memory. You should choose a slightly smaller model size or a lower quantization level if you need long chat sessions.

For models that exceed 16 GB, you can offload part of the workload to your system RAM. This process requires at least 32 GB of system RAM to function. For example, Solar Pro and Codestral 22B require 16.1 GB at Q4_K_M quantization, which uses 18.1 GB of system RAM. Larger models like Gemma 3 27B require 19.8 GB at Q4_K_M quantization and use 21.8 GB of system RAM.

Offloading allows you to run massive models, but it comes with a speed cost. System RAM is much slower than the GDDR7 memory on your graphics card. While models like Mistral Small 3.2 and Aria will run, their generation speed will decrease significantly. For the fastest performance, select models that fit entirely within the native video memory of your hardware.