Best local AI models for NVIDIA RTX 5080

16 GB GDDR7. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
gpt-oss-20b21BQ4_K_M15.4 GB
Reka Flash 321BQ4_K_M15.4 GB
Qwen-Image20BQ4_K_M14.6 GB
Qwen-Image-Edit20BQ4_K_M14.6 GB
CogVLM219BQ4_K_M13.9 GB
HunyuanImage 2.1 / 3.017BQ5_K_M14.5 GB
Ling-Coder-Lite16.8BQ5_K_M14.3 GB
DeepSeek-Coder-V2 16B / 236B16BQ6_K15.7 GB
Kimi-VL A3B16BQ6_K15.7 GB
Apriel-1.5-15B-Thinker15BQ6_K14.8 GB
StarCoder2 3B / 7B / 15B15BQ6_K14.8 GB
Qwen2.5 14B14.7BQ6_K15.3 GB
Phi-3 Medium14BQ6_K13.8 GB
Phi-414BQ6_K13.8 GB
Phi-4-reasoning / -plus14BQ6_K13.8 GB
Wan 2.2 T2I14BQ6_K13.8 GB
Wan 2.1 (1.3B / 14B)14BQ6_K13.8 GB
SkyReels V214BQ6_K13.8 GB
Vicuna 13B13BQ6_K12.8 GB
HunyuanVideo13BQ6_K12.8 GB
HunyuanVideo-Avatar13BQ6_K12.8 GB
LTX-Video / LTX-213BQ6_K12.8 GB
FramePack13BQ6_K12.8 GB
FLUX.1 dev12BFP8 / optimized14.4 GB
Gemma 3 12B12BQ8_015.3 GB
Gemma 4 12B12BQ8_015.3 GB
Mistral NeMo 12B12BQ8_015.3 GB
Pixtral 12B12BQ8_015.3 GB
FLUX.1 schnell12BQ8_015.3 GB
FLUX.1 Kontext dev12BQ8_015.3 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at Q4_K_MSystem RAM at 4k
Solar Pro22B16.1 GB needed18.1 GB
Codestral 22B22B16.1 GB needed18.1 GB
Mistral Small 3.224B17.6 GB needed19.6 GB
Magistral Small24B17.6 GB needed19.6 GB
Devstral Small 1.124B17.6 GB needed19.6 GB
Aria25B18.3 GB needed20.3 GB
Gemma 4 26B-A4B26B19 GB needed21 GB
Gemma 4 (all sizes)26B19 GB needed21 GB
Gemma 3 27B27B19.8 GB needed21.8 GB
Gemma 3 4B/12B/27B (vision)27B19.8 GB needed21.8 GB

How to read this

The NVIDIA RTX 5080 graphics card features 16 GB of GDDR7 memory. This dedicated memory determines the maximum size of the artificial intelligence models you can run locally. To load a model entirely onto the graphics card, the total size of the model files must remain below this limit. Keeping the model inside the graphics memory ensures the fastest processing speeds for your tasks.

The quantization column shows the compression level applied to each model. Quantization reduces the size of a model so it uses less memory. For example, the Q4_K_M quant represents a medium four bit quantization. The Q6_K quant represents a six bit quantization. The Q8_0 quant represents an eight bit quantization. Higher quantization levels like Q8_0 preserve more original model accuracy but require more memory.

With 16 GB of memory, you can run several large models locally. The gpt-oss-20b model and Reka Flash 3 are 21B models that fit using the Q4_K_M quant at 15.4 GB. Vision models like Qwen-Image and Qwen-Image-Edit are 20B models that use 14.6 GB at Q4_K_M. CogVLM2 is a 19B model using 13.9 GB at Q4_K_M. HunyuanImage 2.1 / 3.0 is a 17B model using 14.5 GB at the Q5_K_M quant.

You can also run highly accurate smaller models at higher quantization levels. The DeepSeek-Coder-V2 16B / 236B and Kimi-VL A3B models are 16B models using 15.7 GB at Q6_K. The Qwen2.5 14B model uses 15.3 GB at Q6_K. Models like Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, and Pixtral 12B are 12B models that fit at the Q8_0 quant using 15.3 GB.

When a model exceeds 16 GB, you must offload some parts to your system RAM. This offloading process allows you to run larger models but reduces processing speed. For a system with 32 GB of system RAM, you can run Solar Pro or Codestral 22B at Q4_K_M which needs 16.1 GB of memory and 18.1 GB of system RAM. Gemma 3 27B needs 19.8 GB at Q4_K_M and 21.8 GB of system RAM.

Memory calculations assume a standard four kilobyte context window. As you type longer prompts or generate longer responses, the context window expands. This expansion uses additional graphics memory. If your context window grows large, you may need to select a smaller model or a lower quantization level to prevent memory errors.