Best local AI models for NVIDIA RTX 5080 Laptop
16 GB GDDR7. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| gpt-oss-20b | 21B | Q4_K_M | 15.4 GB |
| Reka Flash 3 | 21B | Q4_K_M | 15.4 GB |
| Qwen-Image | 20B | Q4_K_M | 14.6 GB |
| Qwen-Image-Edit | 20B | Q4_K_M | 14.6 GB |
| CogVLM2 | 19B | Q4_K_M | 13.9 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q5_K_M | 14.5 GB |
| Ling-Coder-Lite | 16.8B | Q5_K_M | 14.3 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q6_K | 15.7 GB |
| Kimi-VL A3B | 16B | Q6_K | 15.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q6_K | 14.8 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q6_K | 14.8 GB |
| Qwen2.5 14B | 14.7B | Q6_K | 15.3 GB |
| Phi-3 Medium | 14B | Q6_K | 13.8 GB |
| Phi-4 | 14B | Q6_K | 13.8 GB |
| Phi-4-reasoning / -plus | 14B | Q6_K | 13.8 GB |
| Wan 2.2 T2I | 14B | Q6_K | 13.8 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q6_K | 13.8 GB |
| SkyReels V2 | 14B | Q6_K | 13.8 GB |
| Vicuna 13B | 13B | Q6_K | 12.8 GB |
| HunyuanVideo | 13B | Q6_K | 12.8 GB |
| HunyuanVideo-Avatar | 13B | Q6_K | 12.8 GB |
| LTX-Video / LTX-2 | 13B | Q6_K | 12.8 GB |
| FramePack | 13B | Q6_K | 12.8 GB |
| FLUX.1 dev | 12B | FP8 / optimized | 14.4 GB |
| Gemma 3 12B | 12B | Q8_0 | 15.3 GB |
| Gemma 4 12B | 12B | Q8_0 | 15.3 GB |
| Mistral NeMo 12B | 12B | Q8_0 | 15.3 GB |
| Pixtral 12B | 12B | Q8_0 | 15.3 GB |
| FLUX.1 schnell | 12B | Q8_0 | 15.3 GB |
| FLUX.1 Kontext dev | 12B | Q8_0 | 15.3 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
| Mistral Small 3.2 | 24B | 17.6 GB needed | 19.6 GB |
| Magistral Small | 24B | 17.6 GB needed | 19.6 GB |
| Devstral Small 1.1 | 24B | 17.6 GB needed | 19.6 GB |
| Aria | 25B | 18.3 GB needed | 20.3 GB |
| Gemma 4 26B-A4B | 26B | 19 GB needed | 21 GB |
| Gemma 4 (all sizes) | 26B | 19 GB needed | 21 GB |
| Gemma 3 27B | 27B | 19.8 GB needed | 21.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | 19.8 GB needed | 21.8 GB |
How to read this
The NVIDIA RTX 5080 Laptop GPU features 16 GB of GDDR7 memory. This dedicated video memory determines the maximum size of the artificial intelligence models you can run locally. To run a model entirely on your graphics hardware, the total memory used by the model must remain under this 16 GB limit. Keeping the model fully inside the video memory ensures the fastest possible processing speeds.
The quantization column shows the compression level applied to each model. Raw models are often too large for consumer hardware, so they are compressed into smaller formats called quants. For example, the 21B gpt-oss-20b and Reka Flash 3 models fit into 15.4 GB of memory when using the Q4_K_M quantization. Models like the 15B StarCoder2 3B / 7B / 15B and Apriel-1.5-15B-Thinker use the Q6_K quantization, which requires 14.8 GB of memory.
Larger models can still run by offloading some data to your system RAM. If you have 32 GB of system RAM, you can run models that exceed the 16 GB video memory limit. For instance, the 22B Codestral 22B and Solar Pro models need 16.1 GB at Q4_K_M, which requires 18.1 GB of system RAM. The 24B Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 models need 17.6 GB at Q4_K_M, requiring 19.6 GB of system RAM.
Higher offload demands will utilize more of your system memory. The 25B Aria model needs 18.3 GB at Q4_K_M and uses 20.3 GB of system RAM. The 26B Gemma 4 26B-A4B and Gemma 4 (all sizes) models need 19 GB at Q4_K_M, which uses 21 GB of system RAM. The largest offload options include the 27B Gemma 3 27B and Gemma 3 4B/12B/27B (vision) models, which need 19.8 GB at Q4_K_M and utilize 21.8 GB of system RAM. Offloading data to system RAM slows down generation speeds.
When selecting a model, you must also consider the context window. The memory figures listed for these models are calculated using a standard 4k context window. If you increase the context length to process longer documents or larger chat histories, the memory usage will rise. You may need to select a smaller model size or a lower quantization level to prevent running out of video memory during long conversations.