Best local AI models for NVIDIA RTX 4080 SUPER
16 GB GDDR6X. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| gpt-oss-20b | 21B | Q4_K_M | 15.4 GB |
| Reka Flash 3 | 21B | Q4_K_M | 15.4 GB |
| Qwen-Image | 20B | Q4_K_M | 14.6 GB |
| Qwen-Image-Edit | 20B | Q4_K_M | 14.6 GB |
| CogVLM2 | 19B | Q4_K_M | 13.9 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q5_K_M | 14.5 GB |
| Ling-Coder-Lite | 16.8B | Q5_K_M | 14.3 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q6_K | 15.7 GB |
| Kimi-VL A3B | 16B | Q6_K | 15.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q6_K | 14.8 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q6_K | 14.8 GB |
| Qwen2.5 14B | 14.7B | Q6_K | 15.3 GB |
| Phi-3 Medium | 14B | Q6_K | 13.8 GB |
| Phi-4 | 14B | Q6_K | 13.8 GB |
| Phi-4-reasoning / -plus | 14B | Q6_K | 13.8 GB |
| Wan 2.2 T2I | 14B | Q6_K | 13.8 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q6_K | 13.8 GB |
| SkyReels V2 | 14B | Q6_K | 13.8 GB |
| Vicuna 13B | 13B | Q6_K | 12.8 GB |
| HunyuanVideo | 13B | Q6_K | 12.8 GB |
| HunyuanVideo-Avatar | 13B | Q6_K | 12.8 GB |
| LTX-Video / LTX-2 | 13B | Q6_K | 12.8 GB |
| FramePack | 13B | Q6_K | 12.8 GB |
| FLUX.1 dev | 12B | FP8 / optimized | 14.4 GB |
| Gemma 3 12B | 12B | Q8_0 | 15.3 GB |
| Gemma 4 12B | 12B | Q8_0 | 15.3 GB |
| Mistral NeMo 12B | 12B | Q8_0 | 15.3 GB |
| Pixtral 12B | 12B | Q8_0 | 15.3 GB |
| FLUX.1 schnell | 12B | Q8_0 | 15.3 GB |
| FLUX.1 Kontext dev | 12B | Q8_0 | 15.3 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
| Mistral Small 3.2 | 24B | 17.6 GB needed | 19.6 GB |
| Magistral Small | 24B | 17.6 GB needed | 19.6 GB |
| Devstral Small 1.1 | 24B | 17.6 GB needed | 19.6 GB |
| Aria | 25B | 18.3 GB needed | 20.3 GB |
| Gemma 4 26B-A4B | 26B | 19 GB needed | 21 GB |
| Gemma 4 (all sizes) | 26B | 19 GB needed | 21 GB |
| Gemma 3 27B | 27B | 19.8 GB needed | 21.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | 19.8 GB needed | 21.8 GB |
How to read this
The NVIDIA RTX 4080 SUPER features 16 GB of GDDR6X memory. This dedicated video memory determines which local AI models you can run entirely on your graphics hardware. To run a model at full speed, both the model weights and the active context window must fit within this 16 GB limit. When a model fits completely in your video memory, you get the fastest possible generation speeds.
The quant column shows the quantization level used to compress each model. Raw models are often too large for consumer hardware, so they are compressed to smaller bit widths. For example, the Q4_K_M quant uses approximately four bits per weight, while Q8_0 uses eight bits. Higher quants like Q6_K and Q8_0 preserve more original model accuracy but require more memory. Lower quants like Q4_K_M allow larger models to fit into your 16 GB limit.
The largest models that fit completely within your video memory include gpt-oss-20b and Reka Flash 3. Both are 21B models that use 15.4 GB of memory at the Q4_K_M quant. You can also run Qwen-Image and Qwen-Image-Edit at Q4_K_M, which use 14.6 GB. The 19B CogVLM2 model fits at Q4_K_M using 13.9 GB. HunyuanImage 2.1 / 3.0 fits at Q5_K_M using 14.5 GB, and Ling-Coder-Lite fits at Q5_K_M using 14.3 GB.
For maximum precision, you can run slightly smaller models at higher quants. DeepSeek-Coder-V2 16B / 236B and Kimi-VL A3B fit at Q6_K using 15.7 GB. Apriel-1.5-15B-Thinker and StarCoder2 3B / 7B / 15B fit at Q6_K using 14.8 GB. Qwen2.5 14B fits at Q6_K using 15.3 GB. Phi-3 Medium, Phi-4, Phi-4-reasoning / -plus, Wan 2.2 T2I, Wan 2.1 (1.3B / 14B), and SkyReels V2 all fit at Q6_K using 13.8 GB. Vicuna 13B, HunyuanVideo, HunyuanVideo-Avatar, LTX-Video / LTX-2, and FramePack fit at Q6_K using 12.8 GB.
Popular 12B models can run at the high Q8_0 quant. Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, and FLUX.1 Kontext dev all use 15.3 GB at Q8_0. FLUX.1 dev fits at an optimized FP8 quant using 14.4 GB. Note that these memory figures are calculated with a standard 4k context window. If you increase the context window to process longer documents, the system will require additional video memory.
If you want to run larger models, you can offload some layers to your system RAM. This requires at least 32 GB of system memory. Offloading allows you to run Solar Pro or Codestral 22B, which need 16.1 GB at Q4_K_M and 18.1 GB of system RAM. Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 need 17.6 GB at Q4_K_M and 19.6 GB of system RAM. Aria needs 18.3 GB at Q4_K_M and 20.3 GB of system RAM.
The largest offload options include Gemma 4 26B-A4B and Gemma 4 (all sizes), which need 19 GB at Q4_K_M and 21 GB of system RAM. Gemma 3 27B and Gemma 3 4B/12B/27B (vision) need 19.8 GB at Q4_K_M and 21.8 GB of system RAM. While offloading lets you run these larger models, it comes with a cost. Moving data between system RAM and video memory is slow, which significantly reduces your generation speed.