Best local AI models for NVIDIA RTX 4090 Laptop
16 GB GDDR6. At a 4k context, 155 of the 233 models in our catalog with verified parameter counts fit fully, up to gpt-oss-20b at 21B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 155 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| gpt-oss-20b | 21B | Q4_K_M | 15.4 GB |
| Reka Flash 3 | 21B | Q4_K_M | 15.4 GB |
| Qwen-Image | 20B | Q4_K_M | 14.6 GB |
| Qwen-Image-Edit | 20B | Q4_K_M | 14.6 GB |
| CogVLM2 | 19B | Q4_K_M | 13.9 GB |
| HunyuanImage 2.1 / 3.0 | 17B | Q5_K_M | 14.5 GB |
| Ling-Coder-Lite | 16.8B | Q5_K_M | 14.3 GB |
| DeepSeek-Coder-V2 16B / 236B | 16B | Q6_K | 15.7 GB |
| Kimi-VL A3B | 16B | Q6_K | 15.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q6_K | 14.8 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q6_K | 14.8 GB |
| Qwen2.5 14B | 14.7B | Q6_K | 15.3 GB |
| Phi-3 Medium | 14B | Q6_K | 13.8 GB |
| Phi-4 | 14B | Q6_K | 13.8 GB |
| Phi-4-reasoning / -plus | 14B | Q6_K | 13.8 GB |
| Wan 2.2 T2I | 14B | Q6_K | 13.8 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q6_K | 13.8 GB |
| SkyReels V2 | 14B | Q6_K | 13.8 GB |
| Vicuna 13B | 13B | Q6_K | 12.8 GB |
| HunyuanVideo | 13B | Q6_K | 12.8 GB |
| HunyuanVideo-Avatar | 13B | Q6_K | 12.8 GB |
| LTX-Video / LTX-2 | 13B | Q6_K | 12.8 GB |
| FramePack | 13B | Q6_K | 12.8 GB |
| FLUX.1 dev | 12B | FP8 / optimized | 14.4 GB |
| Gemma 3 12B | 12B | Q8_0 | 15.3 GB |
| Gemma 4 12B | 12B | Q8_0 | 15.3 GB |
| Mistral NeMo 12B | 12B | Q8_0 | 15.3 GB |
| Pixtral 12B | 12B | Q8_0 | 15.3 GB |
| FLUX.1 schnell | 12B | Q8_0 | 15.3 GB |
| FLUX.1 Kontext dev | 12B | Q8_0 | 15.3 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
| Mistral Small 3.2 | 24B | 17.6 GB needed | 19.6 GB |
| Magistral Small | 24B | 17.6 GB needed | 19.6 GB |
| Devstral Small 1.1 | 24B | 17.6 GB needed | 19.6 GB |
| Aria | 25B | 18.3 GB needed | 20.3 GB |
| Gemma 4 26B-A4B | 26B | 19 GB needed | 21 GB |
| Gemma 4 (all sizes) | 26B | 19 GB needed | 21 GB |
| Gemma 3 27B | 27B | 19.8 GB needed | 21.8 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | 19.8 GB needed | 21.8 GB |
How to read this
The NVIDIA RTX 4090 Laptop GPU features 16 GB of GDDR6 memory. This dedicated memory determines the maximum size of the AI model you can run locally. To run a model entirely on the graphics card, the model files and the context data must fit within this 16 GB limit. Keeping the entire model on the GPU ensures the fastest processing speeds.
The quantization column shows the compression level used to shrink these models. Quantization reduces the precision of the model weights to save space. For example, the 21B gpt-oss-20b and Reka Flash 3 models fit in 15.4 GB of memory using the Q4_K_M quantization. Other models like Qwen-Image and Qwen-Image-Edit use 14.6 GB at Q4_K_M. CogVLM2 fits in 13.9 GB at the same Q4_K_M level.
Higher precision quantizations are possible with slightly smaller models. HunyuanImage 2.1 / 3.0 uses 14.5 GB at Q5_K_M. Ling-Coder-Lite uses 14.3 GB at Q5_K_M. DeepSeek-Coder-V2 16B / 236B, Kimi-VL A3B, Apriel-1.5-15B-Thinker, and StarCoder2 3B / 7B / 15B can run at Q6_K precision. These Q6_K models use between 14.8 GB and 15.7 GB of memory.
Many popular 14B and 13B models also run at Q6_K precision. Qwen2.5 14B uses 15.3 GB. Phi-3 Medium, Phi-4, Phi-4-reasoning / -plus, Wan 2.2 T2I, Wan 2.1 (1.3B / 14B), and SkyReels V2 use 13.8 GB. Vicuna 13B, HunyuanVideo, HunyuanVideo-Avatar, LTX-Video / LTX-2, and FramePack use 12.8 GB. At the 12B size, FLUX.1 dev uses 14.4 GB at FP8 / optimized, while Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, and FLUX.1 Kontext dev use 15.3 GB at Q8_0.
When a model exceeds the 16 GB graphics memory, you must offload parts of it to your system RAM. This offloading allows you to run larger models but slows down the generation speed significantly. For these cases, we assume a system with 32 GB of system RAM. Solar Pro and Codestral 22B require 16.1 GB at Q4_K_M and use 18.1 GB of system RAM.
Larger offloaded models require even more system memory. Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 need 17.6 GB at Q4_K_M and use 19.6 GB of system RAM. Aria needs 18.3 GB at Q4_K_M and uses 20.3 GB of system RAM. Gemma 4 26B-A4B and Gemma 4 (all sizes) need 19 GB at Q4_K_M and use 21 GB of system RAM. Gemma 3 27B and Gemma 3 4B/12B/27B (vision) need 19.8 GB at Q4_K_M and use 21.8 GB of system RAM.
All memory calculations in our catalog assume a standard 4k context window. If you increase the context window to process longer texts, the model will require more memory. This extra memory usage might force you to use a lower quantization level or offload the model to system RAM.