Best local AI models for NVIDIA RTX 3060 12GB
12 GB GDDR6. At a 4k context, 147 of the 233 models in our catalog with verified parameter counts fit fully, up to DeepSeek-Coder-V2 16B / 236B at 16B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 147 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| DeepSeek-Coder-V2 16B / 236B | 16B | Q4_K_M | 11.7 GB |
| Kimi-VL A3B | 16B | Q4_K_M | 11.7 GB |
| Apriel-1.5-15B-Thinker | 15B | Q4_K_M | 11 GB |
| StarCoder2 3B / 7B / 15B | 15B | Q4_K_M | 11 GB |
| Qwen2.5 14B | 14.7B | Q4_K_M | 11.6 GB |
| Phi-3 Medium | 14B | Q5_K_M | 11.9 GB |
| Phi-4 | 14B | Q5_K_M | 11.9 GB |
| Phi-4-reasoning / -plus | 14B | Q5_K_M | 11.9 GB |
| Wan 2.2 T2I | 14B | Q5_K_M | 11.9 GB |
| Wan 2.1 (1.3B / 14B) | 14B | Q5_K_M | 11.9 GB |
| SkyReels V2 | 14B | Q5_K_M | 11.9 GB |
| Vicuna 13B | 13B | Q5_K_M | 11.1 GB |
| HunyuanVideo | 13B | Q5_K_M | 11.1 GB |
| HunyuanVideo-Avatar | 13B | Q5_K_M | 11.1 GB |
| LTX-Video / LTX-2 | 13B | Q5_K_M | 11.1 GB |
| FramePack | 13B | Q5_K_M | 11.1 GB |
| Gemma 3 12B | 12B | Q6_K | 11.8 GB |
| Gemma 4 12B | 12B | Q6_K | 11.8 GB |
| Mistral NeMo 12B | 12B | Q6_K | 11.8 GB |
| Pixtral 12B | 12B | Q6_K | 11.8 GB |
| FLUX.1 schnell | 12B | Q6_K | 11.8 GB |
| FLUX.1 Kontext dev | 12B | Q6_K | 11.8 GB |
| FLUX.1 Krea dev | 12B | Q6_K | 11.8 GB |
| Open-Sora 2.0 | 11B | Q6_K | 10.8 GB |
| Mochi 1 | 10B | Q6_K | 9.8 GB |
| Gemma 2 9B | 9B | Q6_K | 10.3 GB |
| Nemotron Nano 4B / 9B | 9B | Q8_0 | 11.4 GB |
| GLM-4 9B / GLM-4.5-Air | 9B | Q8_0 | 11.4 GB |
| Yi-Coder 1.5B / 9B | 9B | Q8_0 | 11.4 GB |
| GLM-4-9B-Chat / CodeGeeX4 | 9B | Q8_0 | 11.4 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at FP8 / optimized | System RAM at 4k |
|---|---|---|---|
| FLUX.1 dev | 12B | 14.4 GB needed | 16.4 GB |
| Ling-Coder-Lite | 16.8B | 12.3 GB needed | 14.3 GB |
| HunyuanImage 2.1 / 3.0 | 17B | 12.4 GB needed | 14.4 GB |
| CogVLM2 | 19B | 13.9 GB needed | 15.9 GB |
| Qwen-Image | 20B | 14.6 GB needed | 16.6 GB |
| Qwen-Image-Edit | 20B | 14.6 GB needed | 16.6 GB |
| gpt-oss-20b | 21B | 15.4 GB needed | 17.4 GB |
| Reka Flash 3 | 21B | 15.4 GB needed | 17.4 GB |
| Solar Pro | 22B | 16.1 GB needed | 18.1 GB |
| Codestral 22B | 22B | 16.1 GB needed | 18.1 GB |
How to read this
The NVIDIA RTX 3060 graphics card features 12 GB of GDDR6 memory. This onboard memory determines the size of the artificial intelligence models you can run locally. To achieve fast processing speeds, the entire model must fit directly inside this video memory. If a model exceeds this limit, your system must transfer data between the graphics card and system memory, which slows down performance.
Quantization is a method that compresses model files to save space. The quantization column shows the best format to balance size and quality. For example, the 16B DeepSeek-Coder-V2 and Kimi-VL A3B models fit within 11.7 GB using the Q4_K_M quantization. The 15B Apriel-1.5-15B-Thinker and StarCoder2 models use 11 GB with the same Q4_K_M format. Qwen2.5 14B uses 11.6 GB, while Phi-3 Medium, Phi-4, Phi-4-reasoning / -plus, Wan 2.2 T2I, Wan 2.1 (1.3B / 14B), and SkyReels V2 use 11.9 GB at the Q5_K_M quantization level.
Other models fit comfortably within the 12 GB limit at higher quantization levels. Vicuna 13B, HunyuanVideo, HunyuanVideo-Avatar, LTX-Video / LTX-2, and FramePack use 11.1 GB at Q5_K_M. Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, FLUX.1 Kontext dev, and FLUX.1 Krea dev use 11.8 GB at Q6_K. Open-Sora 2.0 uses 10.8 GB at Q6_K, while Mochi 1 uses 9.8 GB. Gemma 2 9B fits at Q6_K using 10.3 GB. Nemotron Nano 4B / 9B, GLM-4 9B / GLM-4.5-Air, Yi-Coder 1.5B / 9B, and GLM-4-9B-Chat / CodeGeeX4 use 11.4 GB at the Q8_0 quantization level.
When a model is too large for the 12 GB of video memory, you can offload parts of it to your system RAM. This process requires at least 32 GB of system RAM to work. For instance, FLUX.1 dev requires 14.4 GB at FP8 or optimized settings, which uses 16.4 GB of system RAM. Ling-Coder-Lite needs 12.3 GB at Q4_K_M and uses 14.3 GB of system RAM. HunyuanImage 2.1 / 3.0 needs 12.4 GB at Q4_K_M and uses 14.4 GB of system RAM. CogVLM2 needs 13.9 GB at Q4_K_M and uses 15.9 GB of system RAM.
Larger models demand even more system memory during offloading. Qwen-Image and Qwen-Image-Edit need 14.6 GB at Q4_K_M, using 16.6 GB of system RAM. Both gpt-oss-20b and Reka Flash 3 need 15.4 GB at Q4_K_M, using 17.4 GB of system RAM. Solar Pro and Codestral 22B need 16.1 GB at Q4_K_M, which uses 18.1 GB of system RAM. Offloading allows these large models to run, but the transfer of data between components reduces the overall generation speed.
You must also consider the memory cost of context. The listed memory figures assume a standard 4k context window. As your conversation grows longer, the model requires more memory to remember the history. Running a model very close to the 12 GB limit of your RTX 3060 may cause out of memory errors if you exceed this context limit.