Best local AI models for NVIDIA RTX 3060 12GB

12 GB GDDR6. At a 4k context, 147 of the 233 models in our catalog with verified parameter counts fit fully, up to DeepSeek-Coder-V2 16B / 236B at 16B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 147 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
DeepSeek-Coder-V2 16B / 236B16BQ4_K_M11.7 GB
Kimi-VL A3B16BQ4_K_M11.7 GB
Apriel-1.5-15B-Thinker15BQ4_K_M11 GB
StarCoder2 3B / 7B / 15B15BQ4_K_M11 GB
Qwen2.5 14B14.7BQ4_K_M11.6 GB
Phi-3 Medium14BQ5_K_M11.9 GB
Phi-414BQ5_K_M11.9 GB
Phi-4-reasoning / -plus14BQ5_K_M11.9 GB
Wan 2.2 T2I14BQ5_K_M11.9 GB
Wan 2.1 (1.3B / 14B)14BQ5_K_M11.9 GB
SkyReels V214BQ5_K_M11.9 GB
Vicuna 13B13BQ5_K_M11.1 GB
HunyuanVideo13BQ5_K_M11.1 GB
HunyuanVideo-Avatar13BQ5_K_M11.1 GB
LTX-Video / LTX-213BQ5_K_M11.1 GB
FramePack13BQ5_K_M11.1 GB
Gemma 3 12B12BQ6_K11.8 GB
Gemma 4 12B12BQ6_K11.8 GB
Mistral NeMo 12B12BQ6_K11.8 GB
Pixtral 12B12BQ6_K11.8 GB
FLUX.1 schnell12BQ6_K11.8 GB
FLUX.1 Kontext dev12BQ6_K11.8 GB
FLUX.1 Krea dev12BQ6_K11.8 GB
Open-Sora 2.011BQ6_K10.8 GB
Mochi 110BQ6_K9.8 GB
Gemma 2 9B9BQ6_K10.3 GB
Nemotron Nano 4B / 9B9BQ8_011.4 GB
GLM-4 9B / GLM-4.5-Air9BQ8_011.4 GB
Yi-Coder 1.5B / 9B9BQ8_011.4 GB
GLM-4-9B-Chat / CodeGeeX49BQ8_011.4 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
FLUX.1 dev12B14.4 GB needed16.4 GB
Ling-Coder-Lite16.8B12.3 GB needed14.3 GB
HunyuanImage 2.1 / 3.017B12.4 GB needed14.4 GB
CogVLM219B13.9 GB needed15.9 GB
Qwen-Image20B14.6 GB needed16.6 GB
Qwen-Image-Edit20B14.6 GB needed16.6 GB
gpt-oss-20b21B15.4 GB needed17.4 GB
Reka Flash 321B15.4 GB needed17.4 GB
Solar Pro22B16.1 GB needed18.1 GB
Codestral 22B22B16.1 GB needed18.1 GB

How to read this

The NVIDIA RTX 3060 graphics card features 12 GB of GDDR6 memory. This onboard memory determines the size of the artificial intelligence models you can run locally. To achieve fast processing speeds, the entire model must fit directly inside this video memory. If a model exceeds this limit, your system must transfer data between the graphics card and system memory, which slows down performance.

Quantization is a method that compresses model files to save space. The quantization column shows the best format to balance size and quality. For example, the 16B DeepSeek-Coder-V2 and Kimi-VL A3B models fit within 11.7 GB using the Q4_K_M quantization. The 15B Apriel-1.5-15B-Thinker and StarCoder2 models use 11 GB with the same Q4_K_M format. Qwen2.5 14B uses 11.6 GB, while Phi-3 Medium, Phi-4, Phi-4-reasoning / -plus, Wan 2.2 T2I, Wan 2.1 (1.3B / 14B), and SkyReels V2 use 11.9 GB at the Q5_K_M quantization level.

Other models fit comfortably within the 12 GB limit at higher quantization levels. Vicuna 13B, HunyuanVideo, HunyuanVideo-Avatar, LTX-Video / LTX-2, and FramePack use 11.1 GB at Q5_K_M. Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, FLUX.1 Kontext dev, and FLUX.1 Krea dev use 11.8 GB at Q6_K. Open-Sora 2.0 uses 10.8 GB at Q6_K, while Mochi 1 uses 9.8 GB. Gemma 2 9B fits at Q6_K using 10.3 GB. Nemotron Nano 4B / 9B, GLM-4 9B / GLM-4.5-Air, Yi-Coder 1.5B / 9B, and GLM-4-9B-Chat / CodeGeeX4 use 11.4 GB at the Q8_0 quantization level.

When a model is too large for the 12 GB of video memory, you can offload parts of it to your system RAM. This process requires at least 32 GB of system RAM to work. For instance, FLUX.1 dev requires 14.4 GB at FP8 or optimized settings, which uses 16.4 GB of system RAM. Ling-Coder-Lite needs 12.3 GB at Q4_K_M and uses 14.3 GB of system RAM. HunyuanImage 2.1 / 3.0 needs 12.4 GB at Q4_K_M and uses 14.4 GB of system RAM. CogVLM2 needs 13.9 GB at Q4_K_M and uses 15.9 GB of system RAM.

Larger models demand even more system memory during offloading. Qwen-Image and Qwen-Image-Edit need 14.6 GB at Q4_K_M, using 16.6 GB of system RAM. Both gpt-oss-20b and Reka Flash 3 need 15.4 GB at Q4_K_M, using 17.4 GB of system RAM. Solar Pro and Codestral 22B need 16.1 GB at Q4_K_M, which uses 18.1 GB of system RAM. Offloading allows these large models to run, but the transfer of data between components reduces the overall generation speed.

You must also consider the memory cost of context. The listed memory figures assume a standard 4k context window. As your conversation grows longer, the model requires more memory to remember the history. Running a model very close to the 12 GB limit of your RTX 3060 may cause out of memory errors if you exceed this context limit.