Best local AI models for NVIDIA Quadro P3200 MAX-Q

6 GB GDDR5. At a 4k context, 114 of the 233 models in our catalog with verified parameter counts fit fully, up to Granite 3.3 2B / 8B at 8B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 114 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Granite 3.3 2B / 8B8BQ4_K_M5.9 GB
Ministral 3B / 8B8BQ4_K_M5.9 GB
InternLM 3 8B8BQ4_K_M5.9 GB
OpenCoder 1.5B / 8B8BQ4_K_M5.9 GB
Seed-Coder 8B8BQ4_K_M5.9 GB
MiniCPM-V 2.6 / MiniCPM-o 2.68BQ4_K_M5.9 GB
Idefics 3 8B8BQ4_K_M5.9 GB
Fuyu-8B8BQ4_K_M5.9 GB
Emu38BQ4_K_M5.9 GB
Stable Diffusion 3.5 Large / Turbo8BQ4_K_M5.9 GB
EXAONE 3.5 2.4B / 7.8B7.8BQ4_K_M5.7 GB
Mistral 7B7BQ4_K_M5.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7BQ5_K_M6 GB
OLMo 2 1B / 7B7BQ5_K_M6 GB
Falcon 3 1B / 3B / 7B7BQ5_K_M6 GB
Command R7B7BQ5_K_M6 GB
OpenHermes 2.57BQ5_K_M6 GB
Zephyr 7B Beta7BQ5_K_M6 GB
OpenChat 3.57BQ5_K_M6 GB
Starling LM 7B7BQ5_K_M6 GB
Codestral Mamba 7B7BQ5_K_M6 GB
CodeGemma 2B / 7B7BQ5_K_M6 GB
aiXcoder-7B7BQ5_K_M6 GB
Nxcode / CodeQwen 1.5 7B7BQ5_K_M6 GB
Janus-Pro 1B / 7B7BQ5_K_M6 GB
Ruyi-Mini-7B7BQ5_K_M6 GB
Qwen2-Audio 7B7BQ5_K_M6 GB
Qwen2.5-Omni 3B / 7B7BQ5_K_M6 GB
YuE7BQ5_K_M6 GB
Magicoder-S-DS 6.7B6.7BQ5_K_M5.7 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at Q4_K_MSystem RAM at 4k
Llama 3.1 8B8B6.4 GB needed8.4 GB
Chroma8.9B6.5 GB needed8.5 GB
Gemma 2 9B9B8 GB needed10 GB
Nemotron Nano 4B / 9B9B6.6 GB needed8.6 GB
GLM-4 9B / GLM-4.5-Air9B6.6 GB needed8.6 GB
Yi-Coder 1.5B / 9B9B6.6 GB needed8.6 GB
GLM-4-9B-Chat / CodeGeeX49B6.6 GB needed8.6 GB
GLM-4V-9B / GLM-4.1V-Thinking9B6.6 GB needed8.6 GB
Mochi 110B7.3 GB needed9.3 GB
Open-Sora 2.011B8.1 GB needed10.1 GB

How to read this

The NVIDIA Quadro P3200 MAX-Q is a mobile workstation graphics card equipped with 6 GB of GDDR5 video memory. This memory size determines the maximum size of the artificial intelligence models you can run entirely on your hardware. To run a model smoothly without system slowdowns, the model files and their operational memory must fit within this 6 GB limit.

The quantization column indicates the compression level applied to each model. For example, the Q4_K_M and Q5_K_M formats compress the model weights to four or five bits per parameter. This compression allows larger models to fit into your video memory. A Q4_K_M quant offers a great balance between model accuracy and memory savings.

Many capable models fit completely inside your 6 GB video memory. You can run Granite 3.3 8B, Ministral 8B, InternLM 3 8B, OpenCoder 8B, Seed-Coder 8B, MiniCPM-V 2.6 8B, Idefics 3 8B, Fuyu-8B, Emu3, and Stable Diffusion 3.5 Large at the Q4_K_M quant using 5.9 GB of memory. EXAONE 3.5 7.8B and Mistral 7B fit at Q4_K_M using 5.7 GB of memory. Magicoder-S-DS 6.7B also fits at Q5_K_M using 5.7 GB of memory.

You can also run several 7B models at the higher Q5_K_M quant which uses exactly 6 GB of memory. These models include Qwen2.5 7B, OLMo 2 7B, Falcon 3 7B, Command R7B, OpenHermes 2.5, Zephyr 7B Beta, OpenChat 3.5, Starling LM 7B, Codestral Mamba 7B, CodeGemma 7B, aiXcoder-7B, Nxcode 7B, Janus-Pro 7B, Ruyi-Mini-7B, Qwen2-Audio 7B, Qwen2.5-Omni 7B, and YuE.

When a model is too large for your 6 GB video memory, you can offload parts of it to your system RAM. This offloading requires a system with 32 GB of system RAM. For example, Llama 3.1 8B needs 6.4 GB at Q4_K_M which requires 8.4 GB of system RAM. Chroma 8.9B, Nemotron Nano 9B, GLM-4 9B, Yi-Coder 9B, GLM-4-9B-Chat, and GLM-4V-9B need 6.6 GB at Q4_K_M which requires 8.6 GB of system RAM. Gemma 2 9B needs 8 GB at Q4_K_M and 10 GB of system RAM. Mochi 1 10B needs 7.3 GB at Q4_K_M and 9.3 GB of system RAM. Open-Sora 2.0 11B needs 8.1 GB at Q4_K_M and 10.1 GB of system RAM.

Offloading models to system RAM comes with a performance cost. System RAM is much slower than the GDDR5 memory on your graphics card. This speed difference will slow down the generation speed of your model. Additionally, these memory calculations are based on a standard 4k context window. If you increase the context window to process longer texts, the model will require more memory and may exceed your limits.