Best local AI models for NVIDIA P106-100
6 GB GDDR5. At a 4k context, 114 of the 233 models in our catalog with verified parameter counts fit fully, up to Granite 3.3 2B / 8B at 8B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 114 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Granite 3.3 2B / 8B | 8B | Q4_K_M | 5.9 GB |
| Ministral 3B / 8B | 8B | Q4_K_M | 5.9 GB |
| InternLM 3 8B | 8B | Q4_K_M | 5.9 GB |
| OpenCoder 1.5B / 8B | 8B | Q4_K_M | 5.9 GB |
| Seed-Coder 8B | 8B | Q4_K_M | 5.9 GB |
| MiniCPM-V 2.6 / MiniCPM-o 2.6 | 8B | Q4_K_M | 5.9 GB |
| Idefics 3 8B | 8B | Q4_K_M | 5.9 GB |
| Fuyu-8B | 8B | Q4_K_M | 5.9 GB |
| Emu3 | 8B | Q4_K_M | 5.9 GB |
| Stable Diffusion 3.5 Large / Turbo | 8B | Q4_K_M | 5.9 GB |
| EXAONE 3.5 2.4B / 7.8B | 7.8B | Q4_K_M | 5.7 GB |
| Mistral 7B | 7B | Q4_K_M | 5.7 GB |
| Qwen2.5 0.5B / 1.5B / 3B / 7B | 7B | Q5_K_M | 6 GB |
| OLMo 2 1B / 7B | 7B | Q5_K_M | 6 GB |
| Falcon 3 1B / 3B / 7B | 7B | Q5_K_M | 6 GB |
| Command R7B | 7B | Q5_K_M | 6 GB |
| OpenHermes 2.5 | 7B | Q5_K_M | 6 GB |
| Zephyr 7B Beta | 7B | Q5_K_M | 6 GB |
| OpenChat 3.5 | 7B | Q5_K_M | 6 GB |
| Starling LM 7B | 7B | Q5_K_M | 6 GB |
| Codestral Mamba 7B | 7B | Q5_K_M | 6 GB |
| CodeGemma 2B / 7B | 7B | Q5_K_M | 6 GB |
| aiXcoder-7B | 7B | Q5_K_M | 6 GB |
| Nxcode / CodeQwen 1.5 7B | 7B | Q5_K_M | 6 GB |
| Janus-Pro 1B / 7B | 7B | Q5_K_M | 6 GB |
| Ruyi-Mini-7B | 7B | Q5_K_M | 6 GB |
| Qwen2-Audio 7B | 7B | Q5_K_M | 6 GB |
| Qwen2.5-Omni 3B / 7B | 7B | Q5_K_M | 6 GB |
| YuE | 7B | Q5_K_M | 6 GB |
| Magicoder-S-DS 6.7B | 6.7B | Q5_K_M | 5.7 GB |
Close, but only with CPU offload
These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.
| Model | Parameters | Memory at Q4_K_M | System RAM at 4k |
|---|---|---|---|
| Llama 3.1 8B | 8B | 6.4 GB needed | 8.4 GB |
| Chroma | 8.9B | 6.5 GB needed | 8.5 GB |
| Gemma 2 9B | 9B | 8 GB needed | 10 GB |
| Nemotron Nano 4B / 9B | 9B | 6.6 GB needed | 8.6 GB |
| GLM-4 9B / GLM-4.5-Air | 9B | 6.6 GB needed | 8.6 GB |
| Yi-Coder 1.5B / 9B | 9B | 6.6 GB needed | 8.6 GB |
| GLM-4-9B-Chat / CodeGeeX4 | 9B | 6.6 GB needed | 8.6 GB |
| GLM-4V-9B / GLM-4.1V-Thinking | 9B | 6.6 GB needed | 8.6 GB |
| Mochi 1 | 10B | 7.3 GB needed | 9.3 GB |
| Open-Sora 2.0 | 11B | 8.1 GB needed | 10.1 GB |
How to read this
The NVIDIA P106-100 is a specialized graphics card equipped with 6 GB GDDR5 video memory. This memory size determines the maximum size of the artificial intelligence models you can run directly on the hardware. To fit models within this limit, you must use quantized versions. Quantization reduces the precision of model weights to save space. The best quant column shows the optimal compromise between model accuracy and memory usage for this specific card.
For models that fit entirely within the 6 GB video memory, several options exist. You can run Granite 3.3 8B, Ministral 8B, InternLM 3 8B, OpenCoder 8B, Seed-Coder 8B, MiniCPM-V 2.6, MiniCPM-o 2.6, Idefics 3 8B, Fuyu-8B, Emu3, and Stable Diffusion 3.5 Large at the Q4_K_M quantization level. Each of these models uses 5.9 GB of video memory. EXAONE 3.5 7.8B and Mistral 7B also fit at the Q4_K_M quantization level, using 5.7 GB of video memory.
Other models can run at the higher quality Q5_K_M quantization level. Qwen2.5 7B, OLMo 2 7B, Falcon 3 7B, Command R7B, OpenHermes 2.5, Zephyr 7B Beta, OpenChat 3.5, Starling LM 7B, Codestral Mamba 7B, CodeGemma 7B, aiXcoder-7B, Nxcode, CodeQwen 1.5 7B, Janus-Pro 7B, Ruyi-Mini-7B, Qwen2-Audio 7B, Qwen2.5-Omni 7B, and YuE all utilize exactly 6 GB of video memory. Magicoder-S-DS 6.7B requires 5.7 GB of video memory at the Q5_K_M quantization level.
When a model is too large for the 6 GB video memory, you can offload parts of it to your system RAM. This process assumes you have 32 GB of system RAM available. Offloading allows you to run larger models, but it costs performance because system RAM is much slower than the GDDR5 memory on the card. This speed penalty will slow down the generation of text or images.
Using CPU offload, you can run Llama 3.1 8B, which needs 6.4 GB at Q4_K_M and 8.4 GB of system RAM. Chroma requires 6.5 GB at Q4_K_M and 8.5 GB of system RAM. Gemma 2 9B needs 8 GB at Q4_K_M and 10 GB of system RAM. Nemotron Nano 9B, GLM-4 9B, GLM-4.5-Air, Yi-Coder 9B, GLM-4-9B-Chat, CodeGeeX4, GLM-4V-9B, and GLM-4.1V-Thinking all need 6.6 GB at Q4_K_M and 8.6 GB of system RAM. Mochi 1 needs 7.3 GB at Q4_K_M and 9.3 GB of system RAM. Open-Sora 2.0 needs 8.1 GB at Q4_K_M and 10.1 GB of system RAM.
All memory calculations are based on a standard 4k context window. If you increase the context window to process longer texts, the memory requirements will rise. This extra memory usage might force you to use a lower quantization level or offload more layers to your system RAM.