Best local AI models for NVIDIA Quadro T2000 MAX-Q

4 GB GDDR5. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Lumina-Next / Lumina-Image 2.05BQ4_K_M3.7 GB
CogVideoX 2B / 5B5BQ4_K_M3.7 GB
DeepSeek-VL24.5BQ5_K_M3.8 GB
DeepFloyd IF4.3BQ5_K_M3.7 GB
Phi-3.5-vision4.2BQ5_K_M3.6 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-4-mini-instruct3.8BQ6_K3.7 GB
Phi-3.5 Mini3.8BQ6_K3.7 GB
OmniGen / OmniGen23.8BQ6_K3.7 GB
SD Cascade (Würstchen v3)3.6BQ6_K3.5 GB
SDXL Turbo3.5BQ6_K3.4 GB
SDXL Lightning3.5BQ6_K3.4 GB
ACE-Step3.5BQ6_K3.4 GB
MusicGen small/medium/large3.3BQ6_K3.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB
Open-Sora Plan2.7BQ8_03.4 GB
LFM2 1.2B / 2.6B2.6BQ8_03.3 GB
Playground v2.52.6BQ8_03.3 GB
Stable Diffusion 3.5 Medium2.5BQ8_03.2 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB

How to read this

The NVIDIA Quadro T2000 MAX-Q is an efficient mobile workstation graphics card equipped with 4 GB of GDDR5 video memory. This dedicated VRAM pool determines the maximum size of the neural network models you can run entirely on your GPU. Keeping a model fully inside the 4 GB VRAM limit ensures the fastest processing speeds for text generation, image creation, and audio synthesis.

To fit inside this hardware limit, models use quantization to reduce their file size. The quant column shows the specific compression level used to make a model fit your GPU. For example, a Q6_K quant represents a highly accurate six bit quantization, while a Q8_0 quant offers near lossless quality at eight bits. Lower quantization levels like Q4_K_M compress the model further to let larger architectures fit into your available memory.

Several capable models fit completely within your 4 GB VRAM limit. For image generation, Lumina-Next or Lumina-Image 2.0 and CogVideoX 2B or 5B fit at 3.7 GB using the Q4_K_M quant. Vision tasks can be handled by DeepSeek-VL2 using 3.8 GB at Q5_K_M, or Phi-3.5-vision using 3.6 GB at Q5_K_M. For text, Qwen3 4B, Gemma 3 4B, Gemma 4 E4B, MiniCPM 3 4B, and Danube 3 4B all run at Q6_K using 3.9 GB of VRAM.

When a model exceeds 4 GB, you can offload parts of it to your system RAM. This CPU offload process requires a host system with 32 GB of system RAM. Offloading lets you run larger models like Mistral 7B, which needs 5.7 GB at Q4_K_M and 7.7 GB of system RAM. Other offload options include Qwen2.5 7B, OLMo 2 7B, Falcon 3 7B, Command R7B, and OpenHermes 2.5, which all require 5.1 GB at Q4_K_M and 7.1 GB of system RAM.

Running models close to your VRAM limit requires careful attention to your context window settings. The memory figures listed here assume a standard 4k context window. Increasing your context length beyond this limit will consume additional VRAM for the key value cache, which can cause out of memory errors or force your system to slow down significantly during long conversations.