Best local AI models for NVIDIA P104-100

4 GB GDDR5X. At a 4k context, 81 of the 233 models in our catalog with verified parameter counts fit fully, up to Lumina-Next / Lumina-Image 2.0 at 5B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 81 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Lumina-Next / Lumina-Image 2.05BQ4_K_M3.7 GB
CogVideoX 2B / 5B5BQ4_K_M3.7 GB
DeepSeek-VL24.5BQ5_K_M3.8 GB
DeepFloyd IF4.3BQ5_K_M3.7 GB
Phi-3.5-vision4.2BQ5_K_M3.6 GB
Qwen3 4B4BQ6_K3.9 GB
Gemma 3 4B4BQ6_K3.9 GB
Gemma 4 E4B4BQ6_K3.9 GB
MiniCPM 3 4B4BQ6_K3.9 GB
Danube 3 4B4BQ6_K3.9 GB
Fish Speech 1.5 / OpenAudio S14BQ6_K3.9 GB
Phi-4-mini-instruct3.8BQ6_K3.7 GB
Phi-3.5 Mini3.8BQ6_K3.7 GB
OmniGen / OmniGen23.8BQ6_K3.7 GB
SD Cascade (Würstchen v3)3.6BQ6_K3.5 GB
SDXL Turbo3.5BQ6_K3.4 GB
SDXL Lightning3.5BQ6_K3.4 GB
ACE-Step3.5BQ6_K3.4 GB
MusicGen small/medium/large3.3BQ6_K3.2 GB
SmolLM3 3B3BQ8_03.8 GB
Replit Code v1.5 3B3BQ8_03.8 GB
Kandinsky 3.13BQ8_03.8 GB
Voxtral Mini / Small3BQ8_03.8 GB
Orpheus TTS3BQ8_03.8 GB
Higgs Audio v23BQ8_03.8 GB
Allegro2.8BQ8_03.6 GB
Open-Sora Plan2.7BQ8_03.4 GB
LFM2 1.2B / 2.6B2.6BQ8_03.3 GB
Playground v2.52.6BQ8_03.3 GB
Stable Diffusion 3.5 Medium2.5BQ8_03.2 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
Stable Diffusion XL3.417B4.1 GB needed6.1 GB
Phi-3 Mini3.8B4.4 GB needed6.4 GB
Phi-4-multimodal5.6B4.1 GB needed6.1 GB
Magicoder-S-DS 6.7B6.7B4.9 GB needed6.9 GB
Mistral 7B7B5.7 GB needed7.7 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7B5.1 GB needed7.1 GB
OLMo 2 1B / 7B7B5.1 GB needed7.1 GB
Falcon 3 1B / 3B / 7B7B5.1 GB needed7.1 GB
Command R7B7B5.1 GB needed7.1 GB
OpenHermes 2.57B5.1 GB needed7.1 GB

How to read this

The NVIDIA P104-100 is a specialized mining card repurposed for compute tasks. It features 4 GB of GDDR5X memory. This memory limit is the main factor when running local AI models. To fit a model entirely on this hardware, the model size must stay under the 4 GB threshold. This page helps you choose the best models and quantization levels to maximize the hardware.

Quantization is a method that compresses model weights to save memory. The quant column shows the best format that fits within your 4 GB limit. For example, Q6_K and Q8_0 are high quality quantization levels. Using Q6_K allows you to run the 4B Gemma 3 4B or Qwen3 4B with 3.9 GB used. Smaller models like the 3B SmolLM3 3B can run at Q8_0 precision while using 3.8 GB of memory.

You can run larger models by using CPU offload if your computer has 32 GB of system RAM. This process splits the model between your graphics card and system memory. Offload allows you to run the 7B Mistral 7B using 5.7 GB of memory at Q4_K_M with 7.7 GB of system RAM. It also enables running the 7B Qwen2.5 0.5B / 1.5B / 3B / 7B which needs 5.1 GB at Q4_K_M and 7.1 GB of system RAM.

CPU offload comes with a performance cost. Moving data between the system RAM and the graphics card slows down generation speeds. While you can run the 7B Falcon 3 1B / 3B / 7B or Command R7B this way, the processing speed will be much slower than running a model fully on the graphics card. Fully local execution is always faster.

Context window size also impacts your memory usage. Running a model with a standard 4k context window requires extra memory space. If you fill the context window, the model might exceed the 4 GB limit and fail. You must budget your memory to leave room for both the model weights and the active conversation history.