Best local AI models for NVIDIA L40S
48 GB GDDR6. At a 4k context, 186 of the 233 models in our catalog with verified parameter counts fit fully, up to Jamba 1.5 Mini / Large at 52B parameters.
Check your own machine against every model →The largest models that fit fully
The 30 largest of the 186 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.
| Model | Parameters | Best quant that fits | Memory used at 4k |
|---|---|---|---|
| Jamba 1.5 Mini / Large | 52B | Q5_K_M | 44.3 GB |
| Llama 3.1 Nemotron 51B | 51B | Q5_K_M | 43.5 GB |
| Mixtral 8x7B | 47B | Q6_K | 46.2 GB |
| Seed-OSS 36B | 36B | Q8_0 | 45.8 GB |
| Qwen3.6-35B-A3B | 35B | Q8_0 | 44.5 GB |
| Command R (35B) | 35B | Q8_0 | 44.5 GB |
| Yi 1.5 9B / 34B | 34B | Q8_0 | 43.2 GB |
| Granite Code 3B to 34B | 34B | Q8_0 | 43.2 GB |
| LLaVA 1.5 / 1.6 (7B to 34B) | 34B | Q8_0 | 43.2 GB |
| Ovis 2 | 34B | Q8_0 | 43.2 GB |
| DeepSeek-Coder 1.3B / 6.7B / 33B | 33B | Q8_0 | 42 GB |
| WizardCoder 33B | 33B | Q8_0 | 42 GB |
| OTel 2.0 LLM 31B IT | 32.1B | Q8_0 | 44.9 GB |
| Qwen3 8B / 14B / 32B | 32B | Q8_0 | 40.7 GB |
| Qwen3.5 (dense variants) | 32B | Q8_0 | 40.7 GB |
| Aya Expanse 8B / 32B | 32B | Q8_0 | 40.7 GB |
| Granite 4.0 Small/Tiny | 32B | Q8_0 | 40.7 GB |
| Qwen2.5-Coder 0.5B to 32B | 32B | Q8_0 | 40.7 GB |
| Qwen3-30B-A3B | 30B | Q8_0 | 38.2 GB |
| Qwen3-Coder 30B-A3B | 30B | Q8_0 | 38.2 GB |
| Gemma 3 27B | 27B | Q8_0 | 34.3 GB |
| Gemma 3 4B/12B/27B (vision) | 27B | Q8_0 | 34.3 GB |
| Wan 2.2 / 2.5 | 27B | Q8_0 | 34.3 GB |
| Gemma 4 26B-A4B | 26B | Q8_0 | 33.1 GB |
| Gemma 4 (all sizes) | 26B | Q8_0 | 33.1 GB |
| Aria | 25B | Q8_0 | 31.8 GB |
| Mistral Small 3.2 | 24B | Q8_0 | 30.5 GB |
| Magistral Small | 24B | Q8_0 | 30.5 GB |
| Devstral Small 1.1 | 24B | Q8_0 | 30.5 GB |
| Solar Pro | 22B | Q8_0 | 28 GB |
How to read this
The NVIDIA L40S graphics card features 48 GB of GDDR6 memory. This dedicated memory determines the maximum size of the artificial intelligence models you can run locally. To load and run a model entirely on the graphics hardware, the total size of the model files must fit within this 48 GB limit.
The quantization column indicates the compression level used for each model. Quantization reduces the precision of model weights to save memory. A Q8_0 quant represents eight bit quantization, which preserves high output quality while reducing the footprint. For larger models, a Q5_K_M or Q6_K quant is used to fit the model within the hardware limits.
The Jamba 1.5 Mini / Large 52B model is the largest option that fits, using a Q5_K_M quant that requires 44.3 GB of memory. Other large options include the Llama 3.1 Nemotron 51B at Q5_K_M requiring 43.5 GB, and the Mixtral 8x7B 47B model at Q6_K requiring 46.2 GB. These configurations maximize the available memory of the card.
Several models fit comfortably using the high quality Q8_0 quantization level. The Seed-OSS 36B uses 45.8 GB, while the Qwen3.6-35B-A3B and Command R (35B) both use 44.5 GB. The Yi 1.5 9B / 34B, Granite Code 3B to 34B, LLaVA 1.5 / 1.6 (7B to 34B), and Ovis 2 models all use 43.2 GB at this quantization level.
Other notable Q8_0 options include the DeepSeek-Coder 1.3B / 6.7B / 33B and WizardCoder 33B at 42 GB. The OTel 2.0 LLM 31B IT uses 44.9 GB. The Qwen3 8B / 14B / 32B, Qwen3.5 (dense variants), Aya Expanse 8B / 32B, Granite 4.0 Small/Tiny, and Qwen2.5-Coder 0.5B to 32B all require 40.7 GB. The Gemma 3 27B and Wan 2.2 / 2.5 models require 34.3 GB.
Running models entirely on the graphics card memory ensures the fastest processing speeds. Offloading parts of a model to system RAM is not required for these configurations, which is beneficial because system memory transfer speeds are much slower. Keeping the entire model on the 48 GB GDDR6 memory prevents performance drops.
The memory usage figures are calculated based on a standard 4k context window. If you increase the context window to process longer documents, the memory required for the active session will increase. You must leave some of the 48 GB memory free to accommodate this additional context data during generation.