Best local AI models for NVIDIA L4

24 GB GDDR6. At a 4k context, 173 of the 233 models in our catalog with verified parameter counts fit fully, up to Qwen3 8B / 14B / 32B at 32B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 173 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Qwen3 8B / 14B / 32B32BQ4_K_M23.4 GB
Qwen3.5 (dense variants)32BQ4_K_M23.4 GB
Aya Expanse 8B / 32B32BQ4_K_M23.4 GB
Granite 4.0 Small/Tiny32BQ4_K_M23.4 GB
Qwen2.5-Coder 0.5B to 32B32BQ4_K_M23.4 GB
Qwen3-30B-A3B30BQ4_K_M22 GB
Qwen3-Coder 30B-A3B30BQ4_K_M22 GB
Gemma 3 27B27BQ5_K_M23 GB
Gemma 3 4B/12B/27B (vision)27BQ5_K_M23 GB
Wan 2.2 / 2.527BQ5_K_M23 GB
Gemma 4 26B-A4B26BQ5_K_M22.2 GB
Gemma 4 (all sizes)26BQ5_K_M22.2 GB
Aria25BQ5_K_M21.3 GB
Mistral Small 3.224BQ6_K23.6 GB
Magistral Small24BQ6_K23.6 GB
Devstral Small 1.124BQ6_K23.6 GB
Solar Pro22BQ6_K21.6 GB
Codestral 22B22BQ6_K21.6 GB
gpt-oss-20b21BQ6_K20.7 GB
Reka Flash 321BQ6_K20.7 GB
Qwen-Image20BQ6_K19.7 GB
Qwen-Image-Edit20BQ6_K19.7 GB
CogVLM219BQ6_K18.7 GB
HunyuanImage 2.1 / 3.017BQ8_021.6 GB
Ling-Coder-Lite16.8BQ8_021.4 GB
DeepSeek-Coder-V2 16B / 236B16BQ8_020.4 GB
Kimi-VL A3B16BQ8_020.4 GB
Apriel-1.5-15B-Thinker15BQ8_019.1 GB
StarCoder2 3B / 7B / 15B15BQ8_019.1 GB
Qwen2.5 14B14.7BQ8_019.5 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at Q4_K_MSystem RAM at 4k
OTel 2.0 LLM 31B IT32.1B27.5 GB needed29.5 GB
DeepSeek-Coder 1.3B / 6.7B / 33B33B24.2 GB needed26.2 GB
WizardCoder 33B33B24.2 GB needed26.2 GB
Yi 1.5 9B / 34B34B24.9 GB needed26.9 GB
Granite Code 3B to 34B34B24.9 GB needed26.9 GB
LLaVA 1.5 / 1.6 (7B to 34B)34B24.9 GB needed26.9 GB
Ovis 234B24.9 GB needed26.9 GB
Qwen3.6-35B-A3B35B25.6 GB needed27.6 GB
Command R (35B)35B25.6 GB needed27.6 GB
Seed-OSS 36B36B26.4 GB needed28.4 GB

How to read this

The NVIDIA L4 graphics card features 24 GB of GDDR6 memory. This dedicated memory size determines which artificial intelligence models can run entirely on the hardware. When a model fits completely within this 24 GB limit, it runs at maximum speed because the processor has direct and fast access to all parameters.

The quantization column shows the best compression format for each model. Quantization reduces the size of model weights to save memory. For example, the Q4_K_M format allows 32B models like Qwen3, Qwen3.5 dense variants, Aya Expanse, Granite 4.0, and Qwen2.5-Coder to fit inside 23.4 GB of memory. The Q5_K_M format fits the 27B Gemma 3, Gemma 3 vision, and Wan 2.2 or 2.5 models inside 23 GB. The Q6_K format fits the 24B Mistral Small 3.2, Magistral Small, and Devstral Small 1.1 inside 23.6 GB. The Q8_0 format fits the 17B HunyuanImage 2.1 or 3.0 inside 21.6 GB.

Models that require more than 24 GB of memory cannot fit entirely on the NVIDIA L4. These models require CPU offloading to run on a system with 32 GB of system RAM. Offloading means some model parts are stored in system memory instead of graphics memory. This process allows you to run larger models, but it costs significant processing speed because system RAM is much slower than GDDR6 memory.

Examples of offloaded models include the OTel 2.0 LLM 31B IT which needs 27.5 GB of memory at Q4_K_M and uses 29.5 GB of system RAM. The DeepSeek-Coder 33B and WizardCoder 33B models need 24.2 GB at Q4_K_M and use 26.2 GB of system RAM. The Yi 1.5 34B, Granite Code 34B, LLaVA 1.5 or 1.6 34B, and Ovis 2 models need 24.9 GB at Q4_K_M and use 26.9 GB of system RAM. The Qwen3.6-35B-A3B and Command R 35B models need 25.6 GB at Q4_K_M and use 27.6 GB of system RAM. The Seed-OSS 36B needs 26.4 GB at Q4_K_M and uses 28.4 GB of system RAM.

All listed memory figures are calculated using a standard context window of 4k tokens. If you increase the context window to process longer documents, the model will require more memory. This extra memory usage can cause a model that normally fits on the card to overflow into system RAM, which will decrease performance.