Best local AI models for AMD PRO W7900

48 GB GDDR6. At a 4k context, 186 of the 233 models in our catalog with verified parameter counts fit fully, up to Jamba 1.5 Mini / Large at 52B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 186 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Jamba 1.5 Mini / Large52BQ5_K_M44.3 GB
Llama 3.1 Nemotron 51B51BQ5_K_M43.5 GB
Mixtral 8x7B47BQ6_K46.2 GB
Seed-OSS 36B36BQ8_045.8 GB
Qwen3.6-35B-A3B35BQ8_044.5 GB
Command R (35B)35BQ8_044.5 GB
Yi 1.5 9B / 34B34BQ8_043.2 GB
Granite Code 3B to 34B34BQ8_043.2 GB
LLaVA 1.5 / 1.6 (7B to 34B)34BQ8_043.2 GB
Ovis 234BQ8_043.2 GB
DeepSeek-Coder 1.3B / 6.7B / 33B33BQ8_042 GB
WizardCoder 33B33BQ8_042 GB
OTel 2.0 LLM 31B IT32.1BQ8_044.9 GB
Qwen3 8B / 14B / 32B32BQ8_040.7 GB
Qwen3.5 (dense variants)32BQ8_040.7 GB
Aya Expanse 8B / 32B32BQ8_040.7 GB
Granite 4.0 Small/Tiny32BQ8_040.7 GB
Qwen2.5-Coder 0.5B to 32B32BQ8_040.7 GB
Qwen3-30B-A3B30BQ8_038.2 GB
Qwen3-Coder 30B-A3B30BQ8_038.2 GB
Gemma 3 27B27BQ8_034.3 GB
Gemma 3 4B/12B/27B (vision)27BQ8_034.3 GB
Wan 2.2 / 2.527BQ8_034.3 GB
Gemma 4 26B-A4B26BQ8_033.1 GB
Gemma 4 (all sizes)26BQ8_033.1 GB
Aria25BQ8_031.8 GB
Mistral Small 3.224BQ8_030.5 GB
Magistral Small24BQ8_030.5 GB
Devstral Small 1.124BQ8_030.5 GB
Solar Pro22BQ8_028 GB

How to read this

The AMD Radeon PRO W7900 workstation graphics card features 48 GB of GDDR6 onboard memory. This memory capacity determines the maximum size of the artificial intelligence models you can run locally. To run a model entirely on the graphics hardware, the model files and the active processing data must fit within this 48 GB limit. Keeping the entire model in the graphics memory ensures the fastest possible processing speeds.

Models are often compressed using a method called quantization to save space. The best quant column shows the highest quality compression level that still fits comfortably inside the 48 GB memory space. For example, a Q8_0 quantization represents an eight bit format that preserves high accuracy. Larger models like the Jamba 1.5 Mini or Large 52B model require a slightly higher Q5_K_M compression to fit within 44.3 GB of memory.

When a model exceeds the available graphics memory, some data must be offloaded to the system memory. For this hardware configuration with 32 GB of system RAM, there are no recommended CPU offload cases. Running models with CPU offload dramatically slows down processing speeds. Keeping the models fully loaded on the 48 GB GDDR6 memory of the AMD PRO W7900 avoids these performance penalties entirely.

The memory usage figures listed for each model assume a standard 4k context window. The context window is the total amount of text the model can read and write at one time. If you increase the context window beyond 4000 tokens, the model will require significantly more memory. You must leave some free space in the 48 GB memory to accommodate this extra context data during active use.

Many powerful models fit completely within the 48 GB limit of this card. The Mixtral 8x7B model fits at a high quality Q6_K quantization using 46.2 GB of memory. Other models like the Llama 3.1 Nemotron 51B fit at Q5_K_M using 43.5 GB. You can also run the Seed-OSS 36B model at Q8_0 quantization using 45.8 GB, or the Qwen3.6-35B-A3B and Command R 35B models at Q8_0 using 44.5 GB.

Smaller models leave even more room for extended context windows. The Gemma 3 27B model fits at Q8_0 quantization using 34.3 GB of memory. The Mistral Small 3.2 24B model fits at Q8_0 using 30.5 GB of memory. Choosing these slightly smaller models allows you to run complex tasks without risking memory overflow on your AMD PRO W7900 hardware.