Best local AI models for AMD RX 7900 XTX

24 GB GDDR6. At a 4k context, 173 of the 233 models in our catalog with verified parameter counts fit fully, up to Qwen3 8B / 14B / 32B at 32B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 173 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Qwen3 8B / 14B / 32B32BQ4_K_M23.4 GB
Qwen3.5 (dense variants)32BQ4_K_M23.4 GB
Aya Expanse 8B / 32B32BQ4_K_M23.4 GB
Granite 4.0 Small/Tiny32BQ4_K_M23.4 GB
Qwen2.5-Coder 0.5B to 32B32BQ4_K_M23.4 GB
Qwen3-30B-A3B30BQ4_K_M22 GB
Qwen3-Coder 30B-A3B30BQ4_K_M22 GB
Gemma 3 27B27BQ5_K_M23 GB
Gemma 3 4B/12B/27B (vision)27BQ5_K_M23 GB
Wan 2.2 / 2.527BQ5_K_M23 GB
Gemma 4 26B-A4B26BQ5_K_M22.2 GB
Gemma 4 (all sizes)26BQ5_K_M22.2 GB
Aria25BQ5_K_M21.3 GB
Mistral Small 3.224BQ6_K23.6 GB
Magistral Small24BQ6_K23.6 GB
Devstral Small 1.124BQ6_K23.6 GB
Solar Pro22BQ6_K21.6 GB
Codestral 22B22BQ6_K21.6 GB
gpt-oss-20b21BQ6_K20.7 GB
Reka Flash 321BQ6_K20.7 GB
Qwen-Image20BQ6_K19.7 GB
Qwen-Image-Edit20BQ6_K19.7 GB
CogVLM219BQ6_K18.7 GB
HunyuanImage 2.1 / 3.017BQ8_021.6 GB
Ling-Coder-Lite16.8BQ8_021.4 GB
DeepSeek-Coder-V2 16B / 236B16BQ8_020.4 GB
Kimi-VL A3B16BQ8_020.4 GB
Apriel-1.5-15B-Thinker15BQ8_019.1 GB
StarCoder2 3B / 7B / 15B15BQ8_019.1 GB
Qwen2.5 14B14.7BQ8_019.5 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at Q4_K_MSystem RAM at 4k
OTel 2.0 LLM 31B IT32.1B27.5 GB needed29.5 GB
DeepSeek-Coder 1.3B / 6.7B / 33B33B24.2 GB needed26.2 GB
WizardCoder 33B33B24.2 GB needed26.2 GB
Yi 1.5 9B / 34B34B24.9 GB needed26.9 GB
Granite Code 3B to 34B34B24.9 GB needed26.9 GB
LLaVA 1.5 / 1.6 (7B to 34B)34B24.9 GB needed26.9 GB
Ovis 234B24.9 GB needed26.9 GB
Qwen3.6-35B-A3B35B25.6 GB needed27.6 GB
Command R (35B)35B25.6 GB needed27.6 GB
Seed-OSS 36B36B26.4 GB needed28.4 GB

How to read this

The AMD RX 7900 XTX graphics card features 24 GB of GDDR6 memory. This onboard memory size determines which artificial intelligence models can run entirely on your local hardware. When a model fits completely within this 24 GB limit, the system processes tokens at maximum speed. If a model exceeds this capacity, the system must transfer data to your system memory, which slows down performance.

The quantization column indicates the compression level applied to each model. Quantization reduces the size of a model so it requires less memory. For example, the Q4_K_M quantization represents a four bit format that balances size and accuracy. The Q5_K_M and Q6_K formats offer higher precision but require more memory. The Q8_0 format provides the highest precision in this selection but uses the most memory per parameter.

For maximum local performance, several 32B models fit within the 24 GB memory limit. The Qwen3 32B, Qwen3.5 dense variants, Aya Expanse 32B, Granite 4.0 Small or Tiny, and Qwen2.5-Coder 32B models all run locally using the Q4_K_M quantization, which consumes 23.4 GB of memory. Other options include the Gemma 3 27B and Wan 2.2 or 2.5 models at Q5_K_M quantization, which require 23 GB of memory.

You can run larger models by offloading part of the workload to your system RAM, assuming your computer has 32 GB of system memory. This offloading process allows you to run the Command R 35B model, which requires 25.6 GB of memory at Q4_K_M quantization and 27.6 GB of system RAM. Similarly, the Yi 1.5 34B model requires 24.9 GB of memory at Q4_K_M quantization and 26.9 GB of system RAM. Offloading enables these larger models to run, but it reduces processing speeds.

When selecting a model, you must account for the context window size. The listed memory usage figures assume a standard 4k context window. If you increase the context window to process longer documents or conversations, the system requires additional memory for the context cache. This extra memory requirement may force you to choose a smaller model size or a lower quantization level to prevent system slowdowns.