Best local AI models for AMD RX 7700S

8 GB GDDR6. At a 4k context, 123 of the 233 models in our catalog with verified parameter counts fit fully, up to Mochi 1 at 10B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 123 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Mochi 110BQ4_K_M7.3 GB
Gemma 2 9B9BQ4_K_M8 GB
Nemotron Nano 4B / 9B9BQ5_K_M7.7 GB
GLM-4 9B / GLM-4.5-Air9BQ5_K_M7.7 GB
Yi-Coder 1.5B / 9B9BQ5_K_M7.7 GB
GLM-4-9B-Chat / CodeGeeX49BQ5_K_M7.7 GB
GLM-4V-9B / GLM-4.1V-Thinking9BQ5_K_M7.7 GB
Chroma8.9BQ5_K_M7.6 GB
Llama 3.1 8B8BQ5_K_M7.4 GB
Granite 3.3 2B / 8B8BQ6_K7.9 GB
Ministral 3B / 8B8BQ6_K7.9 GB
InternLM 3 8B8BQ6_K7.9 GB
OpenCoder 1.5B / 8B8BQ6_K7.9 GB
Seed-Coder 8B8BQ6_K7.9 GB
MiniCPM-V 2.6 / MiniCPM-o 2.68BQ6_K7.9 GB
Idefics 3 8B8BQ6_K7.9 GB
Fuyu-8B8BQ6_K7.9 GB
Emu38BQ6_K7.9 GB
Stable Diffusion 3.5 Large / Turbo8BQ6_K7.9 GB
EXAONE 3.5 2.4B / 7.8B7.8BQ6_K7.7 GB
Mistral 7B7BQ6_K7.4 GB
Qwen2.5 0.5B / 1.5B / 3B / 7B7BQ6_K6.9 GB
OLMo 2 1B / 7B7BQ6_K6.9 GB
Falcon 3 1B / 3B / 7B7BQ6_K6.9 GB
Command R7B7BQ6_K6.9 GB
OpenHermes 2.57BQ6_K6.9 GB
Zephyr 7B Beta7BQ6_K6.9 GB
OpenChat 3.57BQ6_K6.9 GB
Starling LM 7B7BQ6_K6.9 GB
Codestral Mamba 7B7BQ6_K6.9 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at Q4_K_MSystem RAM at 4k
Open-Sora 2.011B8.1 GB needed10.1 GB
FLUX.1 dev12B14.4 GB needed16.4 GB
Gemma 3 12B12B8.8 GB needed10.8 GB
Gemma 4 12B12B8.8 GB needed10.8 GB
Mistral NeMo 12B12B8.8 GB needed10.8 GB
Pixtral 12B12B8.8 GB needed10.8 GB
FLUX.1 schnell12B8.8 GB needed10.8 GB
FLUX.1 Kontext dev12B8.8 GB needed10.8 GB
FLUX.1 Krea dev12B8.8 GB needed10.8 GB
Vicuna 13B13B9.5 GB needed11.5 GB

How to read this

The AMD Radeon RX 7700S is a mobile graphics processor equipped with 8 GB of GDDR6 dedicated video memory. This hardware limit dictates the maximum size of the artificial intelligence models you can run entirely on the graphics card. To run a model smoothly without performance drops, the model weights and the active working memory must fit within this 8 GB boundary.

The quantization column indicates the compression level applied to each model. Quantization reduces the precision of the model weights to save space. For example, the Q4_K_M and Q5_K_M quants compress weights to approximately four and five bits. Higher quants like Q6_K offer better accuracy but require more memory. On this hardware, Gemma 2 9B fits at Q4_K_M using exactly 8 GB, while Llama 3.1 8B fits at Q5_K_M using 7.4 GB.

When a model exceeds the 8 GB video memory limit, you must use CPU offloading. This process splits the model layers between your graphics card and your system RAM. We assume your system has 32 GB of system RAM for these calculations. For instance, running FLUX.1 dev at FP8 requires 14.4 GB of memory, which uses all your video memory and needs an additional 16.4 GB of system RAM.

CPU offloading allows you to run larger models like Vicuna 13B or Mistral NeMo 12B, but it comes with a speed penalty. Transferring data between the system RAM and the graphics card over the system bus is much slower than using dedicated GDDR6 memory. Models like Gemma 3 12B and Gemma 4 12B require 8.8 GB of memory at Q4_K_M, which forces 10.8 GB of data into your system RAM and slows down generation speeds.

You must also consider the memory required for the context window. The memory figures listed here are calculated using a standard 4k context window. If you increase the context length to process longer documents or conversations, the system will require more memory. This extra demand can push a model that normally fits, such as Mochi 1 10B at 7.3 GB, over the 8 GB limit.