Best local AI models for AMD RX 6700

10 GB GDDR6. At a 4k context, 136 of the 233 models in our catalog with verified parameter counts fit fully, up to Vicuna 13B at 13B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 136 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Vicuna 13B13BQ4_K_M9.5 GB
HunyuanVideo13BQ4_K_M9.5 GB
HunyuanVideo-Avatar13BQ4_K_M9.5 GB
LTX-Video / LTX-213BQ4_K_M9.5 GB
FramePack13BQ4_K_M9.5 GB
Gemma 3 12B12BQ4_K_M8.8 GB
Gemma 4 12B12BQ4_K_M8.8 GB
Mistral NeMo 12B12BQ4_K_M8.8 GB
Pixtral 12B12BQ4_K_M8.8 GB
FLUX.1 schnell12BQ4_K_M8.8 GB
FLUX.1 Kontext dev12BQ4_K_M8.8 GB
FLUX.1 Krea dev12BQ4_K_M8.8 GB
Open-Sora 2.011BQ5_K_M9.4 GB
Mochi 110BQ6_K9.8 GB
Gemma 2 9B9BQ5_K_M9.1 GB
Nemotron Nano 4B / 9B9BQ6_K8.9 GB
GLM-4 9B / GLM-4.5-Air9BQ6_K8.9 GB
Yi-Coder 1.5B / 9B9BQ6_K8.9 GB
GLM-4-9B-Chat / CodeGeeX49BQ6_K8.9 GB
GLM-4V-9B / GLM-4.1V-Thinking9BQ6_K8.9 GB
Chroma8.9BQ6_K8.8 GB
Llama 3.1 8B8BQ6_K8.4 GB
Granite 3.3 2B / 8B8BQ6_K7.9 GB
Ministral 3B / 8B8BQ6_K7.9 GB
InternLM 3 8B8BQ6_K7.9 GB
OpenCoder 1.5B / 8B8BQ6_K7.9 GB
Seed-Coder 8B8BQ6_K7.9 GB
MiniCPM-V 2.6 / MiniCPM-o 2.68BQ6_K7.9 GB
Idefics 3 8B8BQ6_K7.9 GB
Fuyu-8B8BQ6_K7.9 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
FLUX.1 dev12B14.4 GB needed16.4 GB
Phi-3 Medium14B10.2 GB needed12.2 GB
Phi-414B10.2 GB needed12.2 GB
Phi-4-reasoning / -plus14B10.2 GB needed12.2 GB
Wan 2.2 T2I14B10.2 GB needed12.2 GB
Wan 2.1 (1.3B / 14B)14B10.2 GB needed12.2 GB
SkyReels V214B10.2 GB needed12.2 GB
Qwen2.5 14B14.7B11.6 GB needed13.6 GB
Apriel-1.5-15B-Thinker15B11 GB needed13 GB
StarCoder2 3B / 7B / 15B15B11 GB needed13 GB

How to read this

The AMD Radeon RX 6700 graphics card features 10 GB of GDDR6 video memory. This onboard memory capacity determines which artificial intelligence models can run directly on your hardware. For local execution, the size of the model and its memory footprint must align with this limit to ensure fast processing speeds.

To fit larger models into the 10 GB memory space, quantization is used to compress the files. The quantization column shows the best format for each model. For example, Vicuna 13B, HunyuanVideo, HunyuanVideo-Avatar, LTX-Video / LTX-2, and FramePack 13B can run using the Q4_K_M quantization, which uses 9.5 GB of video memory. Similarly, Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, FLUX.1 Kontext dev, and FLUX.1 Krea dev fit within 8.8 GB of video memory using the same Q4_K_M quantization.

Other models use different quantization levels to balance quality and memory usage. Open-Sora 2.0 11B uses 9.4 GB with Q5_K_M. Mochi 1 10B uses 9.8 GB with Q6_K. Gemma 2 9B uses 9.1 GB with Q5_K_M. Nemotron Nano 4B / 9B, GLM-4 9B / GLM-4.5-Air, Yi-Coder 1.5B / 9B, GLM-4-9B-Chat / CodeGeeX4, and GLM-4V-9B / GLM-4.1V-Thinking all run at Q6_K using 8.9 GB of memory. Chroma 8.9B fits in 8.8 GB using Q6_K.

Smaller models leave more room for processing. Llama 3.1 8B uses 8.4 GB of memory with Q6_K. Granite 3.3 2B / 8B, Ministral 3B / 8B, InternLM 3 8B, OpenCoder 1.5B / 8B, Seed-Coder 8B, MiniCPM-V 2.6 / MiniCPM-o 2.6, Idefics 3 8B, and Fuyu-8B all require 7.9 GB of video memory using the Q6_K quantization format.

When a model exceeds the 10 GB limit, you must offload data to your system RAM. This process requires a 32 GB system RAM setup. FLUX.1 dev 12B needs 14.4 GB at FP8 / optimized and uses 16.4 GB of system RAM. Phi-3 Medium 14B, Phi-4 14B, Phi-4-reasoning / -plus 14B, Wan 2.2 T2I 14B, Wan 2.1 (1.3B / 14B) 14B, and SkyReels V2 14B need 10.2 GB at Q4_K_M and use 12.2 GB of system RAM. Qwen2.5 14B needs 11.6 GB at Q4_K_M and uses 13.6 GB of system RAM. Apriel-1.5-15B-Thinker 15B and StarCoder2 3B / 7B / 15B 15B need 11 GB at Q4_K_M and use 13 GB of system RAM. Offloading allows you to run these larger models but reduces generation speed because system RAM is slower than video memory.

All memory calculations are based on a standard 4k context window. If you increase the context window to process longer prompts, the memory usage will rise. This extra memory demand might force you to use a lower quantization level or offload more data to system RAM to prevent out of memory errors.