Best local AI models for AMD RX 6700M

10 GB GDDR6. At a 4k context, 136 of the 233 models in our catalog with verified parameter counts fit fully, up to Vicuna 13B at 13B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 136 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Vicuna 13B13BQ4_K_M9.5 GB
HunyuanVideo13BQ4_K_M9.5 GB
HunyuanVideo-Avatar13BQ4_K_M9.5 GB
LTX-Video / LTX-213BQ4_K_M9.5 GB
FramePack13BQ4_K_M9.5 GB
Gemma 3 12B12BQ4_K_M8.8 GB
Gemma 4 12B12BQ4_K_M8.8 GB
Mistral NeMo 12B12BQ4_K_M8.8 GB
Pixtral 12B12BQ4_K_M8.8 GB
FLUX.1 schnell12BQ4_K_M8.8 GB
FLUX.1 Kontext dev12BQ4_K_M8.8 GB
FLUX.1 Krea dev12BQ4_K_M8.8 GB
Open-Sora 2.011BQ5_K_M9.4 GB
Mochi 110BQ6_K9.8 GB
Gemma 2 9B9BQ5_K_M9.1 GB
Nemotron Nano 4B / 9B9BQ6_K8.9 GB
GLM-4 9B / GLM-4.5-Air9BQ6_K8.9 GB
Yi-Coder 1.5B / 9B9BQ6_K8.9 GB
GLM-4-9B-Chat / CodeGeeX49BQ6_K8.9 GB
GLM-4V-9B / GLM-4.1V-Thinking9BQ6_K8.9 GB
Chroma8.9BQ6_K8.8 GB
Llama 3.1 8B8BQ6_K8.4 GB
Granite 3.3 2B / 8B8BQ6_K7.9 GB
Ministral 3B / 8B8BQ6_K7.9 GB
InternLM 3 8B8BQ6_K7.9 GB
OpenCoder 1.5B / 8B8BQ6_K7.9 GB
Seed-Coder 8B8BQ6_K7.9 GB
MiniCPM-V 2.6 / MiniCPM-o 2.68BQ6_K7.9 GB
Idefics 3 8B8BQ6_K7.9 GB
Fuyu-8B8BQ6_K7.9 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
FLUX.1 dev12B14.4 GB needed16.4 GB
Phi-3 Medium14B10.2 GB needed12.2 GB
Phi-414B10.2 GB needed12.2 GB
Phi-4-reasoning / -plus14B10.2 GB needed12.2 GB
Wan 2.2 T2I14B10.2 GB needed12.2 GB
Wan 2.1 (1.3B / 14B)14B10.2 GB needed12.2 GB
SkyReels V214B10.2 GB needed12.2 GB
Qwen2.5 14B14.7B11.6 GB needed13.6 GB
Apriel-1.5-15B-Thinker15B11 GB needed13 GB
StarCoder2 3B / 7B / 15B15B11 GB needed13 GB

How to read this

The AMD Radeon RX 6700M is a mobile graphics card equipped with 10 GB of GDDR6 dedicated video memory. When running artificial intelligence models locally, this hardware memory limit determines which model sizes can run entirely on your graphics processor. Keeping the model files inside the high speed video memory is critical for achieving fast generation speeds.

The best quant column shows the optimal quantization level for each model. Quantization is a compression technique that reduces the precision of model weights to save space. For the 10 GB memory capacity of this card, the Q4_K_M quantization is the best choice for 13B and 12B models. This includes Vicuna 13B, HunyuanVideo, HunyuanVideo-Avatar, LTX-Video / LTX-2, and FramePack, which all use 9.5 GB of video memory. It also applies to Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, FLUX.1 Kontext dev, and FLUX.1 Krea dev, which use 8.8 GB of video memory.

Models with slightly smaller parameter counts can run at higher precision levels. The Open-Sora 2.0 model at 11B runs at Q5_K_M using 9.4 GB. The Mochi 1 model at 10B runs at Q6_K using 9.8 GB. Several 9B models like Gemma 2 9B, Nemotron Nano 4B / 9B, GLM-4 9B / GLM-4.5-Air, Yi-Coder 1.5B / 9B, GLM-4-9B-Chat / CodeGeeX4, and GLM-4V-9B / GLM-4.1V-Thinking fit well at Q6_K or Q5_K_M. The Chroma model at 8.9B uses 8.8 GB at Q6_K.

Many popular 8B models fit comfortably in the 10 GB limit at Q6_K precision. This group includes Llama 3.1 8B using 8.4 GB. It also includes Granite 3.3 2B / 8B, Ministral 3B / 8B, InternLM 3 8B, OpenCoder 1.5B / 8B, Seed-Coder 8B, MiniCPM-V 2.6 / MiniCPM-o 2.6, Idefics 3 8B, and Fuyu-8B, which all use 7.9 GB of video memory. Running these models at Q6_K leaves a small amount of headroom for context data.

When a model exceeds the 10 GB video memory limit, you must use CPU offloading. This process splits the model between your graphics card and your system RAM. Offloading allows you to run larger models but it reduces your generation speed. For these cases, we assume your computer has 32 GB of system RAM. This setup allows you to run FLUX.1 dev at FP8 / optimized, which needs 14.4 GB of video memory and 16.4 GB of system RAM.

Other offload options include 14B models like Phi-3 Medium, Phi-4, Phi-4-reasoning / -plus, Wan 2.2 T2I, Wan 2.1 (1.3B / 14B), and SkyReels V2. These models require 10.2 GB at Q4_K_M and 12.2 GB of system RAM. The Qwen2.5 14B model requires 11.6 GB at Q4_K_M and 13.6 GB of system RAM. The Apriel-1.5-15B-Thinker and StarCoder2 3B / 7B / 15B models require 11 GB at Q4_K_M and 13 GB of system RAM.

You must monitor your context window usage when running models close to your memory limit. The memory figures listed here are calculated using a standard 4k context window. If you increase the context length to process longer documents or chat histories, the memory usage will grow. This extra memory demand can exceed the 10 GB limit and trigger slow system RAM offloading.