Best local AI models for AMD RX 6750 GRE 10GB

10 GB GDDR6. At a 4k context, 136 of the 233 models in our catalog with verified parameter counts fit fully, up to Vicuna 13B at 13B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 136 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Vicuna 13B13BQ4_K_M9.5 GB
HunyuanVideo13BQ4_K_M9.5 GB
HunyuanVideo-Avatar13BQ4_K_M9.5 GB
LTX-Video / LTX-213BQ4_K_M9.5 GB
FramePack13BQ4_K_M9.5 GB
Gemma 3 12B12BQ4_K_M8.8 GB
Gemma 4 12B12BQ4_K_M8.8 GB
Mistral NeMo 12B12BQ4_K_M8.8 GB
Pixtral 12B12BQ4_K_M8.8 GB
FLUX.1 schnell12BQ4_K_M8.8 GB
FLUX.1 Kontext dev12BQ4_K_M8.8 GB
FLUX.1 Krea dev12BQ4_K_M8.8 GB
Open-Sora 2.011BQ5_K_M9.4 GB
Mochi 110BQ6_K9.8 GB
Gemma 2 9B9BQ5_K_M9.1 GB
Nemotron Nano 4B / 9B9BQ6_K8.9 GB
GLM-4 9B / GLM-4.5-Air9BQ6_K8.9 GB
Yi-Coder 1.5B / 9B9BQ6_K8.9 GB
GLM-4-9B-Chat / CodeGeeX49BQ6_K8.9 GB
GLM-4V-9B / GLM-4.1V-Thinking9BQ6_K8.9 GB
Chroma8.9BQ6_K8.8 GB
Llama 3.1 8B8BQ6_K8.4 GB
Granite 3.3 2B / 8B8BQ6_K7.9 GB
Ministral 3B / 8B8BQ6_K7.9 GB
InternLM 3 8B8BQ6_K7.9 GB
OpenCoder 1.5B / 8B8BQ6_K7.9 GB
Seed-Coder 8B8BQ6_K7.9 GB
MiniCPM-V 2.6 / MiniCPM-o 2.68BQ6_K7.9 GB
Idefics 3 8B8BQ6_K7.9 GB
Fuyu-8B8BQ6_K7.9 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
FLUX.1 dev12B14.4 GB needed16.4 GB
Phi-3 Medium14B10.2 GB needed12.2 GB
Phi-414B10.2 GB needed12.2 GB
Phi-4-reasoning / -plus14B10.2 GB needed12.2 GB
Wan 2.2 T2I14B10.2 GB needed12.2 GB
Wan 2.1 (1.3B / 14B)14B10.2 GB needed12.2 GB
SkyReels V214B10.2 GB needed12.2 GB
Qwen2.5 14B14.7B11.6 GB needed13.6 GB
Apriel-1.5-15B-Thinker15B11 GB needed13 GB
StarCoder2 3B / 7B / 15B15B11 GB needed13 GB

How to read this

The AMD Radeon RX 6750 GRE graphics card features 10 GB of GDDR6 memory. This onboard memory size dictates which artificial intelligence models you can run entirely on the hardware. To run a model smoothly without slowdowns, the model files and active context must fit inside this 10 GB limit.

Quantization is a method that compresses model files to save space. The quant column shows the best balance of size and quality for each model. For example, the Q4_K_M quant represents a four bit quantization level that reduces file size while keeping high accuracy. Higher quants like Q5_K_M or Q6_K provide even better quality but require more memory.

With 10 GB of video memory, you can run several large models locally. Vicuna 13B, HunyuanVideo, HunyuanVideo-Avatar, LTX-Video / LTX-2, and FramePack fit at the Q4_K_M quant using 9.5 GB of memory. Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, FLUX.1 Kontext dev, and FLUX.1 Krea dev fit at Q4_K_M using 8.8 GB of memory. Open-Sora 2.0 fits at Q5_K_M using 9.4 GB of memory.

Slightly smaller models can run at higher quants for better precision. Mochi 1 fits at Q6_K using 9.8 GB of memory. Gemma 2 9B fits at Q5_K_M using 9.1 GB of memory. Nemotron Nano 4B / 9B, GLM-4 9B / GLM-4.5-Air, Yi-Coder 1.5B / 9B, GLM-4-9B-Chat / CodeGeeX4, GLM-4V-9B / GLM-4.1V-Thinking, and Chroma fit at Q6_K using 8.8 to 8.9 GB of memory. Llama 3.1 8B fits at Q6_K using 8.4 GB of memory. Granite 3.3 2B / 8B, Ministral 3B / 8B, InternLM 3 8B, OpenCoder 1.5B / 8B, Seed-Coder 8B, MiniCPM-V 2.6 / MiniCPM-o 2.6, Idefics 3 8B, and Fuyu-8B fit at Q6_K using 7.9 GB of memory.

When a model is too large for the 10 GB video memory, you can offload parts of it to your system RAM. This offload process allows you to run larger models but slows down generation speeds. If you have 32 GB of system RAM, you can run FLUX.1 dev at FP8 or optimized settings which needs 14.4 GB of space and uses 16.4 GB of system RAM. Phi-3 Medium, Phi-4, Phi-4-reasoning / -plus, Wan 2.2 T2I, Wan 2.1 (1.3B / 14B), and SkyReels V2 fit at Q4_K_M using 10.2 GB of space and 12.2 GB of system RAM. Qwen2.5 14B needs 11.6 GB at Q4_K_M and uses 13.6 GB of system RAM. Apriel-1.5-15B-Thinker and StarCoder2 3B / 7B / 15B need 11 GB at Q4_K_M and use 13 GB of system RAM.

Memory calculations assume a standard context length of 4k tokens. If you increase the context window to process longer documents, the model will require significantly more memory. This extra memory usage might force you to use a lower quant or rely on slower system RAM offloading.