Best local AI models for Intel Arc B570

10 GB GDDR6. At a 4k context, 136 of the 233 models in our catalog with verified parameter counts fit fully, up to Vicuna 13B at 13B parameters.

Check your own machine against every model →

The largest models that fit fully

The 30 largest of the 136 models that fit; every smaller model in the catalog fits too. Best quant means the highest quality compression whose weights and 4k context both sit inside the memory.

ModelParametersBest quant that fitsMemory used at 4k
Vicuna 13B13BQ4_K_M9.5 GB
HunyuanVideo13BQ4_K_M9.5 GB
HunyuanVideo-Avatar13BQ4_K_M9.5 GB
LTX-Video / LTX-213BQ4_K_M9.5 GB
FramePack13BQ4_K_M9.5 GB
Gemma 3 12B12BQ4_K_M8.8 GB
Gemma 4 12B12BQ4_K_M8.8 GB
Mistral NeMo 12B12BQ4_K_M8.8 GB
Pixtral 12B12BQ4_K_M8.8 GB
FLUX.1 schnell12BQ4_K_M8.8 GB
FLUX.1 Kontext dev12BQ4_K_M8.8 GB
FLUX.1 Krea dev12BQ4_K_M8.8 GB
Open-Sora 2.011BQ5_K_M9.4 GB
Mochi 110BQ6_K9.8 GB
Gemma 2 9B9BQ5_K_M9.1 GB
Nemotron Nano 4B / 9B9BQ6_K8.9 GB
GLM-4 9B / GLM-4.5-Air9BQ6_K8.9 GB
Yi-Coder 1.5B / 9B9BQ6_K8.9 GB
GLM-4-9B-Chat / CodeGeeX49BQ6_K8.9 GB
GLM-4V-9B / GLM-4.1V-Thinking9BQ6_K8.9 GB
Chroma8.9BQ6_K8.8 GB
Llama 3.1 8B8BQ6_K8.4 GB
Granite 3.3 2B / 8B8BQ6_K7.9 GB
Ministral 3B / 8B8BQ6_K7.9 GB
InternLM 3 8B8BQ6_K7.9 GB
OpenCoder 1.5B / 8B8BQ6_K7.9 GB
Seed-Coder 8B8BQ6_K7.9 GB
MiniCPM-V 2.6 / MiniCPM-o 2.68BQ6_K7.9 GB
Idefics 3 8B8BQ6_K7.9 GB
Fuyu-8B8BQ6_K7.9 GB

Close, but only with CPU offload

These need more than the card holds at their smallest practical quant, so part of the model runs from system memory (figures assume 32 GB of it). They work, several times slower.

ModelParametersMemory at FP8 / optimizedSystem RAM at 4k
FLUX.1 dev12B14.4 GB needed16.4 GB
Phi-3 Medium14B10.2 GB needed12.2 GB
Phi-414B10.2 GB needed12.2 GB
Phi-4-reasoning / -plus14B10.2 GB needed12.2 GB
Wan 2.2 T2I14B10.2 GB needed12.2 GB
Wan 2.1 (1.3B / 14B)14B10.2 GB needed12.2 GB
SkyReels V214B10.2 GB needed12.2 GB
Qwen2.5 14B14.7B11.6 GB needed13.6 GB
Apriel-1.5-15B-Thinker15B11 GB needed13 GB
StarCoder2 3B / 7B / 15B15B11 GB needed13 GB

How to read this

The Intel Arc B570 graphics card features 10 GB GDDR6 memory. This dedicated memory size determines which artificial intelligence models can run entirely on your hardware. To run a model smoothly the model files and the active context data must fit within this 10 GB limit. If a model exceeds this capacity it will fail to load or run very slowly.

The best quant column shows the optimal quantization level for each model. Quantization reduces the precision of model weights to save memory. For this hardware the Q4_K_M quant represents a four bit quantization that balances file size and output quality. The Q5_K_M and Q6_K quants offer higher precision at five bits and six bits which require more memory but deliver better accuracy.

Several large models fit completely within the 10 GB GDDR6 memory of the Intel Arc B570. The 13B models like Vicuna 13B, HunyuanVideo, HunyuanVideo-Avatar, LTX-Video / LTX-2, and FramePack use 9.5 GB of memory at the Q4_K_M quant. The 12B models including Gemma 3 12B, Gemma 4 12B, Mistral NeMo 12B, Pixtral 12B, FLUX.1 schnell, FLUX.1 Kontext dev, and FLUX.1 Krea dev use 8.8 GB of memory at the Q4_K_M quant.

Other models fit well within the local memory limits. Open-Sora 2.0 is an 11B model using 9.4 GB at the Q5_K_M quant. Mochi 1 is a 10B model using 9.8 GB at the Q6_K quant. The 9B models like Gemma 2 9B, Nemotron Nano 4B / 9B, GLM-4 9B / GLM-4.5-Air, Yi-Coder 1.5B / 9B, GLM-4-9B-Chat / CodeGeeX4, and GLM-4V-9B / GLM-4.1V-Thinking use between 8.9 GB and 9.1 GB of memory. Chroma is an 8.9B model using 8.8 GB of memory.

Smaller 8B models run comfortably on this hardware. Llama 3.1 8B uses 8.4 GB at the Q6_K quant. Granite 3.3 2B / 8B, Ministral 3B / 8B, InternLM 3 8B, OpenCoder 1.5B / 8B, Seed-Coder 8B, MiniCPM-V 2.6 / MiniCPM-o 2.6, Idefics 3 8B, and Fuyu-8B all use 7.9 GB of memory at the Q6_K quant. These models leave a small buffer of memory for system tasks.

When a model is too large for the 10 GB graphics memory you can offload parts of it to your system RAM. This offload process requires a system with at least 32 GB system RAM. FLUX.1 dev is a 12B model that needs 14.4 GB at FP8 or optimized settings which requires 16.4 GB system RAM. The 14B models like Phi-3 Medium, Phi-4, Phi-4-reasoning / -plus, Wan 2.2 T2I, Wan 2.1 (1.3B / 14B), and SkyReels V2 need 10.2 GB at the Q4_K_M quant and require 12.2 GB system RAM. Qwen2.5 14B needs 11.6 GB at Q4_K_M and requires 13.6 GB system RAM. Apriel-1.5-15B-Thinker and StarCoder2 3B / 7B / 15B need 11 GB at Q4_K_M and require 13 GB system RAM.

Be aware of the context limit when running these models. The memory numbers listed here assume a standard 4k context window. If you increase the context window to process longer documents or larger chat histories the memory usage will increase. Running close to the 10 GB limit with a large context window can cause the system to run out of memory.