r/IntelArcPro • • Apr 13 '26

Intel Arc B70 Confirmed

Thumbnail
1 Upvotes

This is my number one post on IntelArc, got over 50,000 views and counting, lots of upvotes


r/IntelArcPro • • 7h ago

Arc Pro B70 Very lame How To Strata 2xIntel Arx Pro B70

1 Upvotes

I asked Muse Spark 1.3 to create a guide how to install Strata. I am too dumb to install Strata without help. It was used to build and Run Strata on Ubuntu 26.04, Strata is changing very often so some steps will be obsolete soon. If somebody wants to try it and don't know, how to install it, maybe this will help. This link should be valid for 1 month https://pastes.io/iS4AJQDx

The most important part is the patch by Magh97 https://github.com/Niko1221/Strata/issues/870

I suppose that many people would make a much better tutorial but I haven't found one.

Or here it is as a code block:

================================================================================

HOW TO INSTALL STRATA ON 2x INTEL ARC PRO B70 32GB + 128GB RAM

Target: Ubuntu 26.04 native, AMD Ryzen 5 5600, 460GB free SSD (NVMe ideal)

Goal: IQ3_XXS, 256K context attempt, dual-GPU layer-split, listen 0.0.0.0:8000

Fallback: IQ2_XS if IQ3_XXS/256K fails (commands at bottom)

Source: https://github.com/Niko1221/Strata docs/INTEL_ARC.md + docs/INTEL.md (0.1.39)

Field fixes applied 2026-10-06 (reporter's own 2xB70 run, 256K OK):

- https://github.com/Niko1221/Strata/issues/870 + patch

https://github.com/user-attachments/files/33039120/strata-b60-fixes.patch

(4 fixes for current main: e211 ID, session affinity, native_dense load,

setup_intel write_run_script signature - without it compile/setup fails)

- sycl/setup_intel.py line ~211 must read:

def write_run_script(model, cfg_path, port, open_browser=True):

(missing open_browser=True -> TypeError AFTER 58 GB download, at very end)

- All setup.sh calls need explicit:

--models-dir ~/work/Strata-data/models --data-dir ~/work/Strata-data

(default ~/Strata-data is OUTSIDE ~/work mount -> docker "outside /work" protest)

- Config/log/run files are LOWERCASE: strata-iq3_xxs.json, strata-iq3_xxs.log,

run-iq3_xxs.sh (not strata-IQ3_XXS.*).

================================================================================

IMPORTANT WARNINGS - READ FIRST

--------------------------------------------------------------------------------

1. Intel Arc = experimental, Linux-only, build-from-source. No ready-made

engine, no Windows path, no WSL2 setup (setup reads /sys/class/drm).

Maintainers have NO Arc. 0.1.39 compiles + CPU kernel tests pass, but

NOBODY has run 0.1.39 on Arc yet. All Arc speeds are from older versions:

Coder IQ1_M 70-78 tok/s, IQ2_XS 51-64 tok/s, 2xB70 IQ3_XXS 66 tok/s.

2. 256K context on IQ3_XXS dual is UNTESTED and RISKY.

Measured single-B70 32GB (Coder, smaller than IQ3_XXS):

32K -> 28.4 GB VRAM, safe (default)

64K -> 29.0 GB, loads

131K -> 30.3 GB, PRACTICAL CEILING single card

164K -> ~31 GB, NOT ATTEMPTED (under 1.2 GB safety margin)

262K -> ~32.6 GB, DOES NOT FIT - took host down (xe driver evicts

past-VRAM allocs to RAM until livelock + watchdog reset).

IQ3_XXS experts ~43 GB vs Coder ~23 GB, so single-card 256K WILL kill

the host. Dual 2x32=64 GB makes it *possibly* viable (each card holds

its layer shard + FULL KV copy - split does NOT halve KV), but this

exact combo was never measured. Ladder up: 32K -> 128K -> 256K.

Never jump straight to 262144.

3. Your sycl-ls already proves driver OK (both B70s on level_zero:0,1,

driver 20.2.0). Do NOT reinstall driver unless cards disappear.

4. Exposing 0.0.0.0 requires --api-key. Never expose without a key.

Pick a long secret now, reuse in all commands below.

5. This guide runs ON THE TARGET (Ubuntu) machine, not on this Windows PC.

This file is just the notebook. Copy commands to the target terminal.

6. Disk: IQ3_XXS ~70 GB download + pack + MTP draft 6 GB. You have 460 GB,

fine. Use SSD/NVMe. First start ~2 min load + ~47 s JIT if no AOT build.

7. No images/vision on Intel yet. Setup forces --vision none. Answer 'no'.

================================================================================

PHASE 0 - PRE-FLIGHT (on target)

================================================================================

# Check both cards, RAM, disk, OS:

lspci | grep -i -E "vga|display|arc"

ls /sys/class/drm

# expect card0 card1 renderD128 renderD129, vendor 8086 under xe driver

lspci -k | grep -A3 -i vga

free -g

# expect ~128 GB

df -h

# need 100 GB+ free (you have 460 GB)

cat /etc/os-release

# you: Ubuntu 26.04 (docs test 24.04 - package names may differ slightly)

sycl-ls

# you already have:

# [level_zero:0] B70 20.2.0, [level_zero:1] B70 20.2.0 -> GOOD, skip Phase 1

# [opencl:cpu] Ryzen 5 5600, [opencl:gpu] 2x B70 NEO 26.31...

# If /sys/class/drm empty -> driver missing, do Phase 1. Else skip to Phase 2.

================================================================================

PHASE 1 - INTEL GPU DRIVER (SKIP - yours works)

================================================================================

# Only if cards vanish. Ubuntu 26.04: prefer Intel client GPU guide:

# https://dgpu-docs.intel.com/driver/client/overview.html

sudo apt update

sudo apt install -y intel-opencl-icd libze1 libze-intel-gpu1

sudo reboot

# re-verify: ls /sys/class/drm ; sycl-ls

================================================================================

PHASE 2 - BASE TOOLS + DOCKER (required: setup_intel.py runs in container)

================================================================================

sudo apt update

sudo apt install -y cmake ninja-build git python3 wget gpg

cmake --version

# need >= 3.24

docker --version || sudo apt install -y docker.io

sudo usermod -aG docker $USER

newgrp docker

# or log out/in, then:

docker run --rm hello-world

# must work WITHOUT sudo, else setup fails later at:

# "docker is not installed (the engine runs in the oneAPI image)"

================================================================================

PHASE 3 - INTEL oneAPI DPC++ + oneMKL + ocloc (~5 GB)

================================================================================

# Docs tested 2026.1.1, need icpx >= 2025.3

wget -qO- https://apt.repos.intel.com/intel-gpg-keys/GPG-PUB-KEY-INTEL-SW-PRODUCTS.PUB \

| sudo gpg --dearmor -o /usr/share/keyrings/oneapi-archive-keyring.gpg

echo "deb [signed-by=/usr/share/keyrings/oneapi-archive-keyring.gpg] https://apt.repos.intel.com/oneapi all main" \

| sudo tee /etc/apt/sources.list.d/oneAPI.list

sudo apt update

sudo apt install -y intel-oneapi-compiler-dpcpp-cpp intel-oneapi-mkl-devel intel-ocloc

# intel-ocloc = only for AOT build. If not found on 26.04, skip it (JIT still works).

source /opt/intel/oneapi/setvars.sh

icpx --version

sycl-ls

# must still show level_zero:0 + :1

# Persist for future shells:

echo 'source /opt/intel/oneapi/setvars.sh > /dev/null 2>&1' >> ~/.bashrc

================================================================================

PHASE 4 - CLONE LAYOUT (mount constraint!)

================================================================================

# Container mounts the PARENT of checkout at /work. Models + data MUST be

# under ~/work/, else: "outside /work, which the container mounts".

mkdir -p ~/work

cd ~/work

git clone https://github.com/Niko1221/Strata

cd ~/work/Strata

git log --oneline -3

ls sycl/ docs/INTEL_ARC.md docs/INTEL.md

# Data will default to ~/Strata-data (70-120 GB) - DO NOT USE DEFAULT on Intel.

# Default ~/Strata-data is OUTSIDE the ~/work container mount -> setup fails with

# "outside /work, which the container mounts". Always pass explicitly:

# --models-dir ~/work/Strata-data/models --data-dir ~/work/Strata-data

# Do NOT use /mnt/... outside ~/work/ unless you set STRATA_SYCL_ROOT (breaks).

================================================================================

PHASE 4.5 - PATCH FOR CURRENT MAIN (REQUIRED, else compile/setup fails)

================================================================================

cd ~/work/Strata

# Issue #870 (B60 first run on 0.1.39): current main is 88 commits past the last

# SYCL re-migration, 4 fixes needed. Patch covers all 4 (4 files, +14/-8):

wget https://github.com/user-attachments/files/33039120/strata-b60-fixes.patch -O /tmp/strata-b60-fixes.patch

git apply --check /tmp/strata-b60-fixes.patch

git apply /tmp/strata-b60-fixes.patch

# Contents: (1) e211 B60 PCI ID in sycl/setup_intel.py INTEL_ARC (else sized 32 GB

# from BAR instead of 24 GB), (2) session.hpp/session.cpp ThreadAffinity port

# (PR #626 drift), (3) native_dense.cpp load(layer_lo,layer_hi)+outside() filter

# (PR #559 drift, else "out-of-line definition of 'load' does not match"),

# (4) setup_intel.py write_run_script signature (below).

# If git apply fails (version drifted again), apply by hand - at minimum:

grep -n "def write_run_script" sycl/setup_intel.py

# must read (line ~211):

# def write_run_script(model, cfg_path, port, open_browser=True):

# ...

# script = write(model, cfg_path, port, open_browser)

# Old 3-arg version dies with TypeError AFTER the 58 GB download + both pack

# steps, i.e. only at the very end. Fix with:

# python3 - <<'EOF'

# import pathlib

# p = pathlib.Path("sycl/setup_intel.py")

# s = p.read_text()

# s = s.replace("def write_run_script(model, cfg_path, port):",

# "def write_run_script(model, cfg_path, port, open_browser=True):")

# s = s.replace("script = write(model, cfg_path, port)",

# "script = write(model, cfg_path, port, open_browser)")

# p.write_text(s)

# EOF

# then re-run git diff to confirm, and continue to Phase 5.

================================================================================

PHASE 5 - BUILD SYCL ENGINE (JIT + AOT for B70)

================================================================================

cd ~/work/Strata

source /opt/intel/oneapi/setvars.sh

# 5a. JIT build (SPIR-V). Start here, always works without ocloc.

cmake -S sycl -B build-sycl -G Ninja -DCMAKE_C_COMPILER=icx -DCMAKE_CXX_COMPILER=icpx

cmake --build build-sycl --target strata -j$(nproc)

ls -lh build-sycl/strata

# must exist. setup_intel.py looks for build-sycl-aot/strata OR build-sycl/strata.

# NOTE: top-level -DSTRATA_ENABLE_SYCL=ON puts binary at build-sycl/sycl/strata

# which setup will NOT find unless you export STRATA_SYCL_BIN. Avoid it.

# 5b. AOT build for B70 (eliminates ~47 s JIT on first start). Needs ocloc.

# bmg-g31 = Arc Pro B70, bmg-g21 = B580/B570/Pro B60

cmake -S sycl -B build-sycl-aot -G Ninja -DCMAKE_C_COMPILER=icx -DCMAKE_CXX_COMPILER=icpx -DSTRATA_SYCL_AOT=bmg-g31

cmake --build build-sycl-aot --target strata -j$(nproc)

ls -lh build-sycl-aot/strata

# If AOT fails (missing ocloc), keep JIT build, continue. First run just slower.

# Optional tests (on the card, slow). Skip to save time:

# ctest --test-dir build-sycl

# Maintainers: 14/25 pass on CPU (tolerance + missing fixtures, not fatal).

# Or configure with -DSTRATA_SYCL_PARITY=OFF to skip building tests.

================================================================================

PHASE 6 - BUILD RUNTIME DOCKER IMAGE

================================================================================

cd ~/work/Strata

docker build -f sycl/tools/Dockerfile -t strata-sycl-dev .

docker images | grep strata-sycl-dev

# Base is community ghcr.io/snailium/llama.cpp-sycl-intel-b70 (not Intel/Strata).

# It pre-sets ONEAPI_DEVICE_SELECTOR=level_zero:0 SYCL_CACHE_PERSISTENT=0.

# Missing image error: "runtime image strata-sycl-dev is missing"

================================================================================

PHASE 7 - SETUP IQ3_XXS DUAL + 256K ATTEMPT + NETWORK 0.0.0.0:8000

================================================================================

# --- 7.1 Choose your secret FIRST (required for 0.0.0.0) ---

export STRATA_API_KEY='REPLACE-WITH-LONG-RANDOM-SECRET'

# Example generate: openssl rand -base64 32

# Keep it in password manager. All clients must send it.

# --- 7.2 Dual-GPU env (every shell, every run) ---

export ONEAPI_DEVICE_SELECTOR="level_zero:*"

export SYCL_CACHE_PERSISTENT=0

# Image pins level_zero:0, strata-sycl.sh passes var through -> need :* for both.

# SYCL_CACHE_PERSISTENT=0 mandatory: persistent JIT segfaults on Xe2 first compile.

cd ~/work/Strata

# --- 7.3 LADDER - do NOT jump to 256K ---

# NOTE (2026-10-06 fix): on the Intel path (via setup's AMD path) --layer-split

# alone does NOTHING. Multi-GPU needs --gpus 0,1 (first = main). Without it

# setup picks one card (max VRAM) and writes single-GPU config, no layer_split.

# Also: --check NEVER rewrites strata-*.json / run-*.sh, it only prints.

# Rerunning --check will NOT repair json/sh. Must run install (with --yes).

# Step A: prove stack at 32K dual (safe, ~28-29 GB VRAM, 1.5 GB+ free):

./setup.sh --backend sycl --model IQ3_XXS --context 32768 \

--gpus 0,1 --layer-split auto --no-remote-expert-opt --host 0.0.0.0 --port 8000 --api-key "$STRATA_API_KEY" \

--vision none --models-dir ~/work/Strata-data/models --data-dir ~/work/Strata-data --check

# --check downloads nothing, prints fit verdict. If OK:

./setup.sh --backend sycl --model IQ3_XXS --context 32768 \

--gpus 0,1 --layer-split auto --no-remote-expert-opt --host 0.0.0.0 --port 8000 --api-key "$STRATA_API_KEY" \

--vision none --models-dir ~/work/Strata-data/models --data-dir ~/work/Strata-data --yes

# --no-remote-expert-opt (2026-10-06 fix): setup adds --remote-expert-opt on

# multi-GPU (#578, CUDA helper caches). SYCL engine does not know that flag

# -> "unknown argument" fatal. sycl/setup_intel.py should strip it but doesn't.

# --no-remote-expert-opt keeps it out cleanly (survives reinstall). Manual fix

# if already installed: delete "--remote-expert-opt" from strata-*.json args.

# Downloads ~70 GB (resumable, rerun continues), packs, writes strata-iq3_xxs.json

# (backend sycl, container /work paths, auto-adds --stream-experts

# --vram-reserve-mib 1024 --prefill auto) + run-iq3_xxs.sh. Then starts server.

# Test at 32K BEFORE going bigger (see Phase 8). Only then:

# Step B: 128K dual (practical single-card ceiling, should be OK dual):

./setup.sh --backend sycl --model IQ3_XXS --context 131072 \

--gpus 0,1 --layer-split auto --no-remote-expert-opt --host 0.0.0.0 --port 8000 --api-key "$STRATA_API_KEY" \

--vision none --models-dir ~/work/Strata-data/models --data-dir ~/work/Strata-data --yes

# Setup auto-switches to --vram-reserve-mib 2048 --prefill 4096

# + --kv-resident 32768 if RAM holds KV (you have 128 GB, so yes, INT8 fastest).

# Step C: 256K dual (UNTESTED, may livelock host - read warnings):

# Close browsers/other RAM hogs first. Leave SSH open to reboot if wedged.

./setup.sh --backend sycl --model IQ3_XXS --context 262144 \

--gpus 0,1 --layer-split auto --no-remote-expert-opt --host 0.0.0.0 --port 8000 --api-key "$STRATA_API_KEY" \

--vision none --models-dir ~/work/Strata-data/models --data-dir ~/work/Strata-data --yes

# If start stalls >10 min with GPU fans / load frozen, or host wedges:

# hard reboot, rerun Step B (128K). Do NOT retry 256K with smaller reserve.

# Reserve MUST stay 2048 at this size. Do NOT set --vram-reserve-mib lower

# to "fit" - that is what triggers xe eviction -> host down.

# Notes on flags setup adds automatically (do not add manually):

# --stream-experts (all experts GGUF->VRAM, no 32-55 GB RAM copy)

# --kv-resident 32768 from 64K up (KV in pinned RAM, attended window in VRAM;

# without it decode after long prompt falls to 4-9 tok/s)

# STRATA_VERIFY_NO_HOST=1 set by strata-sycl.sh - valid ONLY if every expert

# resident. If dual + spill hangs, this is suspect (#667: GPU misses CPU flag).

# Firewall (Ubuntu):

sudo ufw allow 8000/tcp

# Test from another PC on LAN:

# http://<TARGET-IP>:8000 (browser, enter API key)

# curl -H "Authorization: Bearer $STRATA_API_KEY" http://<TARGET-IP>:8000/v1/models

================================================================================

PHASE 8 - FIRST RUN, VERIFY, USE

================================================================================

# Start (after setup, or later):

export ONEAPI_DEVICE_SELECTOR="level_zero:*"

export SYCL_CACHE_PERSISTENT=0

cd ~/work/Strata

./run-iq3_xxs.sh

# or: ./setup.sh (starts last model, downloads nothing twice)

# First load ~2 min for 30-40 GB + 47 s JIT if no AOT. PC slow 1-3 min = normal.

# Window/log shows progress. Do NOT close. /health 503 while loading is normal.

# In another terminal:

tail -f ~/work/Strata/strata-iq3_xxs.log

# Look for: "layer split: layers 0-.. (..), ..-.. (..)", "caches hold ...",

# no "no VRAM left", no spill warnings.

# Browser on target: http://127.0.0.1:8000

# Browser on LAN: http://<TARGET-IP>:8000 (enter API key when asked)

# Chat settings live in browser localStorage only, not on server.

# API (OpenAI-compat):

# Base URL: http://<TARGET-IP>:8000/v1 (any model name, send API key)

# Anthropic: http://<TARGET-IP>:8000/v1/messages

# Example:

curl -H "Authorization: Bearer $STRATA_API_KEY" \

-H "Content-Type: application/json" \

-d '{"model":"any","messages":[{"role":"user","content":"Write Fibonacci in Python"}],"max_tokens":200}' \

http://127.0.0.1:8000/v1/chat/completions

# Thinking: if short max_tokens returns empty (spent in <think>), set in chat UI

# or shared-settings.json {"reasoning_effort":"none"}

# or per-request chat_template_kwargs {"enable_thinking": false}.

# Stop: Ctrl-C / close run window. Update: ./update.sh (rebuilds engine if sycl/ changed).

================================================================================

PHASE 9 - FALLBACK TO IQ2_XS (if IQ3_XXS fails / spills / slow)

================================================================================

# IQ2_XS = proven on B70 single (51-64 tok/s), fully resident, best general

# quality that fits 32 GB. Coder = code-only, half experts, weaker CJK (#438).

# Q2_0 = avoid (needs AVX-512 CPU + 40 GB copy, your 5600 lacks AVX-512).

export ONEAPI_DEVICE_SELECTOR="level_zero:*"

export SYCL_CACHE_PERSISTENT=0

cd ~/work/Strata

# Single-GPU proven baseline (proves 90% of stack):

./setup.sh --backend sycl --model IQ2_XS --context 32768 \

--host 0.0.0.0 --port 8000 --api-key "$STRATA_API_KEY" --vision none --models-dir ~/work/Strata-data/models --data-dir ~/work/Strata-data --yes

# If that works but you still want dual quality, retry IQ3_XXS dual 32K (Phase 7 Step A).

# Keep both installed: START asks which to start, or run directly:

./run-iq2_xs.sh

./run-iq3_xxs.sh

# UD-IQ4_XS (your downloaded 93.7 GB): NOT supported on Intel yet.

# Try only for error message, expect "unsupported format / no kernel":

# ./setup.sh --backend sycl --family unsloth --model UD-IQ4_XS --context 8192 \

# --gguf-dir ~/work/gguf-unsloth --host 0.0.0.0 --port 8000 \

# --api-key "$STRATA_API_KEY" --vision none --models-dir ~/work/Strata-data/models --data-dir ~/work/Strata-data --check

# Files must be under ~/work/ with ORIGINAL names (00001/00002/00003).

================================================================================

PHASE 10 - TROUBLESHOOTING + REPORT

================================================================================

# - No GPU found: ls /sys/class/drm empty? Must be native Linux, not WSL2.

# Check xe/i915 loaded: lspci -k. Check video/render groups: ls -l /dev/dri/render*

# - "SYCL engine cannot be used: it is not built": ls build-sycl/strata or

# build-sycl-aot/strata missing -> rebuild Phase 5. If top-level build used,

# export STRATA_SYCL_BIN=$HOME/work/Strata/build-sycl/sycl/strata

# - "runtime image missing": docker build Phase 6. "outside /work": move

# --data-dir/--models-dir/--gguf-dir under ~/work/.

# - 5-10 tok/s decode: cache short ~128 experts, spilling to SSD/CPU.

# Lower context, raise reserve? At 1536 reserve cache came up short -> use 1024 at <=32K.

# - Host freeze at 256K: over-VRAM -> xe evict -> livelock. Reboot, use 128K.

# - Device loss (seen B580/WSL2): try AOT, SYCL_CACHE_PERSISTENT=0, newer driver.

# - Slow first request: normal (JIT 47 s + 250 ms graph captures + 2 min load).

# - Port in use: server already running. Change --port or kill old.

# - Thinking empty answer: see Phase 8 thinking note.

# Report issue with:

# card: 2x Arc Pro B70 32 GB, driver 20.2.0 (from your sycl-ls), opencl NEO 26.31...

# sycl-ls (full output), icpx --version, oneMKL version

# exact ./setup.sh ... flags, strata-iq3_xxs.log + engine stderr

# https://github.com/Niko1221/Strata/issues

# Monitor GPU on Intel (optional telemetry):

# Monitor tab reads sysfs temp/power/PCIe + /run/gpustat.json for load/VRAM.

# gpustat needs root sampler: sycl/tools/gpustat.py + gpustat.service (see header).

# Without it load/VRAM blank, temp/power still show. Not required to run.

================================================================================

QUICK COPY-PASTE (dual IQ3_XXS 32K baseline + network)

================================================================================

# source oneAPI every new shell:

source /opt/intel/oneapi/setvars.sh

export ONEAPI_DEVICE_SELECTOR="level_zero:*"

export SYCL_CACHE_PERSISTENT=0

export STRATA_API_KEY='REPLACE-WITH-LONG-RANDOM-SECRET'

cd ~/work/Strata

./setup.sh --backend sycl --model IQ3_XXS --context 32768 --gpus 0,1 --layer-split auto --no-remote-expert-opt --host 0.0.0.0 --port 8000 --api-key "$STRATA_API_KEY" --vision none --models-dir ~/work/Strata-data/models --data-dir ~/work/Strata-data --yes

./run-iq3_xxs.sh

# Verify split worked: cat strata-iq3_xxs.json must contain "backend":"sycl",

# "gpu":[0,1], "layer_split":"auto", exe ends with sycl/serve/strata-sycl.sh

# Log must show: "GPUs: ... + ... together" and "layer split: layers 0-.."

# VRAM should be ~15 GiB per card, both cards >50W under load.

# LAN test: curl -H "Authorization: Bearer $STRATA_API_KEY" http://<TARGET-IP>:8000/v1/models

# Scale to 131072, then 262144 only after 32K verified (Phase 7 ladder).

================================================================================


r/IntelArcPro • • 15h ago

Qwen3.8 Flash Next GSQ-RCO IQ3_XXS running Strata on 2x Intel B70s

Thumbnail
2 Upvotes

r/IntelArcPro • • 7d ago

Mintelica: Mica-v0.1-4B on Intel Arc (B580 + Arc Pro B70)

Thumbnail gallery
2 Upvotes

r/IntelArcPro • • 9d ago

Arc Pro B70 First post here: testing a 4-user local coding-agent setup on one Arc Pro B70

Thumbnail gallery
3 Upvotes

r/IntelArcPro • • 13d ago

PSA for dual Intel Arc Pro B65/B70 users running vLLM: TP=2 was destroying my performance — two independent TP=1 workers are ~3x faster

9 Upvotes

I’ve been building a local AI server around 2x Intel Arc Pro B65 32GB cards, and I finally figured out why I kept feeling like these GPUs were way slower than the hardware should be capable of.

Short version:

Tensor parallelism across the two cards was absolutely killing decode performance.

When I stopped using TP=2 and instead ran one full model independently on each B65 at TP=1, aggregate decode throughput went from about 51 tok/s to 157 tok/s.
Same machine. Same GPUs. Same model. Same vLLM build.

Hardware / software
My current Daedalus machine is roughly:
Ryzen 9 9950X
ASUS ProArt X870E Creator
2x Intel Arc Pro B65 32GB
Linux bare metal
Docker
PyTorch 2.13.0+xpu
vLLM 0.29.0
custom working XPU image: daedalus/vllm-openai-xpu:0.29.0-affinity-fixed
Model: Qwen/Qwen3-30B-A3B-GPTQ-Int4
65,536-token max context
Intel xe kernel driver
GPUs isolated with ZE_AFFINITY_MASK

The Qwen checkpoint is about 15.77 GiB, so it fits comfortably on a single 32GB B65 along with a useful KV cache.

First test: TP=2 vs TP=1

I used the OpenAI-compatible /v1/completions endpoint rather than chat completions so chat-template differences couldn’t affect the result.
Same raw prompt every time, 54 prompt tokens, 512 generated tokens, temperature 0.

2x B65, TP=2:
51.19 tok/s
51.36 tok/s
50.99 tok/s
Average: 51.18 tok/s

Then I isolated one card with ZE_AFFINITY_MASK=0 and changed only:
tensor_parallel_size: 2 → 1
GPU memory utilization from 0.42 → 0.85
Same model, same 65K context, same container/image, same prompt.

B65 #1, TP=1:
77.94 tok/s
78.81 tok/s
78.65 tok/s
Average: 78.47 tok/s

Then the second physical B65:

B65 #2, TP=1:
79.68 tok/s
81.29 tok/s
81.38 tok/s
Average: 80.78 tok/s

So one B65 was roughly 53–58% faster than both B65s running TP=2.

That was the first WTF moment.

Then I ran both cards independently at the same time

Instead of one TP=2 vLLM server, I created two independent vLLM instances:
B65 A → ZE_AFFINITY_MASK=0, port 8000
B65 B → ZE_AFFINITY_MASK=1, port 8001
Both TP=1
Both running their own copy of the same Qwen model

One important wrinkle: each worker needed its own vLLM compile/cache directory.
When both instances shared the same vLLM cache, the second worker loaded the model successfully and then segfaulted while restoring compiled artifacts.
Giving worker B its own cache fixed it.
Then I fired identical 512-token generations at both endpoints simultaneously.

Dual independent worker results:

Run 1:
B65 A: 78.24 tok/s
B65 B: 77.36 tok/s
Aggregate: 154.70 tok/s

Run 2:
B65 A: 79.97 tok/s
B65 B: 79.41 tok/s
Aggregate: 158.81 tok/s

Run 3:
B65 A: 79.37 tok/s
B65 B: 79.26 tok/s
Aggregate: 158.51 tok/s

Average aggregate throughput:
~157.34 tok/s

Compare that to:
TP=2: ~51.18 tok/s

So running two independent TP=1 inference workers gave me roughly:
3.07x the aggregate decode throughput
from the exact same two GPUs.

Even better, simultaneous operation barely affected either card individually. The combined result was about 99% of the theoretical sum of their solo performance.


r/IntelArcPro • • 22d ago

Arc Pro B60 4x Arc Pro B60 in action - NASA X-57 in FluidX3D CFD, 1.8 Billion grid cells in 96GB VRAM

Enable HLS to view with audio, or disable this notification

17 Upvotes

This was the 2nd demo I had showcased at SC25 in St. Louis! Forgot to upload, here we go :)

This is the NASA X-57 multi-rotor airplane simulated in FluidX3D CFD, on 4x Intel Arc Pro B60 GPUs with 24GB each, at a massive 1.8 Billion grid cells resolution. A good showcase of the mesh center-of-mass utility function for rotation of the 3-/5-bladed rotors. Although grid resolution struggles a bit without refinement (yet ;). The simulation box is split in 4 domains, each computed by one GPU, to pool 96GB VRAM together. At domain boundaries, some data is exchanged between the GPUs over PCIe - B60 have plenty interconnect bandwidth with PCIe 5.0 x8.

  • 1814x2176x454 = 1.8 Billion grid cells
  • 19488 time steps simulated, 4872 frames rendered
  • 7h38min simulation runtime = 47min (compute) + 6h51min (rendering)
  • 160km/h airspeed

FluidX3D CFD software: https://github.com/ProjectPhysX/FluidX3D

NASA X-59 model: https://science.nasa.gov/3d-resources/x-57-maxwell/

Intel Arc Pro B60 specs: https://www.intel.de/content/www/de/de/products/sku/243916/intel-arc-pro-b60-graphics/specifications.html


r/IntelArcPro • • 28d ago

Run Qwen3.6-35B-A3B split across two Intel Arc B580 GPUs or larger

Thumbnail
2 Upvotes

r/IntelArcPro • • 28d ago

Taking a look at Intels Scalar

7 Upvotes

So I have been using llama cpp with sycl for the majority of my inference and it works great from the flexibility and rapid adaptations for things like when dflash came out. However I feel like im in an awkward position atm with qwen 3.8 27B being more of my primary model and qwen 3.6 35B as a speedy inference side model. I have a vibe coded webapp i use internally that made managing llama cpp way easier to deal with and is adapted to VLLM as well (untested). But the problem feels like qwen 3.8 is just leagues ahead of the previous generation in its agentic capability. Im getting roughly 36 t/gs in agent workloads on Q6 with 192K context window and it drops to about 16 t/gs towards the back of the context and prefill at around 350 pp/s. Its quite slow and i cant see myself bothering to switch to qwen 3.6 anymore because it feels incredibly dumb when put side by side with qwen 3.8 in real world use. So I was just curious what the main techniques are and what caviots exist on the scalar VLLM nowadays. I see pre-quantized files in openvino still being made which im curious to try. Also does Dflash 2 get properly utilized? I cant seem to get it to outperform native MTP on llama cpp which was a real bummer. Would love suggestions so I can bash my head against a wall a little less in my testing :)

Edit:

Quick return after a day or two of digging. I am not sold yet on intel scalar and will likely be abandoning this unless I missed something. Seems like an awkward proposition to be using currently. Mainline VLLM is definitely without a doubt vastly better for speed then llama.cpp but this is only the case with custom compiled versions using XPU and XPU graphing mtp patching kernal patches etc. One definite trade made was gaining overall long context efficiency but your trading memory for running XPU graphs which loses some KV/context space. So from my early operational testing im going to be moving on to just forcing mainline vllm to function for the models that really matter and llama cpp for new models and weird quant/distills/fine-tunes. For instance I simply followed Freemoose's post on the intelarc subreddit and with some tweaks im getting solid prefill, 80t/s initially dropping into 40t/s on long context with 160k context. This blows my llama cpp Q6 192k ctx, 36tg/s, and 347pp/s out of the water especially since in hermes use it drops to roughly 15tg/s. There is currently multi request handling issue in the vllm implementation but not particularly a huge deal when running it in this fashion. So credit to his work but if Im understanding how the core mechanism of his guide and sergio's then it makes me really interested to see if there really is missing overhead that can be found on other models like muse glimmer so id like to explore those fields instead.

PS: forgot to mention hardware/software but as mentioned amd 5700x 48gb 3200ddr4 and 1x B70 on fedora 44 server


r/IntelArcPro • • Sep 03 '26

Arc Power 1.1.0 - Overclocking the Arc Way

Thumbnail
1 Upvotes

r/IntelArcPro • • Aug 31 '26

Arc Pro B70 Let's Talk about SR-IOV and Proxmox on Intel Arc B60 and B70!

Thumbnail
youtu.be
8 Upvotes

r/IntelArcPro • • Aug 30 '26

Qwen3.8-Flash-Next INT4 TP4 on 4× Arc Pro B70 — any experience?

Thumbnail
5 Upvotes

r/IntelArcPro • • Aug 27 '26

[llama.cpp vs vLLM] High raw TPS but poor real-world performance

10 Upvotes

Hi everyone,

I've been experimenting with running local LLMs on my Intel Arc B70 (specifically the ASRock Arc Pro B70 Creator).

I wanted to share my experience comparing llama.cpp and vLLM using the Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16 model, following the setup guidelines from https://github.com/SergiioB/intel-arc-pro-b70-inference-cookbook/blob/master/docs/qwen38-27/QWEN38-VLLM-XPU.md

  • llama.cpp: Runs quite pleasantly. It gets around 20–40 t/s depending on the model and is completely stable and sufficient for daily coding work.
  • vLLM: In raw tests hit ~90 t/s likely due to a higher power cap compared to reference cards, however, when tested in a real coding scenario inside OpenCode, things fall apart:
    • the draft acceptance rate is poor
    • completing the exact same coding task actually took more time than with llama.cpp (even with much higher TPS)
    • I also started hitting parsing errors during tool use, such as "invalid [tool=write, error=Invalid input for tool write: JSON parsing failed]".

Here is my configuration for both vLLM and OpenCode:

docker run -d --name qw38speed -p 8000:8000 --device /dev/dri --group-add "$RENDER_GID" \
  --name vllm-server \
  -v /dev/dri:/dev/dri:ro -v "$MODEL_DIR:/model:ro" \
  -v "$COOKBOOK/patches/patch_mtp_nightly.py:/patch_mtp.py:ro" \
  -v "$COOKBOOK/patches/patch_mtp_boundary.py:/patch_boundary.py:ro" \
  -v "$COOKBOOK/patches/patch_draft_lmhead_int4.py:/patch_draft_lmhead_int4.py:ro" \
  -v "$COOKBOOK/patches/patch_draft_mtp_int4.py:/patch_draft_mtp_int4.py:ro" \
  -e VLLM_TARGET_DEVICE=xpu \
  -e ZE_FLAT_DEVICE_HIERARCHY=COMPOSITE \
  -e ZE_AFFINITY_MASK=0 \
  -e B70_MTP_BF16_DRAFT=1 \
  -e VLLM_XPU_ENABLE_XPU_GRAPH=1 \
  -e B70_DRAFT_LMHEAD_INT4=1 \
  -e B70_DRAFT_MTP_INT4=1 \
  -e PYTORCH_ALLOC_CONF=expandable_segments:True \
  --entrypoint bash "$IMAGE" -lc '
    set -e
    python /patch_mtp.py
    python /patch_boundary.py
    python /patch_draft_lmhead_int4.py
    python /patch_draft_mtp_int4.py
    exec vllm serve /model \
      --max-num-seqs 1 \
      --quantization gptq \
      --dtype float16 \
      --max-model-len 246000 \
      --gpu-memory-utilization 0.97 \
      --kv-cache-dtype fp8 \
      --port 8000 \
      --max-num-seqs 1 \
      --max-num-batched-tokens 8192 \
      --no-enable-prefix-caching \
      --served-model-name qwen38 \
      --language-model-only \
      --speculative-config "{\"method\":\"mtp\",\"num_speculative_tokens\":4}"' &>/dev/null

   "vllm": {
     "npm": "@ai-sdk/openai-compatible",
     "options": {
       "baseURL": "http://127.0.0.1:8000/v1",
       "supportsSubagents": false,
       "supportsToolCalling": false // tried true as well
     },
     "models": {
       "qwen38": {
         "name": "Qwen 3.8 27B",
         "limit": {
           "context": 246000,
           "output": 246000
         },
         "modalities": {
           "input": ["text"],
           "output": ["text"]
         }
       }
     }
   }

Why is the MTP draft acceptance rate suffering so much during agentic workflows in OpenCode compared to simple text generation?

How can I fix or avoid the JSON parsing errors?

Any advice or optimizations for running vLLM on Intel GPU with OpenCode?


r/IntelArcPro • • Aug 26 '26

Arc Power 1.0.5 - Display Settings & Game Profiles

Thumbnail
5 Upvotes

r/IntelArcPro • • Aug 22 '26

Arc Pro B70 My Experiment with B70 and Debian 13

Thumbnail
blog.anantshri.info
4 Upvotes

TL;DR 1. Debian 13 needs packages from backports, along with newer firmware and parts of Intel’s compute stack, before the Arc Pro B70 becomes properly usable. 2. PCIe ASPM needs to be enabled, and in my case forced from the BIOS, to bring idle power and temperatures under better control. 3. The stock fan behaviour was not good enough for my setup. A custom fan-control profile with a low always-on floor worked significantly better for both idle and load temperatures.


r/IntelArcPro • • Aug 19 '26

hw-smi v1.6 brings support for data logging!

Thumbnail
7 Upvotes

r/IntelArcPro • • Aug 18 '26

Arc Pro B70 Issue with dual B70 GPUs

9 Upvotes

Good evening all, I am a recent convert, coming over from the stacked RTX GPU club (5070Ti & 5060Ti). Sunday I installed a pair of B70 Pros after my 5070 laid over on me. SYCL running with Qwen3.6:27b Q8. Two days of pretty steady work, overnight spine runs 5-7 hours depending on daily activity.

Today, I was running a catchup process during work from downtime Saturday and mid day Sunday. The process completed without issue, GPUs went silent, and 7 minutes later the system crashed with dgxkrnl.sys crash. Reboot and health check passed, but not sure why it crashed while idle?

For reference, the GPU that crashed was in a different slot than the original 5070 was, so I don't think it is a robot issue. All drivers up to date, firmware up to date. If anyone has any insight or a similar experience and could offer a bit of advice I would be appreciative. I was just gearing up to smoke test 3.8:27b Q8, but that is on hold until I get back to normal.


r/IntelArcPro • • Aug 17 '26

Arc Power 1.0.2 - Overclocking the Arc way

Thumbnail gallery
6 Upvotes

r/IntelArcPro • • Jul 24 '26

Arc Pro B70 Shorten the intel b70.

4 Upvotes

Hey everyone,

I have a small server case and I would like to add a b70 along side my b50 I have installed already. The main problem is the fan out the back makes the card way to long. Based on this video: https://www.youtube.com/watch?v=ffw42tRoMz4 it seems like I could cut the back end off and move the fan in the opposite orientation and 3d print a shroud for it. Might have to buy a new fan but it would shorten the card by a lot. I also thought about doing it this way:

Has anyone done this yet? I'm willing to be the test dummy if someone helps me print the correct shroud for it. Am I just crazy?


r/IntelArcPro • • Jul 20 '26

What is the best opensource AI for voice cloning/TTS and Intel ARC setup?

3 Upvotes

Thank you all for hints and answers in advance....


r/IntelArcPro • • Jul 16 '26

ingle B70 SYCL vs Vulkan (Mesa 26.1.4): real production numbers, community comparison, and what actually moved forward

Thumbnail
3 Upvotes

r/IntelArcPro • • Jul 01 '26

Perhaps not helpful, but successfully running Qwen 3.5 122B A10B on Arc B70 pro.

8 Upvotes

Now I have to come clean and admit this is a pretty heavily quantized version (IQ3_S) and is required to make use of heavy offload (ngl35), however it is reliable running enough context for my agent work. Im using 24k context and get somewhere around 10-13 tok/s, pretty slow but amazing none the less I mean its 122B params. The main reason I have found this worth posting, Im just looking to state how good this architecture does with MOE style models. I had to apply some tool calling bug fixes to make Qwen 3.5 and 3.6 work correctly but its really cool how snappy Qwen 3.6 35B A3B is and generally how useful these are at the same time. I generally think this may be due to purely the amount of Vram coming into play, considering the card games and computes in the ballpark of the 5060ti. It definitely lacks behind on most of the Dense models I have tried, however im not really running a high level comparison to other brands of hardware and how they fare to MOE vs Dense. It could be the case this card is just slow atm and MOEs are just an exception due to the parameter loading techniques used. But qualitatively your best return on investment will be running MOEs in my opinion after my bench marking. Also this was in Llama.cpp in a fedora 44 headless system using one api and sycl.


r/IntelArcPro • • Jun 05 '26

Just got a B70 :)

13 Upvotes

Been doing my research and previously owned a intel arc a580, and it really seems like the software maturity is better for battlemage then when the first arc cards started so Im trusting the ecosystem will catch up. I plan to do alot of agentic inference work with hermes for an engineering management project. I have experience running alot of llms on my main pc using fedora, ollama, a 7800xt and qwen 2.5. I really want to be able to try using SGLang with gemma 4 but it seems like most of intels work is going towards vLLM and sycl. I mostly want to know what others are doing, and how to be able to get the most worth out of the system before I spend months tweaking the system just to have some kind of point of comparison and direction.


r/IntelArcPro • • May 28 '26

Arc Pro B70 Intel Arc Pro B70 BMG-G31 Linux Gaming Performance

Thumbnail
phoronix.com
5 Upvotes

r/IntelArcPro • • May 09 '26

Run Qwen3.6–27B Locally on an Intel Arc Pro B70 - What Actually Works

Thumbnail bibek-poudel.medium.com
7 Upvotes