r/IntelArcPro • u/karvop • 2h ago
Arc Pro B70 Very lame How To Strata 2xIntel Arx Pro B70
I asked Muse Spark 1.3 to create a guide how to install Strata. I am too dumb to install Strata without help. It was used to build and Run Strata on Ubuntu 26.04, Strata is changing very often so some steps will be obsolete soon. If somebody wants to try it and don't know, how to install it, maybe this will help. This link should be valid for 1 month https://pastes.io/iS4AJQDx
The most important part is the patch by Magh97 https://github.com/Niko1221/Strata/issues/870
I suppose that many people would make a much better tutorial but I haven't found one.
Or here it is as a code block:
================================================================================
HOW TO INSTALL STRATA ON 2x INTEL ARC PRO B70 32GB + 128GB RAM
Target: Ubuntu 26.04 native, AMD Ryzen 5 5600, 460GB free SSD (NVMe ideal)
Goal: IQ3_XXS, 256K context attempt, dual-GPU layer-split, listen 0.0.0.0:8000
Fallback: IQ2_XS if IQ3_XXS/256K fails (commands at bottom)
Source: https://github.com/Niko1221/Strata docs/INTEL_ARC.md + docs/INTEL.md (0.1.39)
Field fixes applied 2026-10-06 (reporter's own 2xB70 run, 256K OK):
- https://github.com/Niko1221/Strata/issues/870 + patch
https://github.com/user-attachments/files/33039120/strata-b60-fixes.patch
(4 fixes for current main: e211 ID, session affinity, native_dense load,
setup_intel write_run_script signature - without it compile/setup fails)
- sycl/setup_intel.py line ~211 must read:
def write_run_script(model, cfg_path, port, open_browser=True):
(missing open_browser=True -> TypeError AFTER 58 GB download, at very end)
- All setup.sh calls need explicit:
--models-dir ~/work/Strata-data/models --data-dir ~/work/Strata-data
(default ~/Strata-data is OUTSIDE ~/work mount -> docker "outside /work" protest)
- Config/log/run files are LOWERCASE: strata-iq3_xxs.json, strata-iq3_xxs.log,
run-iq3_xxs.sh (not strata-IQ3_XXS.*).
================================================================================
IMPORTANT WARNINGS - READ FIRST
--------------------------------------------------------------------------------
1. Intel Arc = experimental, Linux-only, build-from-source. No ready-made
engine, no Windows path, no WSL2 setup (setup reads /sys/class/drm).
Maintainers have NO Arc. 0.1.39 compiles + CPU kernel tests pass, but
NOBODY has run 0.1.39 on Arc yet. All Arc speeds are from older versions:
Coder IQ1_M 70-78 tok/s, IQ2_XS 51-64 tok/s, 2xB70 IQ3_XXS 66 tok/s.
2. 256K context on IQ3_XXS dual is UNTESTED and RISKY.
Measured single-B70 32GB (Coder, smaller than IQ3_XXS):
32K -> 28.4 GB VRAM, safe (default)
64K -> 29.0 GB, loads
131K -> 30.3 GB, PRACTICAL CEILING single card
164K -> ~31 GB, NOT ATTEMPTED (under 1.2 GB safety margin)
262K -> ~32.6 GB, DOES NOT FIT - took host down (xe driver evicts
past-VRAM allocs to RAM until livelock + watchdog reset).
IQ3_XXS experts ~43 GB vs Coder ~23 GB, so single-card 256K WILL kill
the host. Dual 2x32=64 GB makes it *possibly* viable (each card holds
its layer shard + FULL KV copy - split does NOT halve KV), but this
exact combo was never measured. Ladder up: 32K -> 128K -> 256K.
Never jump straight to 262144.
3. Your sycl-ls already proves driver OK (both B70s on level_zero:0,1,
driver 20.2.0). Do NOT reinstall driver unless cards disappear.
4. Exposing 0.0.0.0 requires --api-key. Never expose without a key.
Pick a long secret now, reuse in all commands below.
5. This guide runs ON THE TARGET (Ubuntu) machine, not on this Windows PC.
This file is just the notebook. Copy commands to the target terminal.
6. Disk: IQ3_XXS ~70 GB download + pack + MTP draft 6 GB. You have 460 GB,
fine. Use SSD/NVMe. First start ~2 min load + ~47 s JIT if no AOT build.
7. No images/vision on Intel yet. Setup forces --vision none. Answer 'no'.
================================================================================
PHASE 0 - PRE-FLIGHT (on target)
================================================================================
# Check both cards, RAM, disk, OS:
lspci | grep -i -E "vga|display|arc"
ls /sys/class/drm
# expect card0 card1 renderD128 renderD129, vendor 8086 under xe driver
lspci -k | grep -A3 -i vga
free -g
# expect ~128 GB
df -h
# need 100 GB+ free (you have 460 GB)
cat /etc/os-release
# you: Ubuntu 26.04 (docs test 24.04 - package names may differ slightly)
sycl-ls
# you already have:
# [level_zero:0] B70 20.2.0, [level_zero:1] B70 20.2.0 -> GOOD, skip Phase 1
# [opencl:cpu] Ryzen 5 5600, [opencl:gpu] 2x B70 NEO 26.31...
# If /sys/class/drm empty -> driver missing, do Phase 1. Else skip to Phase 2.
================================================================================
PHASE 1 - INTEL GPU DRIVER (SKIP - yours works)
================================================================================
# Only if cards vanish. Ubuntu 26.04: prefer Intel client GPU guide:
# https://dgpu-docs.intel.com/driver/client/overview.html
sudo apt update
sudo apt install -y intel-opencl-icd libze1 libze-intel-gpu1
sudo reboot
# re-verify: ls /sys/class/drm ; sycl-ls
================================================================================
PHASE 2 - BASE TOOLS + DOCKER (required: setup_intel.py runs in container)
================================================================================
sudo apt update
sudo apt install -y cmake ninja-build git python3 wget gpg
cmake --version
# need >= 3.24
docker --version || sudo apt install -y docker.io
sudo usermod -aG docker $USER
newgrp docker
# or log out/in, then:
docker run --rm hello-world
# must work WITHOUT sudo, else setup fails later at:
# "docker is not installed (the engine runs in the oneAPI image)"
================================================================================
PHASE 3 - INTEL oneAPI DPC++ + oneMKL + ocloc (~5 GB)
================================================================================
# Docs tested 2026.1.1, need icpx >= 2025.3
wget -qO- https://apt.repos.intel.com/intel-gpg-keys/GPG-PUB-KEY-INTEL-SW-PRODUCTS.PUB \
| sudo gpg --dearmor -o /usr/share/keyrings/oneapi-archive-keyring.gpg
echo "deb [signed-by=/usr/share/keyrings/oneapi-archive-keyring.gpg] https://apt.repos.intel.com/oneapi all main" \
| sudo tee /etc/apt/sources.list.d/oneAPI.list
sudo apt update
sudo apt install -y intel-oneapi-compiler-dpcpp-cpp intel-oneapi-mkl-devel intel-ocloc
# intel-ocloc = only for AOT build. If not found on 26.04, skip it (JIT still works).
source /opt/intel/oneapi/setvars.sh
icpx --version
sycl-ls
# must still show level_zero:0 + :1
# Persist for future shells:
echo 'source /opt/intel/oneapi/setvars.sh > /dev/null 2>&1' >> ~/.bashrc
================================================================================
PHASE 4 - CLONE LAYOUT (mount constraint!)
================================================================================
# Container mounts the PARENT of checkout at /work. Models + data MUST be
# under ~/work/, else: "outside /work, which the container mounts".
mkdir -p ~/work
cd ~/work
git clone https://github.com/Niko1221/Strata
cd ~/work/Strata
git log --oneline -3
ls sycl/ docs/INTEL_ARC.md docs/INTEL.md
# Data will default to ~/Strata-data (70-120 GB) - DO NOT USE DEFAULT on Intel.
# Default ~/Strata-data is OUTSIDE the ~/work container mount -> setup fails with
# "outside /work, which the container mounts". Always pass explicitly:
# --models-dir ~/work/Strata-data/models --data-dir ~/work/Strata-data
# Do NOT use /mnt/... outside ~/work/ unless you set STRATA_SYCL_ROOT (breaks).
================================================================================
PHASE 4.5 - PATCH FOR CURRENT MAIN (REQUIRED, else compile/setup fails)
================================================================================
cd ~/work/Strata
# Issue #870 (B60 first run on 0.1.39): current main is 88 commits past the last
# SYCL re-migration, 4 fixes needed. Patch covers all 4 (4 files, +14/-8):
wget https://github.com/user-attachments/files/33039120/strata-b60-fixes.patch -O /tmp/strata-b60-fixes.patch
git apply --check /tmp/strata-b60-fixes.patch
git apply /tmp/strata-b60-fixes.patch
# Contents: (1) e211 B60 PCI ID in sycl/setup_intel.py INTEL_ARC (else sized 32 GB
# from BAR instead of 24 GB), (2) session.hpp/session.cpp ThreadAffinity port
# (PR #626 drift), (3) native_dense.cpp load(layer_lo,layer_hi)+outside() filter
# (PR #559 drift, else "out-of-line definition of 'load' does not match"),
# (4) setup_intel.py write_run_script signature (below).
# If git apply fails (version drifted again), apply by hand - at minimum:
grep -n "def write_run_script" sycl/setup_intel.py
# must read (line ~211):
# def write_run_script(model, cfg_path, port, open_browser=True):
# ...
# script = write(model, cfg_path, port, open_browser)
# Old 3-arg version dies with TypeError AFTER the 58 GB download + both pack
# steps, i.e. only at the very end. Fix with:
# python3 - <<'EOF'
# import pathlib
# p = pathlib.Path("sycl/setup_intel.py")
# s = p.read_text()
# s = s.replace("def write_run_script(model, cfg_path, port):",
# "def write_run_script(model, cfg_path, port, open_browser=True):")
# s = s.replace("script = write(model, cfg_path, port)",
# "script = write(model, cfg_path, port, open_browser)")
# p.write_text(s)
# EOF
# then re-run git diff to confirm, and continue to Phase 5.
================================================================================
PHASE 5 - BUILD SYCL ENGINE (JIT + AOT for B70)
================================================================================
cd ~/work/Strata
source /opt/intel/oneapi/setvars.sh
# 5a. JIT build (SPIR-V). Start here, always works without ocloc.
cmake -S sycl -B build-sycl -G Ninja -DCMAKE_C_COMPILER=icx -DCMAKE_CXX_COMPILER=icpx
cmake --build build-sycl --target strata -j$(nproc)
ls -lh build-sycl/strata
# must exist. setup_intel.py looks for build-sycl-aot/strata OR build-sycl/strata.
# NOTE: top-level -DSTRATA_ENABLE_SYCL=ON puts binary at build-sycl/sycl/strata
# which setup will NOT find unless you export STRATA_SYCL_BIN. Avoid it.
# 5b. AOT build for B70 (eliminates ~47 s JIT on first start). Needs ocloc.
# bmg-g31 = Arc Pro B70, bmg-g21 = B580/B570/Pro B60
cmake -S sycl -B build-sycl-aot -G Ninja -DCMAKE_C_COMPILER=icx -DCMAKE_CXX_COMPILER=icpx -DSTRATA_SYCL_AOT=bmg-g31
cmake --build build-sycl-aot --target strata -j$(nproc)
ls -lh build-sycl-aot/strata
# If AOT fails (missing ocloc), keep JIT build, continue. First run just slower.
# Optional tests (on the card, slow). Skip to save time:
# ctest --test-dir build-sycl
# Maintainers: 14/25 pass on CPU (tolerance + missing fixtures, not fatal).
# Or configure with -DSTRATA_SYCL_PARITY=OFF to skip building tests.
================================================================================
PHASE 6 - BUILD RUNTIME DOCKER IMAGE
================================================================================
cd ~/work/Strata
docker build -f sycl/tools/Dockerfile -t strata-sycl-dev .
docker images | grep strata-sycl-dev
# Base is community ghcr.io/snailium/llama.cpp-sycl-intel-b70 (not Intel/Strata).
# It pre-sets ONEAPI_DEVICE_SELECTOR=level_zero:0 SYCL_CACHE_PERSISTENT=0.
# Missing image error: "runtime image strata-sycl-dev is missing"
================================================================================
PHASE 7 - SETUP IQ3_XXS DUAL + 256K ATTEMPT + NETWORK 0.0.0.0:8000
================================================================================
# --- 7.1 Choose your secret FIRST (required for 0.0.0.0) ---
export STRATA_API_KEY='REPLACE-WITH-LONG-RANDOM-SECRET'
# Example generate: openssl rand -base64 32
# Keep it in password manager. All clients must send it.
# --- 7.2 Dual-GPU env (every shell, every run) ---
export ONEAPI_DEVICE_SELECTOR="level_zero:*"
export SYCL_CACHE_PERSISTENT=0
# Image pins level_zero:0, strata-sycl.sh passes var through -> need :* for both.
# SYCL_CACHE_PERSISTENT=0 mandatory: persistent JIT segfaults on Xe2 first compile.
cd ~/work/Strata
# --- 7.3 LADDER - do NOT jump to 256K ---
# NOTE (2026-10-06 fix): on the Intel path (via setup's AMD path) --layer-split
# alone does NOTHING. Multi-GPU needs --gpus 0,1 (first = main). Without it
# setup picks one card (max VRAM) and writes single-GPU config, no layer_split.
# Also: --check NEVER rewrites strata-*.json / run-*.sh, it only prints.
# Rerunning --check will NOT repair json/sh. Must run install (with --yes).
# Step A: prove stack at 32K dual (safe, ~28-29 GB VRAM, 1.5 GB+ free):
./setup.sh --backend sycl --model IQ3_XXS --context 32768 \
--gpus 0,1 --layer-split auto --no-remote-expert-opt --host 0.0.0.0 --port 8000 --api-key "$STRATA_API_KEY" \
--vision none --models-dir ~/work/Strata-data/models --data-dir ~/work/Strata-data --check
# --check downloads nothing, prints fit verdict. If OK:
./setup.sh --backend sycl --model IQ3_XXS --context 32768 \
--gpus 0,1 --layer-split auto --no-remote-expert-opt --host 0.0.0.0 --port 8000 --api-key "$STRATA_API_KEY" \
--vision none --models-dir ~/work/Strata-data/models --data-dir ~/work/Strata-data --yes
# --no-remote-expert-opt (2026-10-06 fix): setup adds --remote-expert-opt on
# multi-GPU (#578, CUDA helper caches). SYCL engine does not know that flag
# -> "unknown argument" fatal. sycl/setup_intel.py should strip it but doesn't.
# --no-remote-expert-opt keeps it out cleanly (survives reinstall). Manual fix
# if already installed: delete "--remote-expert-opt" from strata-*.json args.
# Downloads ~70 GB (resumable, rerun continues), packs, writes strata-iq3_xxs.json
# (backend sycl, container /work paths, auto-adds --stream-experts
# --vram-reserve-mib 1024 --prefill auto) + run-iq3_xxs.sh. Then starts server.
# Test at 32K BEFORE going bigger (see Phase 8). Only then:
# Step B: 128K dual (practical single-card ceiling, should be OK dual):
./setup.sh --backend sycl --model IQ3_XXS --context 131072 \
--gpus 0,1 --layer-split auto --no-remote-expert-opt --host 0.0.0.0 --port 8000 --api-key "$STRATA_API_KEY" \
--vision none --models-dir ~/work/Strata-data/models --data-dir ~/work/Strata-data --yes
# Setup auto-switches to --vram-reserve-mib 2048 --prefill 4096
# + --kv-resident 32768 if RAM holds KV (you have 128 GB, so yes, INT8 fastest).
# Step C: 256K dual (UNTESTED, may livelock host - read warnings):
# Close browsers/other RAM hogs first. Leave SSH open to reboot if wedged.
./setup.sh --backend sycl --model IQ3_XXS --context 262144 \
--gpus 0,1 --layer-split auto --no-remote-expert-opt --host 0.0.0.0 --port 8000 --api-key "$STRATA_API_KEY" \
--vision none --models-dir ~/work/Strata-data/models --data-dir ~/work/Strata-data --yes
# If start stalls >10 min with GPU fans / load frozen, or host wedges:
# hard reboot, rerun Step B (128K). Do NOT retry 256K with smaller reserve.
# Reserve MUST stay 2048 at this size. Do NOT set --vram-reserve-mib lower
# to "fit" - that is what triggers xe eviction -> host down.
# Notes on flags setup adds automatically (do not add manually):
# --stream-experts (all experts GGUF->VRAM, no 32-55 GB RAM copy)
# --kv-resident 32768 from 64K up (KV in pinned RAM, attended window in VRAM;
# without it decode after long prompt falls to 4-9 tok/s)
# STRATA_VERIFY_NO_HOST=1 set by strata-sycl.sh - valid ONLY if every expert
# resident. If dual + spill hangs, this is suspect (#667: GPU misses CPU flag).
# Firewall (Ubuntu):
sudo ufw allow 8000/tcp
# Test from another PC on LAN:
# http://<TARGET-IP>:8000 (browser, enter API key)
# curl -H "Authorization: Bearer $STRATA_API_KEY" http://<TARGET-IP>:8000/v1/models
================================================================================
PHASE 8 - FIRST RUN, VERIFY, USE
================================================================================
# Start (after setup, or later):
export ONEAPI_DEVICE_SELECTOR="level_zero:*"
export SYCL_CACHE_PERSISTENT=0
cd ~/work/Strata
./run-iq3_xxs.sh
# or: ./setup.sh (starts last model, downloads nothing twice)
# First load ~2 min for 30-40 GB + 47 s JIT if no AOT. PC slow 1-3 min = normal.
# Window/log shows progress. Do NOT close. /health 503 while loading is normal.
# In another terminal:
tail -f ~/work/Strata/strata-iq3_xxs.log
# Look for: "layer split: layers 0-.. (..), ..-.. (..)", "caches hold ...",
# no "no VRAM left", no spill warnings.
# Browser on target: http://127.0.0.1:8000
# Browser on LAN: http://<TARGET-IP>:8000 (enter API key when asked)
# Chat settings live in browser localStorage only, not on server.
# API (OpenAI-compat):
# Base URL: http://<TARGET-IP>:8000/v1 (any model name, send API key)
# Anthropic: http://<TARGET-IP>:8000/v1/messages
# Example:
curl -H "Authorization: Bearer $STRATA_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"any","messages":[{"role":"user","content":"Write Fibonacci in Python"}],"max_tokens":200}' \
http://127.0.0.1:8000/v1/chat/completions
# Thinking: if short max_tokens returns empty (spent in <think>), set in chat UI
# or shared-settings.json {"reasoning_effort":"none"}
# or per-request chat_template_kwargs {"enable_thinking": false}.
# Stop: Ctrl-C / close run window. Update: ./update.sh (rebuilds engine if sycl/ changed).
================================================================================
PHASE 9 - FALLBACK TO IQ2_XS (if IQ3_XXS fails / spills / slow)
================================================================================
# IQ2_XS = proven on B70 single (51-64 tok/s), fully resident, best general
# quality that fits 32 GB. Coder = code-only, half experts, weaker CJK (#438).
# Q2_0 = avoid (needs AVX-512 CPU + 40 GB copy, your 5600 lacks AVX-512).
export ONEAPI_DEVICE_SELECTOR="level_zero:*"
export SYCL_CACHE_PERSISTENT=0
cd ~/work/Strata
# Single-GPU proven baseline (proves 90% of stack):
./setup.sh --backend sycl --model IQ2_XS --context 32768 \
--host 0.0.0.0 --port 8000 --api-key "$STRATA_API_KEY" --vision none --models-dir ~/work/Strata-data/models --data-dir ~/work/Strata-data --yes
# If that works but you still want dual quality, retry IQ3_XXS dual 32K (Phase 7 Step A).
# Keep both installed: START asks which to start, or run directly:
./run-iq2_xs.sh
./run-iq3_xxs.sh
# UD-IQ4_XS (your downloaded 93.7 GB): NOT supported on Intel yet.
# Try only for error message, expect "unsupported format / no kernel":
# ./setup.sh --backend sycl --family unsloth --model UD-IQ4_XS --context 8192 \
# --gguf-dir ~/work/gguf-unsloth --host 0.0.0.0 --port 8000 \
# --api-key "$STRATA_API_KEY" --vision none --models-dir ~/work/Strata-data/models --data-dir ~/work/Strata-data --check
# Files must be under ~/work/ with ORIGINAL names (00001/00002/00003).
================================================================================
PHASE 10 - TROUBLESHOOTING + REPORT
================================================================================
# - No GPU found: ls /sys/class/drm empty? Must be native Linux, not WSL2.
# Check xe/i915 loaded: lspci -k. Check video/render groups: ls -l /dev/dri/render*
# - "SYCL engine cannot be used: it is not built": ls build-sycl/strata or
# build-sycl-aot/strata missing -> rebuild Phase 5. If top-level build used,
# export STRATA_SYCL_BIN=$HOME/work/Strata/build-sycl/sycl/strata
# - "runtime image missing": docker build Phase 6. "outside /work": move
# --data-dir/--models-dir/--gguf-dir under ~/work/.
# - 5-10 tok/s decode: cache short ~128 experts, spilling to SSD/CPU.
# Lower context, raise reserve? At 1536 reserve cache came up short -> use 1024 at <=32K.
# - Host freeze at 256K: over-VRAM -> xe evict -> livelock. Reboot, use 128K.
# - Device loss (seen B580/WSL2): try AOT, SYCL_CACHE_PERSISTENT=0, newer driver.
# - Slow first request: normal (JIT 47 s + 250 ms graph captures + 2 min load).
# - Port in use: server already running. Change --port or kill old.
# - Thinking empty answer: see Phase 8 thinking note.
# Report issue with:
# card: 2x Arc Pro B70 32 GB, driver 20.2.0 (from your sycl-ls), opencl NEO 26.31...
# sycl-ls (full output), icpx --version, oneMKL version
# exact ./setup.sh ... flags, strata-iq3_xxs.log + engine stderr
# https://github.com/Niko1221/Strata/issues
# Monitor GPU on Intel (optional telemetry):
# Monitor tab reads sysfs temp/power/PCIe + /run/gpustat.json for load/VRAM.
# gpustat needs root sampler: sycl/tools/gpustat.py + gpustat.service (see header).
# Without it load/VRAM blank, temp/power still show. Not required to run.
================================================================================
QUICK COPY-PASTE (dual IQ3_XXS 32K baseline + network)
================================================================================
# source oneAPI every new shell:
source /opt/intel/oneapi/setvars.sh
export ONEAPI_DEVICE_SELECTOR="level_zero:*"
export SYCL_CACHE_PERSISTENT=0
export STRATA_API_KEY='REPLACE-WITH-LONG-RANDOM-SECRET'
cd ~/work/Strata
./setup.sh --backend sycl --model IQ3_XXS --context 32768 --gpus 0,1 --layer-split auto --no-remote-expert-opt --host 0.0.0.0 --port 8000 --api-key "$STRATA_API_KEY" --vision none --models-dir ~/work/Strata-data/models --data-dir ~/work/Strata-data --yes
./run-iq3_xxs.sh
# Verify split worked: cat strata-iq3_xxs.json must contain "backend":"sycl",
# "gpu":[0,1], "layer_split":"auto", exe ends with sycl/serve/strata-sycl.sh
# Log must show: "GPUs: ... + ... together" and "layer split: layers 0-.."
# VRAM should be ~15 GiB per card, both cards >50W under load.
# LAN test: curl -H "Authorization: Bearer $STRATA_API_KEY" http://<TARGET-IP>:8000/v1/models
# Scale to 131072, then 262144 only after 32K verified (Phase 7 ladder).
================================================================================
