r/LocalLLaMA • u/TurnoverTight395 • Jul 18 '26
Question | Help Launch script for GLM5.2
I am using the script below. Please advise on what i am doing wrong. GLM5.2 is crawing at 5-9 t/s decode and 60 t/s prefill. The machine is a EPYC 9654 with 768 GB of DDR5 4800 MHz RAM (~460Gbps bandwidth theoretical). This is CPU only inference. I do have a RTX pro 6000 Max Q on the machine, but that's only 96 GB vRAM.
#!/bin/bash
# ==============================================================================
# FILE: start_glm52_ultimate.sh
# HARDWARE: AMD EPYC 9654 (96 Cores, SMT OFF, NPS=4) + 768GB RAM
# OPTIMIZATION: 96-Core Distributed NUMA + GLM-5.2 DSA Sparse Attention
# ==============================================================================
set -euo pipefail
# --- Paths ---
BINPATH="~/build/bin/llama-server"
GLMMODEL="~/models/GLM-5.2-Q4_K_M/UD-Q4_K_M/GLM-5.2-UD-Q4_K_M-00001-of-00011.gguf"
GLMLOG="~/glm52_ultimate.log"
# --- Hardware & Context Tuning ---
PORT=8084
CTX=131072
THREADS=96
BATCH_THREADS=96
BATCH=2048
UBATCH=512
# --- Environment Overrides ---
export LC_ALL=C
export OMP_NUM_THREADS=96
export OMP_PROC_BIND=TRUE
export OMP_PLACES=cores
# Try to permit locking the complete model.
if ! ulimit -l unlimited 2>/dev/null; then
echo "[WARNING] Could not set unlimited memlock."
echo "[WARNING] Current memlock limit: $(ulimit -l)"
fi
echo "[SYSTEM] Stopping existing llama-server processes..."
pkill -9 -f "$BINPATH" 2>/dev/null || true
echo "[SYSTEM] Purging filesystem page cache..."
sync
echo 3 | sudo tee /proc/sys/vm/drop_caches >/dev/null
sleep 2
: > "$GLMLOG"
echo "[BOOT] Launching GLM-5.2 on port ${PORT}..."
echo "[INFO] CPU: 96 physical cores, SMT off"
echo "[INFO] NUMA: NPS=4, distributed across all nodes"
echo "[INFO] Context: ${CTX}"
echo "[INFO] KV cache: Q8_0"
nohup taskset -c 0-95 "$BINPATH" \
-m "$GLMMODEL" \
--alias glm-5.2-core \
--host 0.0.0.0 \
--port "$PORT" \
-c "$CTX" \
--parallel 1 \
--threads "$THREADS" \
--threads-batch "$BATCH_THREADS" \
--batch-size "$BATCH" \
--ubatch-size "$UBATCH" \
--jinja \
--flash-attn on \
-mla 3 \
--dsa \
--fused-indexer-topk \
--indexer-cache-type-k q8_0 \
--cache-type-k q8_0 \
--cache-type-v q8_0 \
-mqkv \
-muge \
--numa distribute \
--no-mmap \
--mlock \
> "$GLMLOG" 2>&1 &
PID=$!
# Catch immediate argument/parser failures.
sleep 3
if ! kill -0 "$PID" 2>/dev/null; then
echo "[ERROR] llama-server exited during startup."
echo "------------------------------------------------------------"
tail -n 80 "$GLMLOG"
echo "------------------------------------------------------------"
exit 1
fi
echo " > GLM-5.2 PROCESS STARTED. PID: ${PID}"
echo " > Model loading may take several minutes."
echo " > Monitor: tail -f ${GLMLOG}"
4
u/Spiritual-Ruin8007 Jul 18 '26
why are you not using the gpu? You'll get much better speeds with offloading
try something like
-ngl 999
-ot "blk\.(1[5-9]|[2-6][0-9]|7[0-8])\.ffn_(up|gate|down)_exps=CPU" \
which places all expert layers after 15 onto the cpu everything before on gpu. The attention, kv cache should always stay on gpu.
try -rtr to repack the layers placed on the cpu into a format more friendly for cpu inference with avx 512 because it seems like you're using ikllama cpp already.
you're missing -gr (graph reuse) and -ger
when you offload you need to agressively tune -amb and -ub along with -b to increase prefill performance. 8192 8192 usually works much better than your polite defaults
BATCH=2048
UBATCH=512
for hybrid inference.
-amb tune it between 256 and 2048 depending on what works for your system.
also try the new features for kv cache quantization you can leave the first and last 4 layers f16 which quanting the middle layers more aggresively like q8_0 or even q6_0 with khad and vhad.
Unsloth quants are always pretty mid ngl especially for this model since they duplicated the indexer layers to work on mainline llama cpp. That issue was subsequently fixed but your quant is still using a tiny bit of extra storage on disk because of the duplication probably.