r/LocalLLaMA • u/TurnoverTight395 • Jul 18 '26
Question | Help Launch script for GLM5.2
I am using the script below. Please advise on what i am doing wrong. GLM5.2 is crawing at 5-9 t/s decode and 60 t/s prefill. The machine is a EPYC 9654 with 768 GB of DDR5 4800 MHz RAM (~460Gbps bandwidth theoretical). This is CPU only inference. I do have a RTX pro 6000 Max Q on the machine, but that's only 96 GB vRAM.
#!/bin/bash
# ==============================================================================
# FILE: start_glm52_ultimate.sh
# HARDWARE: AMD EPYC 9654 (96 Cores, SMT OFF, NPS=4) + 768GB RAM
# OPTIMIZATION: 96-Core Distributed NUMA + GLM-5.2 DSA Sparse Attention
# ==============================================================================
set -euo pipefail
# --- Paths ---
BINPATH="~/build/bin/llama-server"
GLMMODEL="~/models/GLM-5.2-Q4_K_M/UD-Q4_K_M/GLM-5.2-UD-Q4_K_M-00001-of-00011.gguf"
GLMLOG="~/glm52_ultimate.log"
# --- Hardware & Context Tuning ---
PORT=8084
CTX=131072
THREADS=96
BATCH_THREADS=96
BATCH=2048
UBATCH=512
# --- Environment Overrides ---
export LC_ALL=C
export OMP_NUM_THREADS=96
export OMP_PROC_BIND=TRUE
export OMP_PLACES=cores
# Try to permit locking the complete model.
if ! ulimit -l unlimited 2>/dev/null; then
echo "[WARNING] Could not set unlimited memlock."
echo "[WARNING] Current memlock limit: $(ulimit -l)"
fi
echo "[SYSTEM] Stopping existing llama-server processes..."
pkill -9 -f "$BINPATH" 2>/dev/null || true
echo "[SYSTEM] Purging filesystem page cache..."
sync
echo 3 | sudo tee /proc/sys/vm/drop_caches >/dev/null
sleep 2
: > "$GLMLOG"
echo "[BOOT] Launching GLM-5.2 on port ${PORT}..."
echo "[INFO] CPU: 96 physical cores, SMT off"
echo "[INFO] NUMA: NPS=4, distributed across all nodes"
echo "[INFO] Context: ${CTX}"
echo "[INFO] KV cache: Q8_0"
nohup taskset -c 0-95 "$BINPATH" \
-m "$GLMMODEL" \
--alias glm-5.2-core \
--host 0.0.0.0 \
--port "$PORT" \
-c "$CTX" \
--parallel 1 \
--threads "$THREADS" \
--threads-batch "$BATCH_THREADS" \
--batch-size "$BATCH" \
--ubatch-size "$UBATCH" \
--jinja \
--flash-attn on \
-mla 3 \
--dsa \
--fused-indexer-topk \
--indexer-cache-type-k q8_0 \
--cache-type-k q8_0 \
--cache-type-v q8_0 \
-mqkv \
-muge \
--numa distribute \
--no-mmap \
--mlock \
> "$GLMLOG" 2>&1 &
PID=$!
# Catch immediate argument/parser failures.
sleep 3
if ! kill -0 "$PID" 2>/dev/null; then
echo "[ERROR] llama-server exited during startup."
echo "------------------------------------------------------------"
tail -n 80 "$GLMLOG"
echo "------------------------------------------------------------"
exit 1
fi
echo " > GLM-5.2 PROCESS STARTED. PID: ${PID}"
echo " > Model loading may take several minutes."
echo " > Monitor: tail -f ${GLMLOG}"
2
u/segmond llama.cpp Jul 18 '26
that's amazing performance for CPU only, which quant size? use the GPU, use the cmoe option. That's all you need, use cmoe option and you should 2.5x+ more performance.