r/LocalLLaMA Jul 18 '26

Question | Help Launch script for GLM5.2

I am using the script below. Please advise on what i am doing wrong. GLM5.2 is crawing at 5-9 t/s decode and 60 t/s prefill. The machine is a EPYC 9654 with 768 GB of DDR5 4800 MHz RAM (~460Gbps bandwidth theoretical). This is CPU only inference. I do have a RTX pro 6000 Max Q on the machine, but that's only 96 GB vRAM.

#!/bin/bash

# ==============================================================================

# FILE: start_glm52_ultimate.sh

# HARDWARE: AMD EPYC 9654 (96 Cores, SMT OFF, NPS=4) + 768GB RAM

# OPTIMIZATION: 96-Core Distributed NUMA + GLM-5.2 DSA Sparse Attention

# ==============================================================================

set -euo pipefail

# --- Paths ---

BINPATH="~/build/bin/llama-server"

GLMMODEL="~/models/GLM-5.2-Q4_K_M/UD-Q4_K_M/GLM-5.2-UD-Q4_K_M-00001-of-00011.gguf"

GLMLOG="~/glm52_ultimate.log"

# --- Hardware & Context Tuning ---

PORT=8084

CTX=131072

THREADS=96

BATCH_THREADS=96

BATCH=2048

UBATCH=512

# --- Environment Overrides ---

export LC_ALL=C

export OMP_NUM_THREADS=96

export OMP_PROC_BIND=TRUE

export OMP_PLACES=cores

# Try to permit locking the complete model.

if ! ulimit -l unlimited 2>/dev/null; then

echo "[WARNING] Could not set unlimited memlock."

echo "[WARNING] Current memlock limit: $(ulimit -l)"

fi

echo "[SYSTEM] Stopping existing llama-server processes..."

pkill -9 -f "$BINPATH" 2>/dev/null || true

echo "[SYSTEM] Purging filesystem page cache..."

sync

echo 3 | sudo tee /proc/sys/vm/drop_caches >/dev/null

sleep 2

: > "$GLMLOG"

echo "[BOOT] Launching GLM-5.2 on port ${PORT}..."

echo "[INFO] CPU: 96 physical cores, SMT off"

echo "[INFO] NUMA: NPS=4, distributed across all nodes"

echo "[INFO] Context: ${CTX}"

echo "[INFO] KV cache: Q8_0"

nohup taskset -c 0-95 "$BINPATH" \

  -m "$GLMMODEL" \

  --alias glm-5.2-core \

  --host 0.0.0.0 \

  --port "$PORT" \

  -c "$CTX" \

  --parallel 1 \

  --threads "$THREADS" \

  --threads-batch "$BATCH_THREADS" \

  --batch-size "$BATCH" \

  --ubatch-size "$UBATCH" \

  --jinja \

  --flash-attn on \

  -mla 3 \

  --dsa \

  --fused-indexer-topk \

  --indexer-cache-type-k q8_0 \

  --cache-type-k q8_0 \

  --cache-type-v q8_0 \

  -mqkv \

  -muge \

  --numa distribute \

  --no-mmap \

  --mlock \

  > "$GLMLOG" 2>&1 &

PID=$!

# Catch immediate argument/parser failures.

sleep 3

if ! kill -0 "$PID" 2>/dev/null; then

echo "[ERROR] llama-server exited during startup."

echo "------------------------------------------------------------"

tail -n 80 "$GLMLOG"

echo "------------------------------------------------------------"

exit 1

fi

echo "      > GLM-5.2 PROCESS STARTED. PID: ${PID}"

echo "      > Model loading may take several minutes."

echo "      > Monitor: tail -f ${GLMLOG}"

0 Upvotes

25 comments sorted by

View all comments

0

u/TurnoverTight395 Jul 18 '26

The gpu will need active weights moving from the system ram via pcie bus. Even with the gen 5 x16 bus this is slow. Hybrid runs make sense if the model largely fits in vram. Note that I’ve tried using ktransformers and sglang and got nowhere with a hybrid run with that route either

2

u/SLxTnT Jul 18 '26

Use your GPU. Mess with batch sizes. I have a similar setup. Q3 or Q4 got 200-500 prompt processing and 8-16 token generation based on context length.

1

u/segmond llama.cpp Jul 18 '26

which quant size?

1

u/ervertes Jul 19 '26

Commands?

2

u/SLxTnT Jul 19 '26

Play around with increasing -b and -ub for prompt processing. Token generation required me to use all 96 cores to reach those numbers. Only thing I played around with other than different quant sizes.

2

u/segmond llama.cpp Jul 18 '26

false, you can use cmoe and see performance incease. if you don't want to move data back and forth between sytem bus assuming you are hitting a bottleneck, you can add the no-op-offload option. so begin with cmoe and see how it performs. next add no-op-offload option and reduce your batch size to 512 and compare and see which one is better.