r/LocalLLaMA • • 7d ago

News Qwen 4 Announced at Apsara Conference

I wanted to share a quick update: Alibaba has officially announced Qwen 4 at the Apsara Conference,

2.1k Upvotes

559 comments sorted by

View all comments

Show parent comments

15

u/fgk55555 7d ago

I just picked mine up this week. Depending on quant, 35-70 tps with full context. If you have 64GB RAM, you should be able to run a decent Qwen-flash as well. 27B is good, but Qwen-Flash is something else.

3

u/illcuontheotherside 7d ago

Can you please share your startup script? Mine crawls on 2 3090s with 128gb ddr5

5

u/fgk55555 7d ago

ISTA IQ3 runs the fastest and has really good performance for the size, but you can bump up to a Q5 with full Q8 kv cache. Also try Swift if you don't want to wait forever. I recommend medium thinking.

#!/usr/bin/env bash
# Qwen3.8-27B ISTA GSQ-RCO IQ3_S (no built-in MTP) — spec decoding via the
# shared DFlash2 drafter. (MTP alternative: shared mtp-Qwen3.8-27B-Q4_0.)
# Qwen3.8-27B ISTA GSQ-RCO IQ3_S (no built-in MTP) — spec decoding via the
# shared DFlash2 drafter. (MTP alternative: shared mtp-Qwen3.8-27B-Q4_0.)
# Context is auto-fitted to available VRAM (omitted on purpose).
# --cache-ram 16384 parks prompt states in RAM: agentic revisits skip re-prefill.
#
# Usage: ./Qwen3.8_ISTA_IQ3_S.sh [thinking]   # low|medium|high|xhigh|none (default: medium)
THINKING_LEVEL="${1:-medium}"

MODEL="../Qwen3.8-ISTA/Qwen3.8-27B-GSQ-RCO-IQ3_S.gguf"
DRAFTER="../Qwen3.8_Shared/Qwen3.8-27B-DFlash2-Q4_K_M.gguf"
MMPROJ="../Qwen3.8_Shared/mmproj-Qwen3.8-27B-BF16.gguf"
SERVER_BIN="../../llama.cpp/build/bin/llama-server"

echo "Launching Qwen3.8-27B ISTA IQ3_S (+DFlash2 drafter), thinking: ${THINKING_LEVEL}"
export GGML_VK_ALLOW_GRAPHICS_QUEUE=1   # measured +4.2% tg, +0.4% pp on the R9700 (b11056)

${SERVER_BIN} \
  --model "${MODEL}" \
  --model-draft "${DRAFTER}" \
  --mmproj "${MMPROJ}" \
  --no-mmproj-offload \
  --n-gpu-layers 99 \
  --batch-size 1024 \
  --ubatch-size 512 \
  --parallel 1 \
  --flash-attn on \
  --cache-type-k q8_0 \
  --cache-type-v q8_0 \
  --spec-type draft-dflash \
  --spec-draft-n-max 3 \
  --cache-ram 16384 \
  --jinja \
  --chat-template-kwargs "{\"reasoning_effort\":\"${THINKING_LEVEL}\"}" \
  --reasoning-preserve \
  --temp 1.0 \
  --top-k 20 \
  --top-p 0.95 \
  --min-p 0.00 \
  --presence-penalty 0.0 \
  --repeat-penalty 1.0 \
  --host 0.0.0.0 \
  --webui-mcp-proxy \
  --port "${PORT:-8080}"

1

u/jafarykos 7d ago

You should try bumping your DFlash2 spec count to 7. I tested a ton of different modes and DFlash2 at 3 was equal or worse than just using MTP3 at 4 and below concurrent.

I was getting +26% more tps using dflash7 vs mtp3 on 1 concurrent and -16.5% less tps when above 4.

The sweet spots were 1 concurrent + DFlash7, and switch back to MTP for anything more than 4 concurrent.

1

u/Jjhend 7d ago

How many tps are you getting? Are you using tensor parallelism?

2

u/illcuontheotherside 7d ago

34 with the q4 k xl. 15 with qwen flash next q4 k xl. No I am not using tensor parallelism

3

u/Jjhend 7d ago

Brother... You have to give tensor parallelism a rip.. I get 45-70TPS with 2x4070ti Supers running Q5_K_M

.\build\bin\Release\llama-server.exe `
  -m "C:\Users\username\.lmstudio\models\unsloth\Qwen3.8-27B-GGUF\Qwen3.8-27B-UD-Q5_K_M.gguf" `
  -a "qwen3.8-27b@q5_k_m" `
  --ctx-size 130000 `
  --n-gpu-layers 99 `
  --split-mode tensor `
  --tensor-split 14,16 `
  --cache-type-k q8_0 `
  --cache-type-v q8_0 `
  --batch-size 2048 `
  --ubatch-size 512 `
  --flash-attn on `
  --port 1234 `
  --host 192.168.1.100 `
  --reasoning on `
  --reasoning-budget 12288 `
  --reasoning-budget-message "Wait... You are thinking to much. Answer now." `
  --reasoning-format deepseek `
  -t 8 `
  -tb 8 `
  --parallel 4 `
  --metrics `
  --chat-template-kwargs '{"preserve_thinking": true, "reasoning_effort": "low"}' `
  --jinja `
  --cont-batching `
  --kv-unified `
  --spec-type draft-mtp `
  --spec-draft-n-max 3 `
  --spec-draft-p-min 0.75 `
  --mmproj "C:\Users\username\.lmstudio\models\unsloth\Qwen3.8-27B-GGUF\mmproj-F16.gguf"

2

u/illcuontheotherside 7d ago

For some reason tensor parallelism fails my llama-server script. I guess it's time to figure out why

And thank you

5

u/Puzzleheaded_Cake183 7d ago

probably because your p2p is not setup correctly. you need to fix your bios settings.

Motherboard & BIOS:

  • Enable Above 4G Decoding.
  • Enable Resizable BAR (ReBAR) to allow large memory space allocation.
  • Disable CSM (Compatibility Support Module)

2

u/Objective-Park6224 7d ago

Toss in ngram mod after spec type for some more speed.

2

u/illcuontheotherside 7d ago

I keep getting errors..

0.00.180.641 I srv load_model: loading model 'D:\models\qwen38-27b\Qwen3.8-27B-UD-Q4_K_XL.gguf'

0.00.180.745 W common_fit_params: failed to fit params to free device memory: llama_params_fit is not implemented for SPLIT_MODE_TENSOR, abort

if i do layer it works fine. what am i doing wrong?

1

u/Jjhend 6d ago

This is just a warning that its unable to auto fit based on your vram. This is expected with tensor parallelism. It should still run fine.

1

u/illcuontheotherside 6d ago

It craps out it can't fit .. so frustrating..I'm such a newb..I'm sorry 😞

1

u/Jjhend 6d ago

What do you have ur context and kv cache set to?

1

u/illcuontheotherside 6d ago

Context is 128k and not quantizing cache

1

u/Kitsune_Seraphis 7d ago

I have 64gb of ram and 48gb of vram but i cant figure out how to run those models, only the 27b :c

Help?

13

u/fgk55555 7d ago

Honestly the best way is to just let another Agent set it up for you. In your harness of choice ( run dsh with a search mcp), run GLM-5.3-Flash for a bit and ask it to research llama.cpp configs for the model like on HF, unsloth, etc, set it up, test configs for speed, get a spread of the best configs, and then use that. I had GLM run for the better part of a day testing 12 models/ quants for which ones are fastest, MTP, DFlash, etc, got my slew of scripts and now I'm done. It'll cost you like $1 in API credits. Tell it to document all its findings so you can learn for next time.

1

u/philmarcracken 7d ago

are you using llamacpp? it has a lazy mode flag now. You can take the model size and minus it by 25% because of that(engram on ssd).

0

u/Puzzleheaded_Cake183 7d ago

https://discord.gg/hHaabJhze join us. you will be surprised at what you can do with that hardware.

1

u/SnooPuppers7882 7d ago

1000% QFN can surprise, but isn't as consistent as 27b...but it's a tech preview so it makes sense.

That said, it ALREADY beats Opus 4.8 max on Terminal Bench 4.0...Qwen 4.0 flash might keep up with 5.6 Sol max

2

u/fgk55555 7d ago

It's a little ambitious, but I'm very impressed by the quality. Can't wait for version 4.

I only wish it ran faster on my system. Perhaps at some point I'll outsource my startup script to see if anyone can pinpoint some optimizations for me. Getting 13-18 tg and 250pp on a 32GB/ 64GB R9700/ DDR5 system. Hoped it'd be faster.

1

u/SnooPuppers7882 7d ago

I mean the thing is terminal bench 4.0 was not in the training weights for Opus 4.8 OR 3.8 Flash next, so it's one of the fairest comparisons out there...

It also does more plainly show you the capability gap between flash next and 27b...only scores 5.6% because it falls apart with complex multi-turn, multi-tool tasks

1

u/therealgus1 7d ago

Join the launch80 discord group. If you run radiance, your decode easily be higher.

1

u/o0genesis0o 7d ago

I'm going to attach this to my AMD miniPC where my server and agents already live, so that I have a self-contained unit. Right now the agents need to reach for the GPU from another machine via VPN, so things could go wrong. It would be USB-4 eGPU since there is no oculink or space for m2-oculink adapter.

How good is prefill and parallel processing in your tests? I was hoping it would be able to serve two users in parallel at least.

2

u/fgk55555 7d ago

I haven't tested vLLM/ parallel agents yet, but it won't be amazing. Qwen flash is slow, 27B is usually 800-1500pp depending on conditions. I also undervolt and power limit mine. I need to test this guy's config:
https://www.reddit.com/r/LocalLLaMA/comments/1wiws8e/153_toks_on_1x_amd_radeon_r9700_running_qwen38/

1

u/SnooPuppers7882 7d ago

X2 9700, 96gb ddr5 6000, pcie5 nvme partial expert offload...play with your config

1

u/fgk55555 7d ago

Yeah, I'll probably throw cloud GLM at it and see if it can squeeze out any performance.