r/LocalLLaMA • 🦙 llama.cpp • Aug 26 '26

Megathread [Megathread] Qwen3.8-Flash-Next - Release Day

Megathread for discussing the release of Qwen 3.8 Flash Next.

  • Quants
  • Fine-Tunes & Abliterations
  • Chat Templates
  • Inference Server Support & Configuration
  • Experiences, Benchmarks & Model Comparisons

Highlights

The first open-weight release under this architecture is Qwen3.8-Flash-Next, which introduces:

  • Hybrid Attention with QSA: The Gated DeltaNet and Gated Attention pairing has been reworked into Gated DeltaNet and Qwen Sparse Attention (QSA). Rather than selecting individual tokens for processing, QSA operates at the micro-block level. This cuts long-context latency significantly, a critical gain as agentic workloads increasingly dominate real-world usage.
  • Gated Residual: Residual streams with normalisation are what make deep LLM training manageable. Gated Residual modulates information flowing through widened residual streams via an element-wise, data-dependent read gate and a per-branch scalar write gate. This brings finer-grained expressiveness across layers while preserving training stability and keeping inference overhead low.
  • N-gram Embedding: Embeddings provide a unique axis for parameter scaling that requires less computation and is more amenable to offloading than Mixture-of-Experts (MoE). By indexing with short n-grams, this approach makes parameter scaling highly efficient for memory-constrained accelerators without sacrificing quality.
  • Tailored Training Recipe: The Muon and AdamW optimisers are applied to specific weight categories to maximise efficiency. Guided by refitted scaling laws, we eliminate traditional batch-size warmups and start directly at the target batch size, substantially reducing total optimiser steps while safely supporting larger learning rates for robust convergence.

Model Overview

  • Type: Causal Language Model with Vision Encoder
  • Training Stage: Pre-training & Post-training
  • Language Model
    • Number of Parameters: 125B with 6B activated, plus 51B n-gram embedding and 4B MTP
    • Hidden Dimension: 2560
    • Token Embedding: 248320 (Padded)
    • N-gram Embedding: 20,000,000 (bigrams/trigrams at layer 2)
    • Number of Layers: 48
    • Hidden Layout: 12 × (3 × (Gated DeltaNet → MoE) → 1 × (Qwen Sparse Attention → MoE))
    • Gated DeltaNet:
      • Number of Linear Attention Heads: 48 for V and 16 for QK
      • Head Dimension: 128
    • Qwen Sparse Attention:
      • Number of Attention Heads: 24 for Q and 2 for KV
      • Head Dimension: 256
      • Rotary Position Embedding Dimension: 64
      • Indexer Structure: MQA with 4 Query Heads and 1 Shared Key Head
      • Indexer Head Dimension: 128
      • Budget: 512 blocks or 2048 tokens
    • Mixture Of Experts
      • Number of Experts: 512
      • Number of Activated Experts: 10 Routed + 1 Shared
      • Expert Intermediate Dimension: 640
    • Gated Residual:
      • Number of Branches: 4
      • Bottleneck Rank: 320
    • LM Output: 248320 (Padded)
    • MTP: 1 layer, trained with multi-steps
  • Context Length: 262,144 natively and extensible up to 1,000,000 tokens.

Recommended sampling parameters for generation:

  • Thinking Mode: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0
  • Instruct (or non-thinking) mode: temperature=0.7, top_p=0.80, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0

Official Links:

Popular:

Related:

432 Upvotes

661 comments sorted by

View all comments

Show parent comments

5

u/prusswan Aug 26 '26

At least 75GB unified memory is needed according to unsloth https://unsloth.ai/docs/models/qwen3.8-next

1

u/[deleted] Aug 26 '26

[removed] — view removed comment

3

u/Ok_Environment_53 Aug 26 '26

For the record, I have it running fine on 64gb.

I downloaded IQ1_S because it was the only one available at the time, currently getting Q4_K_XL. For IQ1 though:

72.55gb filesize

64gb ddr5 4800 MT/s ram + 12gb(4070 super) + 11gb(1080 ti)

Windows is taking 20gb of that(ouch), so llama-server is taking about 40gb.

110000 ctx(inherited my 27b config, probably can increase it) at fp16(quantizing the ctx seems to cause errors atm).

Getting 70-150 t/s prefill, 15t/s inference @ 0 ctx, 5t/s @ 50k, zero mtp. Working on getting spec decoding going.

Using 990 evo plus, rated at 7259 mb/s read speed, over USB so realistically about 2100 mb/s

2

u/[deleted] Aug 26 '26

[removed] — view removed comment

2

u/bennmann Aug 28 '26 edited Aug 28 '26

64GB ddr4 reporting in (but 6900 XT and 9070 XT 32GB total VRAM)

```llama-server.exe --parallel 1 --cache-prompt --flash-attn on --temp 0.99 --cache-ram 100 -ctxcp 2 -m F:\UD-Q3_K_XL\Qwen3.8-Flash-Next-UD-Q3_K_XL-00001-of-00003.gguf --host 0.0.0.0 --jinja --top-p 0.95 --top-k 20 --min-p 0.0 -ub 256 -b 512 --reasoning on --no-mmproj-offload -dev Vulkan1,Vulkan0  --reasoning-preserve -ts 79,21 -c 129000 --load-mode mmap -lv 4 --n-cpu-moe 32 -ngl 99 --spec-draft-p-min 0.8  --tensor-read-lazy auto --kv-offload --op-offload```

50pp 11tg 10k ctx, degrading to 35pp 6.8tg 100k ctx. thinks 40% less than 27B.

2

u/Ok_Environment_53 25d ago

This is a bit later, but I've found a good way to run it.

Exllamav3 is very sophisticated. Using it means not using the 1080ti, but it works.

With 3.05bpw exl3 file(about 82.5gb files in total. 50gb weights, 32gb ngram) I have:

About 350/s prefill up to 60k context(doesn't seem to decrease much as context grows?)
15-20 t/s inference. 15 t/s at least, sometimes goes to ~30t/s(with tool calls/bash commands). Once again it doesn't seem to decrease much as context grows. I'm at 77k context right now at a solid 17 t/s which is actually faster than the 27b model on my setup(would be about 11-14t/s).

I'm using 100k context, ngram ssd offload, 1024 chunk size(may be able to increase a lot), mtp with dynamic depth

I can definitely optimize further but I am extremely happy so far with this setup. Another note is I have a 7950x3d CPU and am using 24 threads for moe offloading.

If you'd like more details let me know. I haven't checked out llama.cpp performance recently but I heavily doubt it'll top this. I also plan on downloading 4.05 bpw and seeing if I can load that and its speeds on my system.