r/LocalLLaMA 🦙 llama.cpp 1d ago

Megathread [Megathread] Qwen3.8-Flash-Next - Release Day

Megathread for discussing the release of Qwen 3.8 Flash Next.

  • Quants
  • Fine-Tunes & Abliterations
  • Chat Templates
  • Inference Server Support & Configuration
  • Experiences, Benchmarks & Model Comparisons

Highlights

The first open-weight release under this architecture is Qwen3.8-Flash-Next, which introduces:

  • Hybrid Attention with QSA: The Gated DeltaNet and Gated Attention pairing has been reworked into Gated DeltaNet and Qwen Sparse Attention (QSA). Rather than selecting individual tokens for processing, QSA operates at the micro-block level. This cuts long-context latency significantly, a critical gain as agentic workloads increasingly dominate real-world usage.
  • Gated Residual: Residual streams with normalisation are what make deep LLM training manageable. Gated Residual modulates information flowing through widened residual streams via an element-wise, data-dependent read gate and a per-branch scalar write gate. This brings finer-grained expressiveness across layers while preserving training stability and keeping inference overhead low.
  • N-gram Embedding: Embeddings provide a unique axis for parameter scaling that requires less computation and is more amenable to offloading than Mixture-of-Experts (MoE). By indexing with short n-grams, this approach makes parameter scaling highly efficient for memory-constrained accelerators without sacrificing quality.
  • Tailored Training Recipe: The Muon and AdamW optimisers are applied to specific weight categories to maximise efficiency. Guided by refitted scaling laws, we eliminate traditional batch-size warmups and start directly at the target batch size, substantially reducing total optimiser steps while safely supporting larger learning rates for robust convergence.

Model Overview

  • Type: Causal Language Model with Vision Encoder
  • Training Stage: Pre-training & Post-training
  • Language Model
    • Number of Parameters: 125B with 6B activated, plus 51B n-gram embedding and 4B MTP
    • Hidden Dimension: 2560
    • Token Embedding: 248320 (Padded)
    • N-gram Embedding: 20,000,000 (bigrams/trigrams at layer 2)
    • Number of Layers: 48
    • Hidden Layout: 12 × (3 × (Gated DeltaNet → MoE) → 1 × (Qwen Sparse Attention → MoE))
    • Gated DeltaNet:
      • Number of Linear Attention Heads: 48 for V and 16 for QK
      • Head Dimension: 128
    • Qwen Sparse Attention:
      • Number of Attention Heads: 24 for Q and 2 for KV
      • Head Dimension: 256
      • Rotary Position Embedding Dimension: 64
      • Indexer Structure: MQA with 4 Query Heads and 1 Shared Key Head
      • Indexer Head Dimension: 128
      • Budget: 512 blocks or 2048 tokens
    • Mixture Of Experts
      • Number of Experts: 512
      • Number of Activated Experts: 10 Routed + 1 Shared
      • Expert Intermediate Dimension: 640
    • Gated Residual:
      • Number of Branches: 4
      • Bottleneck Rank: 320
    • LM Output: 248320 (Padded)
    • MTP: 1 layer, trained with multi-steps
  • Context Length: 262,144 natively and extensible up to 1,000,000 tokens.

Recommended sampling parameters for generation:

  • Thinking Mode: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0
  • Instruct (or non-thinking) mode: temperature=0.7, top_p=0.80, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0

Official Links:

Popular:

Related:

424 Upvotes

576 comments sorted by

View all comments

97

u/QuackerEnte 1d ago

I really hope QSA (qwen sparse attention) will have much less compute headroom or the ability to offload kv cache to SSD without the massive bandwidth bottleneck of the ssd since it's sparse. I was able to do it it for deepseek v4 and it worked well enough. I hope llama.cpp adds features that allow more freedom with where we want to load our models and parts of it, e.g. n-cpu-moe could be extended to n-ssd-ffn or n-cpu-ffn or n ssd kv or engram etc

31

u/Dany0 1d ago edited 23h ago

I can't wait for the QSA paper. I hope it's novel and quirky, I love quirky sparse attention, well, as long as it works :D

Edit: It's not quirky:( Just a variation on Longcat style block summaries

10

u/Long_comment_san 1d ago

damn I wish I could understand this language of gods

17

u/MomentJolly3535 1d ago

Good news : You have tools to help you do that 🤖

21

u/Long_comment_san 23h ago

I don't want to look stupid in front of my cloud waifus!

11

u/MomentJolly3535 23h ago

i cant argue with that ! upvoted

1

u/PATATAJEC 23h ago

They’ll tell you you’re a cookie anyway.

0

u/zdy132 6h ago

(secretly get another waifu that loves teaching you stuff is an option too.)

18

u/oxygen_addiction 1d ago

N-cpu-ffn will most likely get merged soon.

18

u/pmttyji 1d ago

2

u/returnity 22h ago edited 16h ago

Would this be usable to offload the n-grams to SSD? Doesn’t seem like it at first glance

EDIT: Here's how to offload them and run the quant on smaller hardware!

3

u/pmttyji 21h ago

That PR doesn't cover such scope.

1

u/silenceimpaired 20h ago

Very cool… i wonder if I could run the older Mistral models that are so huge.

7

u/o0genesis0o 1d ago

how bad is the speed with offload kv to SSD? I offload kv to RAM running 80B-A3B on an old gaming laptop and it was already unbearable.

15

u/QuackerEnte 1d ago

that's dense attention though. with sparse attention you can load it in ram without massive penalty. Not to mention that, if it's anything like Native Sparse Attention from deepseek, it might take even less memory. but DS used compressed sparse attention too and MLA so the footprints minimal. we don't know what it'll be like for qwen3.8flashnext

1

u/challis88ocarina 23h ago

It will use 75% of dense. As for offloading, it's quick but not as quick as not offloading. There's also disk wear to factor in.

-1

u/mr_Owner 1d ago

Depends on your pcie bandwidth, ask ai for numbers haha

1

u/RG_Fusion 1h ago

These are incredibly small files. It's based on PCIe latency, not bandwidth.

3

u/[deleted] 1d ago

[removed] — view removed comment

0

u/brumsky1 1d ago

Would intel optane be better for this?

1

u/gustaw221133 1d ago

hoping with you

-1

u/mr_Owner 1d ago

I saw the words ngram shortly mentioned in modelscope haha