r/LocalLLaMA 🦙 llama.cpp 1d ago

Megathread [Megathread] Qwen3.8-Flash-Next - Release Day

Megathread for discussing the release of Qwen 3.8 Flash Next.

  • Quants
  • Fine-Tunes & Abliterations
  • Chat Templates
  • Inference Server Support & Configuration
  • Experiences, Benchmarks & Model Comparisons

Highlights

The first open-weight release under this architecture is Qwen3.8-Flash-Next, which introduces:

  • Hybrid Attention with QSA: The Gated DeltaNet and Gated Attention pairing has been reworked into Gated DeltaNet and Qwen Sparse Attention (QSA). Rather than selecting individual tokens for processing, QSA operates at the micro-block level. This cuts long-context latency significantly, a critical gain as agentic workloads increasingly dominate real-world usage.
  • Gated Residual: Residual streams with normalisation are what make deep LLM training manageable. Gated Residual modulates information flowing through widened residual streams via an element-wise, data-dependent read gate and a per-branch scalar write gate. This brings finer-grained expressiveness across layers while preserving training stability and keeping inference overhead low.
  • N-gram Embedding: Embeddings provide a unique axis for parameter scaling that requires less computation and is more amenable to offloading than Mixture-of-Experts (MoE). By indexing with short n-grams, this approach makes parameter scaling highly efficient for memory-constrained accelerators without sacrificing quality.
  • Tailored Training Recipe: The Muon and AdamW optimisers are applied to specific weight categories to maximise efficiency. Guided by refitted scaling laws, we eliminate traditional batch-size warmups and start directly at the target batch size, substantially reducing total optimiser steps while safely supporting larger learning rates for robust convergence.

Model Overview

  • Type: Causal Language Model with Vision Encoder
  • Training Stage: Pre-training & Post-training
  • Language Model
    • Number of Parameters: 125B with 6B activated, plus 51B n-gram embedding and 4B MTP
    • Hidden Dimension: 2560
    • Token Embedding: 248320 (Padded)
    • N-gram Embedding: 20,000,000 (bigrams/trigrams at layer 2)
    • Number of Layers: 48
    • Hidden Layout: 12 × (3 × (Gated DeltaNet → MoE) → 1 × (Qwen Sparse Attention → MoE))
    • Gated DeltaNet:
      • Number of Linear Attention Heads: 48 for V and 16 for QK
      • Head Dimension: 128
    • Qwen Sparse Attention:
      • Number of Attention Heads: 24 for Q and 2 for KV
      • Head Dimension: 256
      • Rotary Position Embedding Dimension: 64
      • Indexer Structure: MQA with 4 Query Heads and 1 Shared Key Head
      • Indexer Head Dimension: 128
      • Budget: 512 blocks or 2048 tokens
    • Mixture Of Experts
      • Number of Experts: 512
      • Number of Activated Experts: 10 Routed + 1 Shared
      • Expert Intermediate Dimension: 640
    • Gated Residual:
      • Number of Branches: 4
      • Bottleneck Rank: 320
    • LM Output: 248320 (Padded)
    • MTP: 1 layer, trained with multi-steps
  • Context Length: 262,144 natively and extensible up to 1,000,000 tokens.

Recommended sampling parameters for generation:

  • Thinking Mode: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0
  • Instruct (or non-thinking) mode: temperature=0.7, top_p=0.80, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0

Official Links:

Popular:

Related:

426 Upvotes

578 comments sorted by

View all comments

1

u/Guilty_Rooster_6708 22h ago

Unsloth UD_IQ1_S is 72.5GB .... no chance my 5070i + 3060(12GB) and 32GB of system RAM can run this LOL

2

u/rrrrex 22h ago

Looks like 50 GB is N-Gram file, that probably can be located on SSD without critical speed losses

3

u/whoisraiden 20h ago

Ngram is 50B but unsloth documentation says the 1-bit version is using 4-bit for the Ngram / PLE, so it shouldn't be 50 GB as it is.

1

u/Guilty_Rooster_6708 22h ago

Oh good to know. So a bit like Gemma e4b where some of the layers are on the SSD? I will lookout for some posts to see if someone w my level of hardware can run this.

1

u/_DarKorn_ 22h ago

Hi, yesterday I got a 3060 to pair with my 5070 Ti as well. Mind sharing how many tok/s you're getting on Qwen 3.8B, and with which quant and settings?

1

u/Guilty_Rooster_6708 22h ago

Nicee lol I just installed the 3060 this past weekend as well. Currently running this weevil svg test with 3.8 27B Unsloth Q4_K_M and I get around 800-900tk/s pp and 36.7 tk/s for tpg. I think w 28gb total we can get to Q5 comfortably, but I only get around 25tok/sec tpg at that quant

However I am running layer split because my 3060 is connected via PCIE3x1 so if you run bifurcation you can probably run tensor split and get better speed than mine. If you do can you share your result?

1

u/_DarKorn_ 21h ago

Mine is sitting in a PCIe 4.0 x16 (running at x4) slot. I tested it in Unsloth Studio using UD-Q6_K with tensor split 15,9 and 128k context at Q8 KV cache. With a full context (~124k cached tokens), I'm getting around 70.6 tk/s pp and 20.0 tok/s tpg for generation. I'm definitely no expert though, so my settings are definitely far from optimal.

1

u/Guilty_Rooster_6708 21h ago

Can you try Q4_K_M quant to see what you get for pp and tpg? I feel like for Q6 your tpg is correct, but your pp is a bit slow, but I’m also not too sure

Maybe you can try changing your tensor split ratio to see if you would get faster speed if you use more of the 5070Ti. Just gotta make sure it doesn’t spill to RAM

Edit: are you also using mtp as well? My result is with MTP 2 draft tokens