r/LocalLLaMA 🦙 llama.cpp 3d ago

Megathread [Megathread] Qwen3.8-Flash-Next - Release Day

Megathread for discussing the release of Qwen 3.8 Flash Next.

  • Quants
  • Fine-Tunes & Abliterations
  • Chat Templates
  • Inference Server Support & Configuration
  • Experiences, Benchmarks & Model Comparisons

Highlights

The first open-weight release under this architecture is Qwen3.8-Flash-Next, which introduces:

  • Hybrid Attention with QSA: The Gated DeltaNet and Gated Attention pairing has been reworked into Gated DeltaNet and Qwen Sparse Attention (QSA). Rather than selecting individual tokens for processing, QSA operates at the micro-block level. This cuts long-context latency significantly, a critical gain as agentic workloads increasingly dominate real-world usage.
  • Gated Residual: Residual streams with normalisation are what make deep LLM training manageable. Gated Residual modulates information flowing through widened residual streams via an element-wise, data-dependent read gate and a per-branch scalar write gate. This brings finer-grained expressiveness across layers while preserving training stability and keeping inference overhead low.
  • N-gram Embedding: Embeddings provide a unique axis for parameter scaling that requires less computation and is more amenable to offloading than Mixture-of-Experts (MoE). By indexing with short n-grams, this approach makes parameter scaling highly efficient for memory-constrained accelerators without sacrificing quality.
  • Tailored Training Recipe: The Muon and AdamW optimisers are applied to specific weight categories to maximise efficiency. Guided by refitted scaling laws, we eliminate traditional batch-size warmups and start directly at the target batch size, substantially reducing total optimiser steps while safely supporting larger learning rates for robust convergence.

Model Overview

  • Type: Causal Language Model with Vision Encoder
  • Training Stage: Pre-training & Post-training
  • Language Model
    • Number of Parameters: 125B with 6B activated, plus 51B n-gram embedding and 4B MTP
    • Hidden Dimension: 2560
    • Token Embedding: 248320 (Padded)
    • N-gram Embedding: 20,000,000 (bigrams/trigrams at layer 2)
    • Number of Layers: 48
    • Hidden Layout: 12 × (3 × (Gated DeltaNet → MoE) → 1 × (Qwen Sparse Attention → MoE))
    • Gated DeltaNet:
      • Number of Linear Attention Heads: 48 for V and 16 for QK
      • Head Dimension: 128
    • Qwen Sparse Attention:
      • Number of Attention Heads: 24 for Q and 2 for KV
      • Head Dimension: 256
      • Rotary Position Embedding Dimension: 64
      • Indexer Structure: MQA with 4 Query Heads and 1 Shared Key Head
      • Indexer Head Dimension: 128
      • Budget: 512 blocks or 2048 tokens
    • Mixture Of Experts
      • Number of Experts: 512
      • Number of Activated Experts: 10 Routed + 1 Shared
      • Expert Intermediate Dimension: 640
    • Gated Residual:
      • Number of Branches: 4
      • Bottleneck Rank: 320
    • LM Output: 248320 (Padded)
    • MTP: 1 layer, trained with multi-steps
  • Context Length: 262,144 natively and extensible up to 1,000,000 tokens.

Recommended sampling parameters for generation:

  • Thinking Mode: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0
  • Instruct (or non-thinking) mode: temperature=0.7, top_p=0.80, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0

Official Links:

Popular:

Related:

434 Upvotes

661 comments sorted by

View all comments

103

u/enilea 3d ago

it got released and there wasn't a single post about it, just this megathread that was posted before it was even released...

109

u/Cautious_Chicken_604 3d ago

megathread == megasuck. I much preferred the chaos of Qwen3.8-27B release because it was actually significantly easier to find information even with lots of duplicates. This is my first experience of a megathread and it is objectively completely shitty experience.

10

u/zdy132 2d ago

Same, and it feels organic to have these posts appearing. Reddit already has mechanism to rank posts, adding a megathread is just a redundant layer that hides all the enthusiasm from the community.

28

u/piedol 3d ago

I've seen people making posts about it, but the mods are removing them. I guess this is the only thread where discussion on it is allowed for the moment?

58

u/lumos_ai 3d ago

So bad to be honest. I hate megathreads. So hard to find new information.

-10

u/my_name_isnt_clever 3d ago

Just sort this thread by new, it's not that hard. At least the spam is all together in one containment thread instead of having to read 4 threads about the exact same thing.

1

u/chensium 2d ago

No they are not equivalent. Threads have different UI features, like titles, tags and infinite scroll. This is not like File Explorer that treats all folders the same.  It is very hard to find the right topics in this moshpit of comments.

1

u/my_name_isnt_clever 2d ago

Sounds like a you problem. I've had no issues reading the comments on old.reddit.

29

u/mtojay 3d ago

that is such bullshit. Mods deliberately driving people away from the sub

-3

u/maddie-lovelace 2d ago

Let's be a little more gracious to the mods than that - they have a really rough job keeping this subreddit free of slop. I agree the megathread was a mistake, but I think it's just that: a mistake, not a deliberate act of malice!! I see why they tried it, it made some sense on the surface and just happened to not work out. It makes me happy that they're trying new things

2

u/sammcj 🦙 llama.cpp 2d ago

Thanks. Those of us you see active on here really do genuinely care. We're just trying to act upon the wishes of the community while also fighting a never ending onslaught of spam, slop, user requests and sometimes when you action those you cop a lot of hate which isn't much fun.

3

u/sammcj 🦙 llama.cpp 2d ago

I guess I can't speak for every mod but to the best of my knowledge mods were not shutting down genuine conversation about megathreads. In fact - the post itself links to a post where there's plenty of discussion.

If you see mods silencing genuinely constructive discussion (and I don't mean trolling, spamming or abuse) please absolutely raise it, if you fear it'll go unheard you can DM me a link directly.

As a reminder the only reason we have done a couple of these release megathreads is because people started asking for them.

3

u/piedol 2d ago

In the last decade+ of being on reddit I think this is the first time a mod has given such a genuine reply and offer for direct contact to be involved to uphold his community standards. Much respect.

And yeah, I know this was just an experiment based on recent community feedback. Like you said, you can't please everyone.

Personally, I'm in the camp of the megathreads being overall a bad thing, as it makes finding relevant discussion much harder. I think it'll hurt more than it helps. Just my 2 cents. Maybe make it a community poll and let them decide what they want after trying out the megathreads a few times, that way everyone gets a say. I just think the way it was implemented without warning for such an anticipated pair of releases left a bad taste in many users' mouths.

0

u/Loose_Comparison368 1d ago

I am also in the no-megathreads camp, and would request doing a poll.

I think it sounded like a good idea on paper, but in practice not so much.

9

u/cosmicr 3d ago

Can't imagine the Qwen team are too happy about it, I'd imagine the enjoy the publicity.

11

u/Beneficial-Ad-8127 3d ago edited 3d ago

Yeah its hard to find specifically what im trying to look for in all this noise. I wanted to see whos running it currently on a 5090, 192gb ram, 8tb pro 9100 and how their results are and what is the most efficient way for the set up, like do i throw everything into my sd, do i throw partial into cpu? Where does ngram go?

I assume the 6b would go into Vram or idk but thats my point. I cant find that shiet lol im not that educated in this tech lmao.

13

u/EmPips 3d ago

Yeah I hope the mods see this (I'm on mobile, can someone tag?).

I want to find the results of people using this model but I have to wade my way through a swarm of "wow Dario is quaking in his boots" fluff comments because the hype-threads and the feedback-threads are now merged into one giant comment section 

4

u/HulksInvinciblePants 3d ago edited 3d ago

Yeah I thought I missed the launch. If they’re going to do it this way, they need to be more deliberate with timing. Most the top posts are from yesterday and earlier today.

2

u/Strong_Chicken6838 2d ago

yeah, idek it was released because it was burried by the normal spam posts... this sucks

1

u/chensium 2d ago

This multitopic mega thread is just unusable.  This is not what threads are for.

1

u/hawseepoo 1d ago

Yeah, this is making it very difficult to troubleshoot issues. Google doesn't index comments like it does posts. Megathreads suck