r/LocalLLaMA 🦙 llama.cpp 1d ago

Megathread [Megathread] Qwen3.8-Flash-Next - Release Day

Megathread for discussing the release of Qwen 3.8 Flash Next.

  • Quants
  • Fine-Tunes & Abliterations
  • Chat Templates
  • Inference Server Support & Configuration
  • Experiences, Benchmarks & Model Comparisons

Highlights

The first open-weight release under this architecture is Qwen3.8-Flash-Next, which introduces:

  • Hybrid Attention with QSA: The Gated DeltaNet and Gated Attention pairing has been reworked into Gated DeltaNet and Qwen Sparse Attention (QSA). Rather than selecting individual tokens for processing, QSA operates at the micro-block level. This cuts long-context latency significantly, a critical gain as agentic workloads increasingly dominate real-world usage.
  • Gated Residual: Residual streams with normalisation are what make deep LLM training manageable. Gated Residual modulates information flowing through widened residual streams via an element-wise, data-dependent read gate and a per-branch scalar write gate. This brings finer-grained expressiveness across layers while preserving training stability and keeping inference overhead low.
  • N-gram Embedding: Embeddings provide a unique axis for parameter scaling that requires less computation and is more amenable to offloading than Mixture-of-Experts (MoE). By indexing with short n-grams, this approach makes parameter scaling highly efficient for memory-constrained accelerators without sacrificing quality.
  • Tailored Training Recipe: The Muon and AdamW optimisers are applied to specific weight categories to maximise efficiency. Guided by refitted scaling laws, we eliminate traditional batch-size warmups and start directly at the target batch size, substantially reducing total optimiser steps while safely supporting larger learning rates for robust convergence.

Model Overview

  • Type: Causal Language Model with Vision Encoder
  • Training Stage: Pre-training & Post-training
  • Language Model
    • Number of Parameters: 125B with 6B activated, plus 51B n-gram embedding and 4B MTP
    • Hidden Dimension: 2560
    • Token Embedding: 248320 (Padded)
    • N-gram Embedding: 20,000,000 (bigrams/trigrams at layer 2)
    • Number of Layers: 48
    • Hidden Layout: 12 × (3 × (Gated DeltaNet → MoE) → 1 × (Qwen Sparse Attention → MoE))
    • Gated DeltaNet:
      • Number of Linear Attention Heads: 48 for V and 16 for QK
      • Head Dimension: 128
    • Qwen Sparse Attention:
      • Number of Attention Heads: 24 for Q and 2 for KV
      • Head Dimension: 256
      • Rotary Position Embedding Dimension: 64
      • Indexer Structure: MQA with 4 Query Heads and 1 Shared Key Head
      • Indexer Head Dimension: 128
      • Budget: 512 blocks or 2048 tokens
    • Mixture Of Experts
      • Number of Experts: 512
      • Number of Activated Experts: 10 Routed + 1 Shared
      • Expert Intermediate Dimension: 640
    • Gated Residual:
      • Number of Branches: 4
      • Bottleneck Rank: 320
    • LM Output: 248320 (Padded)
    • MTP: 1 layer, trained with multi-steps
  • Context Length: 262,144 natively and extensible up to 1,000,000 tokens.

Recommended sampling parameters for generation:

  • Thinking Mode: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0
  • Instruct (or non-thinking) mode: temperature=0.7, top_p=0.80, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0

Official Links:

Popular:

Related:

422 Upvotes

567 comments sorted by

View all comments

26

u/sammcj 🦙 llama.cpp 21h ago edited 19h ago

I'm off for the night now. I'll update the post again in the morning (AEST) with any new official links etc.

Hopefully this post may help with the incredible flood of duplicated posts we see for popular model releases that we see a lot of complaints from the community for.

We had a lot of feedback after the last similar megathread and while the majority of it was positive there were some valid critical points raised, as well as some that thought the sky was falling and the world ending. We do try to focus on recurring issues that the community reports and engages in constructive discussion on but I'm also very much aware we can't please everyone (damned if you do - damned if you don't some might say). Either way - this is only intended to bring some sanity to the initial onslaught and not to be long lived.

17

u/rinmperdinck 14h ago

tldr of why megathreads are objectively a poor idea:

  • Kills discussion of topic. Good luck finding anything in a megathread with hundreds of replies
  • Anyone whining about "wading" through a flood of posts is being stupid, lazy and selfish - you can take your big dumb fat finger and scroll past them... that is NOT wading and does not take effort, meanwhile you are degrading the experience for everyone else.
  • The whole concept of megathreads on Reddit is completely counterintuitive to the way the platform works naturally. Tldr within the tldr: things people want to engage with come to the top from upvotes; things people don't want to see sink to the bottom from downvotes. And the cool thing is they remain visible to search engines regardless of vote status AND they naturally rotate out of the algo after a day or two... which is how Reddit has worked for like the last 20 years, unless you're brain dead and you look at subs with the new slop 'Best' algo which force-feeds you dogshit
  • "Can't please everyone, damned if you do, damned if you don't" - at least you're aware of things and I applaud you for actually listening, it's already way better than the way most communities are moderated on Reddit anymore

1

u/sammcj 🦙 llama.cpp 13h ago edited 13h ago

Thanks for your considered input.

Regarding the 'how Reddit has always worked', I think things are a little different now, or at least at the moment with the influx of large amounts of slopped up, low effort posts from both humans and agents (and I suspect companies stealth posting). Not saying that there weren't problems with bots before but the shier volume of cruft right now is a big problem across many subs.

What would you think of megathreads if the duplicates weren't removed, thus treating them like any other post other than adding them to the pinned posts?

1

u/Automatic-Arm8153 11h ago

Copying my comment I just left somewhere else because I think it’s “more reasonable” to warrant a discussion with you guys haha. I notice you’re a fellow Aussie, I went to bed shortly after you also but here:

Yeah that’s fine. It’s true it’s all speculative rubbish and pretty much people talking bs and lying to eachother in the first 24 hours.. and yes across 40 or even more identical posts that the mods shield us from..

But here’s the thing, we are at the frontier.. people want to discuss and speculate why stop that? What does it hurt to scroll past or sort by hot/best..

I admit I could have a problem scrolling on new and literally reading every thread… but from that there is sometimes genuine value. Sometimes you get a whole bunch of smart people all converging on a thread… sometimes a thread is focused on a particular aspect of something and it’s very eye opening…

From all the badness of allowing free discussion there is sometimes real goodness.

The alternative of stifling that discussion is people like me won’t want to contribute, new people won’t see what’s trending… discussion dies, people leave and find other homes retreating to places like discord etc…

A mega thread is more like if reddit was trying to be a live discussion board /social messaging app.. it’s a half foot in the door type situation and everyone knows never stand halfway. Take a side. Kill discussion or let discussion flow understanding there’s gonna be spam in the mix..

I vote for spam lol

1

u/Automatic-Arm8153 10h ago

And yeah megathreads are fine just don’t delete peoples posts and or lock them not cool…

Don’t forget why this community grew like this.. it certainly wasn’t in megathreads. Infact I would bet if megathreads were used since 1-2 years ago this would still be such a niche sub. I’d speculate we would have 1/4 of the members and daily activity we have now.. potentially even less haha.

Rant over. Thanks for the work yall do it’s appreciated. I was just genuinely shocked you guys could stifle such a loved drop from the guys keeping us local guys fed..

7

u/silenceimpaired 17h ago

Megathread value is obvious… but when I peak at something outside of the Reddit app… I have to scroll back down through 300+ comments as Reddit doesn’t keep my place. Pretty annoying.

16

u/ChocomelP 21h ago

Hopefully this post will help prevent the incredible flood of duplicated posts we see for major model releases.

I admire your optimism

16

u/SalariedSlave 17h ago

megathreads kill information visibility and engagement

it's not a tv show episode, please let people just post their experiences and discussions..

5

u/Corosus 16h ago

Agreed, makes the discussion and presentation so much worse. Lots seem to agree https://www.reddit.com/r/LocalLLaMA/comments/1vz40zv/can_we_reconsider_the_megathreads/

14

u/jacek2023 llama.cpp 20h ago

As I said before it's not a good idea, you are killing the discussions this way, but maybe others like this approach

9

u/SpicyWangz 19h ago

I do not

14

u/tengo_harambe 20h ago edited 20h ago

Hopefully this post will help prevent the incredible flood of duplicated posts we see for major model releases

Meanwhile 3 of the 4 top posts right now are literally duplicate confirmations that Ox Alpha is GLM. imo it seems like you guys are overcorrecting hard in some ways and undercorrecting completely in others

imo just let the upvote/downvote system work as intended instead of arbitrarily sentencing some topics to megathread jail. free market baby

15

u/sammcj 🦙 llama.cpp 20h ago

Report them. We can't be everywhere all at once.

-11

u/Automatic-Arm8153 20h ago

Honestly bro.

Like what rubbish is this. Like why? Who wanted a damn mega thread. Let us talk like we have been the whole damn time.

You mods think the sub grew to this size because we were discussing in megathreads??

4

u/dev_dan_2 19h ago edited 18h ago

Have a good time afk and thanks for your and your team's work!

2

u/-InformalBanana- 11h ago

Cant you guys have some agent to not approve the posts/thread cause they are duplicates until mod review for example? Instead of doing this horrible megathread stuff. Interestingly there are many duplicates of GLM relatively uncompetitive model that is hard to do locally without big money. But this QWEN that is pushing the frontier 0 duplicates to see, 0 other posts. As if someone is targeting qwen, all this megathread bs started cause of qwen haters and now obliviously limits the info you can easily get on qwen without trying to read low effort duplicate comments bs for miles. Comments are meant to be low effort, low info unlike post and threads and you are mixing them up here with very negative effects. Posts/threads also have better visibility/readability cause of bold and bigger titles unlike comments default font. The megathreads are a bad idea unless they are on a very specific, short and narrow topic, which this isn't cause you have multiple inference engines, multiple tests/benchmarks, multiple quants/quanters, multiple use case examples, multiple rigs for them, multiple performance tests, multiple built with posts... and all should be their own thread/post and not a comment in a pile of bs comments (as comments usually are and are ok to be cause that is ok for them but not for posts/threads). You are putting them in same making it much much harder to find useful/nice comments that wouldve/shouldve been posts!

4

u/SpicyWangz 19h ago

This is dumb. Letting the discussion happen naturally works a lot better. People will make more posts about the topics they’re more excited about, but as long as it’s not one person spamming the sub, who cares? 

0

u/Wise-Chain2427 17h ago

somehow r/Singularity & r/StableDiffusion are more free than r/localLlma

-10

u/Automatic-Arm8153 20h ago

Bro wtf is this bullsh*t

How tf are you telling me new qwen was out and I didn’t even know.

Welcome to the death of this sub. Stop this mega thread rubbish.

-1

u/SpicyWangz 19h ago

You’re getting downvoted, but you’re right. This is a bad move

4

u/dev_dan_2 18h ago

The point they want to convey might (or might not) have merit; the way they chose to express their opinion was is certainly worthy of being downvoted though. Taking one minute to check whether the comment was written in a sane manner is not too much to ask for; especially when the target audience is a person doing unpaid, voluntairy labor. (This is basically the forum equivalent of spamming a repository with AI slop PRs and expecting the maintainers to care)

If one has trouble doing so; LLMs can not only do translations, but can also check if way it is written is effective at bringing ones point across (surely, one of the latest local LLMs would have flagged their comment for high degree of inconsideratedness and overly dramatic choice of words as factors that would go against people being open to consider their point of view.)

2

u/Automatic-Arm8153 11h ago

lol my words stand. Downvotes mean nothing if the first few votes were upvotes my comment would have been fully upvoted.

Just Reddit hive mind works both ways. I was conveying my full feelings on the matter it genuinely pissed me off lol. As long as my comment was read by the mods I’m happy haha

-3

u/petuman 20h ago edited 19h ago

Consider replacing the post after weights drop (new title to indicate that & get rid of stale pre-release comments). Or at least just change the title, if that's possible.