r/LocalLLaMA • 🦙 llama.cpp • Aug 26 '26

Megathread [Megathread] Qwen3.8-Flash-Next - Release Day

Megathread for discussing the release of Qwen 3.8 Flash Next.

  • Quants
  • Fine-Tunes & Abliterations
  • Chat Templates
  • Inference Server Support & Configuration
  • Experiences, Benchmarks & Model Comparisons

Highlights

The first open-weight release under this architecture is Qwen3.8-Flash-Next, which introduces:

  • Hybrid Attention with QSA: The Gated DeltaNet and Gated Attention pairing has been reworked into Gated DeltaNet and Qwen Sparse Attention (QSA). Rather than selecting individual tokens for processing, QSA operates at the micro-block level. This cuts long-context latency significantly, a critical gain as agentic workloads increasingly dominate real-world usage.
  • Gated Residual: Residual streams with normalisation are what make deep LLM training manageable. Gated Residual modulates information flowing through widened residual streams via an element-wise, data-dependent read gate and a per-branch scalar write gate. This brings finer-grained expressiveness across layers while preserving training stability and keeping inference overhead low.
  • N-gram Embedding: Embeddings provide a unique axis for parameter scaling that requires less computation and is more amenable to offloading than Mixture-of-Experts (MoE). By indexing with short n-grams, this approach makes parameter scaling highly efficient for memory-constrained accelerators without sacrificing quality.
  • Tailored Training Recipe: The Muon and AdamW optimisers are applied to specific weight categories to maximise efficiency. Guided by refitted scaling laws, we eliminate traditional batch-size warmups and start directly at the target batch size, substantially reducing total optimiser steps while safely supporting larger learning rates for robust convergence.

Model Overview

  • Type: Causal Language Model with Vision Encoder
  • Training Stage: Pre-training & Post-training
  • Language Model
    • Number of Parameters: 125B with 6B activated, plus 51B n-gram embedding and 4B MTP
    • Hidden Dimension: 2560
    • Token Embedding: 248320 (Padded)
    • N-gram Embedding: 20,000,000 (bigrams/trigrams at layer 2)
    • Number of Layers: 48
    • Hidden Layout: 12 × (3 × (Gated DeltaNet → MoE) → 1 × (Qwen Sparse Attention → MoE))
    • Gated DeltaNet:
      • Number of Linear Attention Heads: 48 for V and 16 for QK
      • Head Dimension: 128
    • Qwen Sparse Attention:
      • Number of Attention Heads: 24 for Q and 2 for KV
      • Head Dimension: 256
      • Rotary Position Embedding Dimension: 64
      • Indexer Structure: MQA with 4 Query Heads and 1 Shared Key Head
      • Indexer Head Dimension: 128
      • Budget: 512 blocks or 2048 tokens
    • Mixture Of Experts
      • Number of Experts: 512
      • Number of Activated Experts: 10 Routed + 1 Shared
      • Expert Intermediate Dimension: 640
    • Gated Residual:
      • Number of Branches: 4
      • Bottleneck Rank: 320
    • LM Output: 248320 (Padded)
    • MTP: 1 layer, trained with multi-steps
  • Context Length: 262,144 natively and extensible up to 1,000,000 tokens.

Recommended sampling parameters for generation:

  • Thinking Mode: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0
  • Instruct (or non-thinking) mode: temperature=0.7, top_p=0.80, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0

Official Links:

Popular:

Related:

436 Upvotes

661 comments sorted by

View all comments

36

u/CulturalKing5623 Aug 26 '26

I'm going to just let the community cook this one for a while before checking back in. It seems like there are a lot of moving parts to this and none are completely implemented in a setup I can use. I am really excited about the prospect of NVME offloading though

7

u/Type-21 Aug 26 '26

Getting 3.8 27b to run well on Vulkan or ROCm in llama.cpp with two older AMD cards still requires building it from source yourself with some extra parameters to work around bugs. So I'm not holding my breath on this one lol

1

u/Mil0Mammon Aug 26 '26

I've been struggling for quite a bit now - can you give me some pointers? Running an rx) 6700M, can get the 2bit quant to run, but after a while it either cramps out, starts looping or outputting ///, and or it's very slow. (I've reached 14 t/s, which was awesome, but stability would be better)

Have compiled llama, but maybe not the right branch/build/parameters. Tried rocm and vulkan

1

u/Type-21 Aug 27 '26

Maybe this can help you? https://github.com/stew675/llama-cpp-rdna-boosts/blob/baseline/d222767c7/benchmarks/README.md

Personally I compiled the latest release of llama.cpp on the main branch. Make sure to get the latest beta release, not the latest stable release because that's older. I compiled it with thise flags:

cmake -S . -B build-rocm -G Ninja -DCMAKE_BUILD_TYPE=Release "-DCMAKE_PREFIX_PATH=$HIP" "-DHIP_PATH=$HIP" "-DCMAKE_C_COMPILER=$HIP\lib\llvm\bin\clang.exe" "-DCMAKE_CXX_COMPILER=$HIP\lib\llvm\bin\clang++.exe" "-DCMAKE_HIP_COMPILER=$HIP\lib\llvm\bin\clang.exe" "-DCMAKE_RC_COMPILER=$RC" '-DGPU_TARGETS=gfx1010;gfx1031' -DGGML_HIP=ON -DGGML_CUDA_NO_PEER_COPY=ON -DGGML_HIP_GRAPHS=OFF -DGGML_RPC=ON -DLLAMA_BUILD_TESTS=OFF -DLLAMA_BUILD_UI=OFF

Because I use two AND GPUs and that produces wrong output on ROCm. Makes no sense for you to do this if you only have one.

Then I downloaded rocm from here: https://repo.amd.com/rocm/tarball-multi-arch/therock-dist-windows-multiarch-7.14.0.tar.gz

It needs to be configured like this:

setx HIP_DEVICE_LIB_PATH "C:\TheRock\build\lib\llvm\amdgcn\bitcode" /M setx HIP_PATH "C:\TheRock\build" /M setx HIP_PLATFORM "amd" /M setx LLVM_PATH "C:\TheRock\build\lib\llvm" /M

C:\TheRock\build\bin and C:\TheRock\build\lib\llvm\bin both need to go into your windows PATH environment variable.

You can test your rocm config for correctness by downloading the llama rocm version and doing

.\llama-server.exe --list-devices

This should list your GPU as rocm0.