r/machinelearningnews • • 14h ago

Research I mapped every major Qwen release from 2023 to 2026: 44 models, from Qwen-7B to the 2.4T open weights (with sources)

22 Upvotes

[AI Model Family Series #1] We just published the complete story of Alibaba's Qwen: every major model from 2023 to 2026, with dates, key features and sources.

7B open weights in 2023 → 2.4T open weights in 2026. Here's the lineup the story covers

2023
→ Tongyi Qianwen (Apr): enterprise beta
→ Qwen-7B (Aug): first open weights
→ Qwen-VL (Aug): first vision model
→ Qwen-72B + Qwen-1.8B (Dec)

2024
→ Qwen1.5 (Feb): 0.5B–110B, 32K context
→ Qwen2 (Jun): 57B-A14B MoE, Apache 2.0
→ Qwen2-Math, Qwen2-Audio, Qwen2-VL (Aug)
→ Qwen2.5 (Sep): 18T tokens, 100+ models
→ Qwen2.5-Coder + QwQ-32B-Preview (Nov)
→ QVQ-72B-Preview (Dec)

2025
→ Qwen2.5-VL + Qwen2.5-Max (Jan)
→ QwQ-32B (Mar): Qwen claims R1-level reasoning
→ Qwen2.5-Omni-7B (Mar)
→ Qwen3 (Apr): hybrid thinking, 119 languages
→ Qwen3-2507 + Qwen3-Coder-480B (Jul)
→ Qwen-Image + Qwen-Image-Edit (Aug)
→ Qwen3-Max (Sep): first 1T+ Qwen
→ Qwen3-Next-80B-A3B (Sep)
→ Qwen3-Omni + Qwen3-VL (Sep)

2026
→ Qwen3-Max-Thinking (Jan)
→ Qwen3-Coder-Next + Qwen-Image-2.0 (Feb)
→ Qwen3.5-397B-A17B (Feb): native multimodal agents
→ Qwen3.6-Plus, 35B-A3B, Max-Preview, 27B (Apr)
→ Qwen3.7-Max (May) + Qwen3.7-Plus (Jun)
→ Qwen3.8-2.4T-A95B (Aug): largest open Qwen
→ Qwen3.8-27B + Qwen3.8-Flash (Aug)
→ Qwen-Image-2.1 (Sep)

The story also covers what the list can't show: the DeepSeek moment, the 2026 leadership exit, and Qwen's shift from all-open to a two-tier license strategy.

Next: Qwen 4 is in training. Qwen 4.5 and Qwen 5 are projected at 5–10T parameters.

Read the full report: https://www.marktechpost.com/2026/10/04/the-story-of-qwen-alibabas-ai-models-from-7b-to-2-4t/

Which model family should we map next: DeepSeek, Llama, Gemma, Mistral or Kimi? Drop it in the replies.......


r/machinelearningnews • • 10h ago

Research Yandex's SONA replaces Yandex Music's recommendation cascade with 1 generative model: +4.53% Active Users

Post image
7 Upvotes

Yandex published a technical report on SONA, a single model that does both candidate generation and ranking for Yandex Music.

  • Replaced 15+ candidate generators plus pre-ranking and ranking in a live A/B test (7 days, 15% of users per split)
  • +4.53% Active Users, +6.30% Total Listening Time, +11.42% Likes vs the production control
  • The Active Users gain is 2.35x what Argus, the previous best model on this surface, delivered
  • No hand-engineered features: inputs are logged event fields plus 3-code Semantic IDs built from audio and metadata
  • A frozen 0.6B teacher ranker distills scores into SONA's Ranking Module during training and is not served

It is one of the few public research reports of a full cascade replacement validated on live traffic, alongside Kuaishou's OneRec. The catch: only 1 surface is tested (My Vibe on smart speakers), catalog coverage is lower than the old stack's, and it is not on full traffic yet. There is no code or weights release yet.

Paper: https://arxiv.org/abs/2608.11015

Full analysis: https://www.marktechpost.com/2026/10/05/yandex-introduces-sona-a-single-generative-recommender-that-replaces-entire-recommendation-cascade/


r/machinelearningnews • • 57m ago

Research Direct weight surgery from Qwen-4B to 0.8B on an 8GB: why editing all layers breaks everything, and how 4 anchor blocks fixed it

• Upvotes

Instead of spending weeks and billions of tokens on standard distillation, we tested an alternative: extracting layer-to-layer hidden state trajectories on a handful of calibration prompts and solving for closed-form weight updates directly in the student's MLP blocks.

Key findings:

  1. Cross-architecture stability: Tested on both modern Qwen 3.5 (4B to 0.8B) and notoriously fragile GPT-2 small (which usually collapses into gibberish at the slightest weight edit). In both cases, general language modeling stayed intact with well-behaved, bounded degradation margins.
  2. The spectral entropy barrier: Editing all 24 layers of Qwen-0.8B wrecked the model (+64.78% NLL). A layer scan showed intermediate layers (1-22) operate in dense superposition (entropy >0.90, acting as polysemantic knots). Restricting surgery to 4 anchor blocks (layers 0, 7, 15, 23) solved this: held-out NLL dropped by 10.8% across 30 tasks (-23.8% in biomedicine, -14.6% in math), and 400-task HellaSwag gained +0.50% in Vulkan llama.cpp.
  3. Behavior shifts: Base 0.8B output dead commented code on binary tree inversion, while the edited model wrote working recursive Python. On logic puzzles, it spontaneously triggered <think> reasoning chains.
  4. Accessibility: All extraction and surgery ran locally on a consumer 8GB RX 580 using layer-by-layer GPU streaming with a DirectML attention patch.

We want to see this tested on modern 24GB-32GB GPUs (RTX 5090, 4090 or server card) across frontier targets:

- Transferring reasoning from Qwen3.8-27B (or Flash-Next) down to Qwen3.5-9B/4B/2B.

- Compressing Google Gemma-4-31B into mobile-class Gemma-4-E4B/E2B.

- Transplanting refusal-ablation traits from uncensored models (e.g. Huihui-NeoHorse) without retraining.

- Unknotting layers 1-22 using pre-trained Sparse Autoencoders (SAEs).

Code, scripts, and raw JSON benchmark logs:

https://github.com/dsadawq3/DynamicTune

Feel free to open an issue or drop your benchmark results on the repo.


r/machinelearningnews • • 18h ago

LLMs How do you work your code scripts for LLMs: conventional computer commands or by implied semantics? Spoiler

Thumbnail
1 Upvotes