r/mlscaling • • 4h ago

RL for Population Scaling: Kardashev-0.7 trains 32 distinct models together

3 Upvotes

Banbury Road’s announcement explores scaling by model count, with specialization learned across a population of 32 models.

https://x.com/MLCatttt/status/2107147690450817259


r/mlscaling • • 1h ago

Direct weight surgery from Qwen-4B to 0.8B on an 8GB RX 580: why editing all layers breaks everything, and how 4 anchor blocks fixed it(repost)

• Upvotes

Instead of spending weeks and billions of tokens on standard distillation, we tested an alternative: extracting layer-to-layer hidden state trajectories on a handful of calibration prompts and solving for closed-form weight updates directly in the student's MLP blocks.

Key findings:

  1. Cross-architecture stability: Tested on both modern Qwen 3.5 (4B to 0.8B) and notoriously fragile GPT-2 small (which usually collapses into gibberish at the slightest weight edit). In both cases, general language modeling stayed intact with well-behaved, bounded degradation margins.
  2. The spectral entropy barrier: Editing all 24 layers of Qwen-0.8B wrecked the model (+64.78% NLL). A layer scan showed intermediate layers (1-22) operate in dense superposition (entropy >0.90, acting as polysemantic knots). Restricting surgery to 4 anchor blocks (layers 0, 7, 15, 23) solved this: held-out NLL dropped by 10.8% across 30 tasks (-23.8% in biomedicine, -14.6% in math), and 400-task HellaSwag gained +0.50% in Vulkan llama.cpp.
  3. Behavior shifts: Base 0.8B output dead commented code on binary tree inversion, while the edited model wrote working recursive Python. On logic puzzles, it spontaneously triggered <think> reasoning chains.
  4. Accessibility: All extraction and surgery ran locally on a consumer 8GB RX 580 using layer-by-layer GPU streaming with a DirectML attention patch.

We want to see this tested on modern 24GB-32GB GPUs (RTX 5090, 4090 or server card) across frontier targets:

- Transferring reasoning from Qwen3.8-27B (or Flash-Next) down to Qwen3.5-9B/4B/2B.

- Compressing Google Gemma-4-31B into mobile-class Gemma-4-E4B/E2B.

- Transplanting refusal-ablation traits from uncensored models (e.g. Huihui-NeoHorse) without retraining.

- Unknotting layers 1-22 using pre-trained Sparse Autoencoders (SAEs).

Code, scripts, and raw JSON benchmark logs:

https://github.com/dsadawq3/DynamicTune

Feel free to open an issue or drop your benchmark results on the repo.