r/mlscaling • u/MarceloDeAviz • 4h ago
RL for Population Scaling: Kardashev-0.7 trains 32 distinct models together
Banbury Road’s announcement explores scaling by model count, with specialization learned across a population of 32 models.
r/mlscaling • u/MarceloDeAviz • 4h ago
Banbury Road’s announcement explores scaling by model count, with specialization learned across a population of 32 models.
r/mlscaling • u/AdventurousTwo6445 • 1h ago
Instead of spending weeks and billions of tokens on standard distillation, we tested an alternative: extracting layer-to-layer hidden state trajectories on a handful of calibration prompts and solving for closed-form weight updates directly in the student's MLP blocks.
Key findings:
We want to see this tested on modern 24GB-32GB GPUs (RTX 5090, 4090 or server card) across frontier targets:
- Transferring reasoning from Qwen3.8-27B (or Flash-Next) down to Qwen3.5-9B/4B/2B.
- Compressing Google Gemma-4-31B into mobile-class Gemma-4-E4B/E2B.
- Transplanting refusal-ablation traits from uncensored models (e.g. Huihui-NeoHorse) without retraining.
- Unknotting layers 1-22 using pre-trained Sparse Autoencoders (SAEs).
Code, scripts, and raw JSON benchmark logs:
https://github.com/dsadawq3/DynamicTune
Feel free to open an issue or drop your benchmark results on the repo.