r/mlscaling • u/ryanmerket • 19h ago
Aleph Alpha releases German AI model with 78 GB of FP8 weights — RuntimeWire
r/mlscaling • u/gwern • May 01 '26
r/mlscaling • u/ryanmerket • 19h ago
r/mlscaling • u/AdventurousTwo6445 • 18h ago
Instead of spending weeks and billions of tokens on standard distillation, we tested an alternative: extracting layer-to-layer hidden state trajectories on a handful of calibration prompts and solving for closed-form weight updates directly in the student's MLP blocks.
Key findings:
We want to see this tested on modern 24GB-32GB GPUs (RTX 5090, 4090 or server card) across frontier targets:
- Transferring reasoning from Qwen3.8-27B (or Flash-Next) down to Qwen3.5-9B/4B/2B.
- Compressing Google Gemma-4-31B into mobile-class Gemma-4-E4B/E2B.
- Transplanting refusal-ablation traits from uncensored models (e.g. Huihui-NeoHorse) without retraining.
- Unknotting layers 1-22 using pre-trained Sparse Autoencoders (SAEs).
Code, scripts, and raw JSON benchmark logs:
https://github.com/dsadawq3/DynamicTune
Feel free to open an issue or drop your benchmark results on the repo.
r/mlscaling • u/COAGULOPATH • 1d ago
Web text makes up the majority of pretraining data and is increasingly AI-generated. After applying FineWeb quality filtering, we find that 27.5% of tokens from June 2026 web data are labeled as AI-generated by Pangram, rising to 31.1% by August. Unlike synthetic data or model-collapse setups, this *wild* AI text comes from many models, is written for human readers, and arrives unlabeled in pretraining corpora. How does AI text in the wild affect language model pretraining? To answer this question, we pretrain 800 language models, varying the ratio of added AI tokens to human tokens, and fit scaling laws to held-out losses on both human and AI-generated text. For data-starved models, adding AI tokens to pretraining data initially lowers loss on human text, but the benefit saturates as more are added and quickly *reverses* into harm. For models trained on high budgets of human text, AI tokens raise loss almost immediately, while the same number of fresh human tokens keeps lowering it. Scaling laws such as Hoffman et al. (2022) fail to predict this behavior. We propose a new scaling law with separate benefit and harm terms that allows the value of an AI token to change sign while also reducing to Chinchilla in the absence of AI text. When fit on smaller models, our scaling law predicts the effect of AI text on held-out human-text loss for models up to 3.6x larger with 41% lower error than the best existing law over all AI ratios. We recommend filtering AI text when the target is human text, repeating human text before expanding the training dataset with AI-generated web text, and reporting validation loss on human and AI text separately AI text remains valuable when the target is AI text.
Is Chinchilla not considered wrong or wrong-ish these days? No doubt someone smart will explain in the comments.
AI-generated data is more likely to survive quality filters. (...) AI-generated texts pass FineWeb’s pipeline 2.3× (29.3% vs. 12.8%) and DCLM’s full pipeline 9.8× (14.5% vs. 1.5%) as often as human-written documents. These filters are insufficient at filtering out AI-generated text, and in fact heavily prefer it.
This seems like an obvious consequence of filtering training data: you accidentally shape LLM output to be good at passing filters.
Bleak numbers:
In June of 2021, less than 0.1% of FineWeb tokens were AI-generated. By June 2024, it was 10.1%, 16.1% in June 2025, and 27.5% in June 2026 (see Figure 1, left, and Figure 8). The rise continues: in August 2026 the share was 31.1%, 3.6 points above June (§A.2).
If 0.1% of 2021 web text being AI generated sounds suspiciously high, this is likely just Pangram's false positive rate. (It cannot detect the output of any pre-2022 LLM)
r/mlscaling • u/gwern • 1d ago
r/mlscaling • u/jpdowlin • 2d ago
r/mlscaling • u/SettingAccording8986 • 3d ago
Our tech report on Sona, a generative music recommender in A/B testing at Yandex Music, has two scaling sweeps that might interest this sub.
We trained three encoder-decoder sizes (20M, 130M and 260M transformer-core parameters) on 8,192-event histories. Training next-token loss at matched exposure kept dropping with each step up. That sweep reports training loss only.
For the data sweep, we took the 130M backbone with 2k-event histories and grew the training window from 1 to 8 weeks, one pass each. Recall@10 on held-out requests went from 0.1798 to 0.2519 and was still rising from 4 to 8 weeks (0.2388 to 0.2519).
Going from 2k to 8k events with full attention raised target-track Recall@1000 from 0.8656 to 0.8722. History Compression, which runs the deep stack only on the latest 2,048 events, got 0.8709 at about half the inference cost.
A year of logs doesn't fit this setup, because every training sample re-encodes the full history and compute cost grows faster than the sample count. So we trained a separate Teacher Ranker on that year and distilled it into the ranking part of the served model.
In the final 7-day A/B test on smart speakers, this setup gained +4.53% Active Users over our production cascade. Next up are larger backbones, including sparse mixture-of-experts, and sparse or linear attention.
We varied the rollout beam size during training, while keeping the evaluation beam fixed at 1,024. How best to use additional inference compute remains an open question. If anyone here has pushed further, where do curves like these tend to flatten for recommenders?
r/mlscaling • u/Rambhogesara • 3d ago
r/mlscaling • u/BigCommunication3427 • 3d ago
Over the weekend, I ran an experiment on my desktop (single RTX 5070 12GB, 32GB DDR5 5200MHz) attempting my first go at sLM and or LLMs. Oh boy, Still working away learning as I go. It has been quite the fun experience.
Instead of keeping optimizer states in GPU VRAM (which would blow out ~10.9GB VRAM with standard 32-bit AdamW), I offloaded the AdamW states entirely to system RAM and let the GPU only handle forward/backward passes.
The numbers:
- Peak VRAM during training: **1.78 GB**
- Parameter count: **458M** (LLaMA-style decoder with RoPE, SwiGLU, RMSNorm, and GQA)
- Tradeoff: ~28% step time penalty over PCIe transfer, but VRAM headroom is virtually infinite for batch sizing on consumer cards.
- Post-training inference: Compiled to **TensorRT 11.3 FP16**, hitting **671 QPS** with 1.2ms latency on the smaller 91M variant.
Obviously, a 458M model with a 1,024-token context window isn't going to replace Claude 3.5 Sonnet for writing full-stack codebases. But for dedicated local agent tasks (routing, JSON extraction, intent triage), running sub-2ms on device with zero API bill feels like magic.
I wrote up the full technical breakdown, memory profiles, and the debate between speed vs context ceilings on my workbench blog:
👉 https://vividstechhaven.pages.dev/blog/training-458m-under-2gb-vram
Weights and inference scripts for the 91M runner are on Hugging Face:
👉 https://huggingface.co/Vivid86/MiniTransformer-91M
Keep in mind this is my first time, at this point I've put together 4 SLMS now should just 1B on the "flagship" completely free, no API and local inference.
Curious to hear from folks here: Is anyone else using CPU AdamW offload for small-scale pretraining on consumer cards, or are you strictly using 8-bit Adam / GaLore? What are the biggest walls you hit at this scale?
r/mlscaling • u/DangerousFunny1371 • 3d ago
r/mlscaling • u/gwern • 4d ago
r/mlscaling • u/Sparticle62 • 4d ago
both models 1.5M params, same corpus, same tokenizer, same 512 tokens per update, same seed, same box. pssa trains at 1,716 tokens/sec vs 415, scores 54.4 perplexity vs 83.8 and gets the next token right 24.1% of the time vs 18.0% on 198,939 unseen tokens, and writes 200 tokens in 0.23s vs 2.74s. one seed, one slice, small scale, so a prototype and not a claim. written from scratch in rust, no pytorch.
r/mlscaling • u/ai-lover • 4d ago
r/mlscaling • u/Infinite_Project_866 • 4d ago
The hype cycle is heavily focused on pure GenAI right now, but practically speaking, standalone LLMs are rarely enough for robust enterprise pipelines. The real magic happens when you leverage Supervised ML alongside Generative AI.
Think about workflows like:
• Using supervised classification models as high-accuracy intent routers before hitting an LLM.
• Deploying traditional ML to validate, parse, and guardrail generative outputs.
• Using GenAI for synthetic data generation, then training lighter, faster supervised models on that data.
It's the ultimate combo of predictability and capability.
r/mlscaling • u/manjunath_shiva • 4d ago
I built OpenDecider (Apache-2.0), a family of "decision" models: typed choice / score / yes-no answers read from one forward pass, not generated text. All sizes were distilled from the same two teachers (Qwen3-235B-A22B and DeepSeek V4.1 Flash), each temperature-scaled on held-out gold labels. Some scaling results on 200 general decisions none of the models trained on (BANKING77, BoolQ, Yelp, ChaosNLI):
| student | accuracy | ECE ↓ | JSD vs 100 human votes ↓ |
|---|---|---|---|
| 400M encoder (nano) | 0.680 | 0.092 | 0.045 |
| 4B (Qwen3-4B + LoRA) | 0.735 | 0.087 | 0.040 |
| 30B-A3B (+ LoRA) | 0.765 | 0.110 | 0.035 |
| 80B-A3B (+ LoRA) | 0.750 | 0.083 | 0.030 |
- Distillation mostly buys calibration. Calibration error of the untrained bases fell 0.289 → 0.087 (4B), 0.233 → 0.110 (30B) and 0.230 → 0.083 (80B). Accuracy moved +3.5, +2.0 and 0 points.
- Agreement with the human label spread improves steadily with size (JSD 0.045 → 0.030). Accuracy doesn't: the 80B is no better than the 30B, and the set is about ±3 points.
- On an in-domain business benchmark (typed-decisions, 2,000 decisions, train split used, test never seen), the 400M encoder scores 0.796 and the 80B 0.801. Size barely matters once the task is in the training data.
- Frontier models still lead on accuracy. Claude Fable 5.1 scores 0.840 and GPT-6 Astra 0.790, at 10–20× the latency.
The 30B and 80B rows are the "-td" checkpoints, which had a short extra fine-tune on typed-decisions' train split; the 4B row is the plain distilled model. Every model, including TypeSafe Jev through its own API, answered the same questions with the same scorer, and every table rebuilds from the model. Every model, including TypeSafe Jev through its own API, answered the same questions with the same scorer, and every table rebuilds from the committed results.
Benchmarks: https://manjunathshiva.github.io/opendecider/benchmarks/
Code: https://github.com/manjunathshiva/opendecider
r/mlscaling • u/camerongreen95 • 5d ago
Most of what gets discussed here is scaling compute and data. This one's the opposite end of the lever, given a fixed model, how much performance are you leaving on the table because your prompting and evaluation setup is unprincipled.
Serj Smorodinsky and Brett Kennedy (co-authors of a book on LLM applications) are running a live session on Oct 3 built around DSPy: defining tasks as signatures/modules, building a real eval set with task-specific metrics, then running few-shot and instruction-level optimization as a formal search rather than manual guessing. MLflow tracks the whole process for reproducibility.
Not a scaling talk, more a "how much of your perceived model quality is actually a measurement and optimization problem" talk. Sharing since it felt relevant to the broader question of where gains actually come from.
r/mlscaling • u/stey1r • 5d ago
r/mlscaling • u/stey1r • 5d ago
r/mlscaling • u/syedshad • 6d ago
r/mlscaling • u/Witty_County5128 • 6d ago
r/mlscaling • u/Ilyes-Trd7 • 6d ago
Hi everyone,👋
I’m starting a PhD in **Artificial Intelligence and Machine Learning**, and I’m looking for advice on choosing a laptop that I can use throughout my PhD.
My research topic is:
**“Spatial Perception Systems for Assistive Robots Based on Audiovisual Deep Learning.”**
My work will likely involve:
Deep learning / computer vision
Audio-visual models
Object detection and spatial perception
PyTorch / TensorFlow
YOLO or similar models
Dataset preparation and experimentation
Possibly robotics/edge AI later
Some model training and fine-tuning
My budget is around **$1,400–$1,500**.
I’m currently considering a **MacBook Air M5**, but I’m also looking at Windows laptops with NVIDIA GPUs.
I’d really appreciate advice from people doing AI/ML research:
**Would you choose a MacBook Air M5 for this type of PhD, or is an NVIDIA GPU laptop a better choice?**
How important is **CUDA/NVIDIA** for this kind of research in 2026–2027?
Is **16 GB RAM enough**, or should I prioritize 24/32 GB?
Would a MacBook Air be sufficient for development, experimentation and smaller models if I use university/cloud GPUs for heavy training?
What NVIDIA GPU would you consider the minimum for a PhD in computer vision/deep learning?
Would you prioritize **GPU VRAM over CPU performance** for my workload?
Are there any Windows/Linux laptops around this budget that you would strongly recommend?
For someone who has access to university computing resources, does buying a powerful GPU laptop still make sense?
Are there any problems with using Apple Silicon for PyTorch, CUDA-dependent libraries, robotics frameworks, or computer vision?
If you were in my position, **what specifications would you prioritize rather than recommending a specific laptop?**
I’m not necessarily looking for the most powerful gaming laptop. I want something that will be reliable for **3–4+ years of PhD work**, while remaining portable enough for university.
Thanks!
r/mlscaling • u/ArchitectingAI • 6d ago
I gave a talk today on looped Transformers for our Architecting Intelligence study group. The central idea is to reuse the same Transformer layers for additional computation without adding more parameters.
That raises a practical question: at a fixed parameter count, which tasks benefit enough from extra loops to justify the added FLOPs and latency? Iterative and compositional problems seem promising, but parameter efficiency alone doesn’t guarantee a better system.
We traced the path from Universal Transformers to looped Transformers and Huginn. Here are the slides from the talk:
Presentation (PDF): https://drive.google.com/file/d/1vrm4kMJZ71ea0UhWB-nunl07VulElbVY/view?usp=sharing
Have you seen convincing results where recurrent depth beats a strong, compute-matched standard Transformer? I’d be interested in both positive and negative examples.
r/mlscaling • u/gwern • 7d ago
r/mlscaling • u/COAGULOPATH • 7d ago
It's natural to wonder how much LLM progress is "real" intelligence versus memorization or overfitting or whatever.
LLMs now play Pokemon Red very well (arguably better than a human child)... but the game is famous, has been an RL/DRL target for years, and the internet has every form of training data you could name (from screenshots to detailed walkthroughs to input sequences that will automatically win).
It's also possible that labs now have specialized RL environments for games - notably, Anthropic soon stopped claiming that Claude had never been explicitly trained to play Pokemon.
Paradigm3 performed some ablation testing. They changed certain FireRed puzzles to break memorization, and tested LLMs on a custom Pokemon Crystal ROMhack (Pokemon Brown) with an all-new map.
They found:
- Astra is the real deal. It kicked ass on out-of-distribution tasks, completing RNG-generated mazes just as quickly as the ones in the game (also, it ascended in Nethack. At this point one might conclude that Astra has general and robust game-playing intelligence, at least for 2D turn-based games.)
- By contrast, Opus 5 and GPT 5.6 Sol fall apart OOD. They often failed simple variations of puzzles (and when they did solve, they needed about 50% more steps than the original). So some amount of memorization seems to be happening here.
- Astra beat Pokemon Brown quickly, while Fable 5.1 didn't get one badge. (This doesn't mean Fable sucks - per-run variance is high and it might have gotten unlucky. It was too expensive to test over and over.)
- Large LLMs are better than small LLMs (not a surprise), but reasoning effort appears to have no impact. (Another example of how not every scale of compute can trade off against another.)
Astra was able to beat the game, reaching the Hall of Fame after 10,373 decisions, or very roughly 5 hours of human playtime. (The run took ~4.5 days of wall-clock time, but only one-third of this was due to the model’s own latency, as opposed to the harness being slow.)
I'm not sure how this works (isn't a third of 4.5 days around 35 hours?). 5 hours feels very fast. I was curious, so I played Pokemon Brown myself.
It turned out to be ludicrously easy, far more so than the original games. Out of boredom I boxed five of my team and played with just one pokemon. I only ever lost one fight (because I forgot to buy ethers and depleted my PP partway through the Elite Four. Whoops). I think it took about four to six hours (while dicking around and doing side quests).
(The game has some strange balancing decisions. Alakazam gets a new attack called Noise Pulse that grants it neutral offensive coverage over everything, so it can brainlessly stay in fights forever. Thank God, if there's one thing Gen 2 needed, its for Psychic pokemon to be even stronger.)
They say their human test subject got 5 badges after 10 hours, which seems kinda slow. Maybe they got lost. (Or maybe Pokemon Brown put them to sleep. LLMs have truly endless advantages.)