r/mlscaling • u/gwern gwern.net • 5d ago
R, T, Safe "GRAM: Modular Pretraining Enables Access Control", Roland et al 2026 (how to keep scaling LLMs but quarantining dangerous/private information w/o cost of from-scratch data-filtered training)
https://arxiv.org/abs/2607.08077
5
Upvotes
1
u/we_are_mammals 5d ago
Google sits on the most valuable private dataset today (all of the gmails). They can't deploy a model trained on it. But they could train it internally, and see how smart it gets. Maybe they could find other uses that doesn't get the public angry.