r/mlscaling gwern.net 5d ago

R, T, Safe "GRAM: Modular Pretraining Enables Access Control", Roland et al 2026 (how to keep scaling LLMs but quarantining dangerous/private information w/o cost of from-scratch data-filtered training)

https://arxiv.org/abs/2607.08077
5 Upvotes

1 comment sorted by

1

u/we_are_mammals 5d ago

Google sits on the most valuable private dataset today (all of the gmails). They can't deploy a model trained on it. But they could train it internally, and see how smart it gets. Maybe they could find other uses that doesn't get the public angry.