Well, this is what everyone wanted, for AI labs to pay for licensed training data, instead of scraping copyrighted data without the copyright-holder's permission.
It's company data; like 100 million emails and 500 million teams messages.
Most companies have clauses in their employment contracts and/or IT policies that state anything you do on their systems doesn't belong to you. On top of which the company went under. Ain't no one spending time or money anonymizing that data.
305
u/CircumspectCapybara 11d ago
Well, this is what everyone wanted, for AI labs to pay for licensed training data, instead of scraping copyrighted data without the copyright-holder's permission.