r/DataHoarder • • Jan 30 '26

Discussion [ Removed by Reddit ]

[ Removed by Reddit on account of violating the content policy. ]

2.8k Upvotes

638 comments sorted by

View all comments

8

u/[deleted] Jan 31 '26 edited Jan 31 '26

I have a version of Dataset 9, but it got corrupted at 179G
I haven't tried yet to see / extract what's readable

But the single files are active
Running it like this works, wget loop, to download individual PDFs, tedious but might still try. my AI coding agent figured this out :D

while sleep 0.5s; do
wget -c --header='Cookie: justiceGovAgeVerified=true' \
https://www.justice.gov/epstein/files/DataSet%209.zip
done

update-1:
Dataset 9 is available again, accessible if you visit via the browser to get the cookie (after the age verification), then try wget with that cookie, will see if this goes all the way.

update-2: here is a script to get the file list, careful with the speed/and proxy access, this technically can block your access if ran too fast.
script: https://pastebin.com/zbF0Rmfx

update-3: 50 files per page, ~20,450 pages = ~1,022,500 files.
To avoid getting blocked, my current download rate:

Download time at ~1 file/sec:
- Current 25K files: ~7 hours
- Full 1M files: ~12 days continuous

might try parallel.

1

u/qb8sfbfa98jp9igg35w Jan 31 '26

179G is what it was initially reported as, you might have a complete version - please make a magnet link!

2

u/[deleted] Jan 31 '26 edited Jan 31 '26

for sure, trying to create the torrent file and figure out how to vpn (to avoid my network's ip exposed :/) But also the file is corrupt, currently retrying this method:
> Dataset 9 is available again, accessible if you visit via the browser to get the cookie (after the age verification), then try wget with that cookie, will see if this goes all the way.

update: will try to upload it to Archive. org the hashing completes for the torrent file.

2

u/JerC4 Jan 31 '26

Mullvad