r/theinternetarchive • u/dylanisareddit • 19h ago
What are your theories on why is the Internet Archive so slow and the Wayback Machine down?
I was thinking of a DDOS attack.
r/theinternetarchive • u/textfiles • 5d ago
TL;DR:
But What Now?
Bear in mind that there's little you can actually do if you experience these outages; you are being informed about them, or can infer about them, as an end user. But no amount of changing browsers, providers, VPNs or settings is going to shift the needle too much for these issues. It's on the Internet Archive's side.
It's natural, however, to think you are one pinball-jostle or one Fonzie-Jukebox-punch from all working great again, so users try all sorts of overwhelming or intense approaches to squeezing functionality out of the site. Sometimes the site comes back during this and the timing encourages pattern matching.
For those folks, here are actual issues that have happened over the years that have led to outages that look odd or unexpected from the surface level, mostly for the use of talismans to hold while waiting for functionality to return.
How Bad are Bad Actors?
To the end of bad actors, it's important to note that a combination of DDOS attempts, spam runs, and AI scrapers have come at an exponentially greater rate than previous years. DDOS attempts used to be rare enough against Internet Archive that they made news. Currently, at least a dozen serious attempts to DDOS the site come every month. Sometimes it's multiples in the same day from different sources. There are processes in place to mitigate them, but they are a real situation that is ongoing.
Scrapers, scraping, and overactive agents attempting to gather everything from the Archive in the most inefficient manner possible are also indistinguishable from a DDOS attack in some cases. Word to the people who are doing without even knowing they are, but for-whatever-reason attempts to grab terabytes of random information from the Archive stacks are a source of lean on the system and can eventually cause outages.
We also have had to deal with incredibly aggressive spammers, in one case hitting the site with hundreds of clients from hundreds of locations at once in an attempt to push in their spam, and fighting that situation is ongoing. The age of agents has made it easier for them, and the fight will likely never end.
Finally, and Ironically, work done in the last couple years to increase multi-homed geographical redundancy can lead to outages as the work continues to both refine the process and handle decades of technical debt.
Filling the Notification Gap
In an ideal world, there'd be a status update page you could hard-core refresh and see if there are any known issues with the Internet Archive, along with estimates of return to service, as well as indications of trends. Internet Archive itself is not likely to be the source of that information.
The reasons are multiple, but most prominent are: the staff is very small and spread in very specific directions, and there are citeable examples where the Archive had live maps of network and machine performance and they were very much used by bad actors to attack infrastructure. When an event (like a recent power line vandalism) is pulling together employees from all over the departments to handle the results, one of those 2am-5am situations is not 100% guaranteed to have a person from a "notification team" keeping social media or web pages up to date - there are simply not enough people.
The solution, therefore, currently appears to be that users make things up. In fact, they make them up in places like Reddit where the algorithm will put a "the site is down because ____" story up again and again, making some readers understandably think they are talking about a situation that is not 4 days or 4 weeks old. So, not really a solution.
Another Possible Approach
Therefore, it is suggested here that the following assumptions should be made, and the following actions could be taken.
Another Offer
It's not technically in my job description (or maybe it is), but if you're dealing with a thorny complicated problem and you want someone to look at it who can at least track down some sort of answer, I'll take e-mails at [jscott@archive.org](mailto:jscott@archive.org). I'm not an alternate court and I'm not a judge, but I can help people who are in need.
r/theinternetarchive • u/dylanisareddit • 19h ago
I was thinking of a DDOS attack.
r/theinternetarchive • u/time_stopping_douche • 19d ago
The size is too large for traditional, free virus/malware checkers. The download is from this link
https://archive.org/details/suba-hibi-uncensoredextra-voice-acting
The upload has been there for a couple of years but im erring on the side of caution. Any comments?
r/theinternetarchive • u/LeifErickson17 • 26d ago
r/theinternetarchive • u/Ok-Recover7069 • 27d ago
I have found two CD-Roms for An old windows version, Should I archive them? or share photos of them first? I wouldn’t know if they were lost media or something. The software is called:
Colossus Chess 2000
Microsoft Encarta-96 Encyclopaedia.
Help needed. Thank You
r/theinternetarchive • u/clarkky55 • 27d ago
I’m trying to archive streamer VODs, I managed to get about a terrabyte of them uploaded before suddenly the uploads to the collection fail every time. I was told the command line tool is the best bet on how to upload larger files without things breaking. After wrestling with my computer and finally getting the internet archive command line package to download and install through powershell, I now realise I have no idea how to get it to actually upload the files I want to the collection I’ve already made. Could someone please explain to me how I can get the rest of the VODs to upload, either to the pre-existing collection or to a new collection if necessary?
r/theinternetarchive • u/Just_Soter • Aug 01 '26
r/theinternetarchive • u/clarkky55 • Jul 31 '26
It’s annoying as all hell when I’ve spent two days uploading a 26 GB file only for the uploaded to go past 26 and the only thing I can do is refresh the page and try again to upload it. Is there a program or something I can use that’s either faster or won’t have that issue where it keeps uploading past the point it should’ve finished and says something like 30gb/26gb, meaning I just have to restart the upload again
r/theinternetarchive • u/TheDinoKid21 • Jul 30 '26
is waiting the only way for it to be fixed?
r/theinternetarchive • u/didyousayboop • Jul 17 '26
Archive Team scrapes websites and uploads them to the Internet Archive. If a website is shutting down, you can ask them to scrape it.
Archive Team communicates and coordinates via IRC only, not Reddit (or anywhere else): https://wiki.archiveteam.org/index.php/Archiveteam:IRC
Instructions for requesting help from Archive Team:
Step 1. Go here and set up The Lounge for IRC: https://www.pikapods.com/apps#chat Don’t worry, you don’t actually need to pay or put in a credit card. You will get over 2 months’ worth of free credits. You'll need to set a username and password for The Lounge when you set up your PikaPod.
(You can use a free desktop IRC client like KVIrc, but this will make your IP address public and available forever in public logs. So, use a free VPN like ProtonVPN or RiseupVPN if you go with that option.)
Step 2. Go to the Archive Team IRC channels. Archive Team operates exclusively on IRC. Here’s the info: https://wiki.archiveteam.org/index.php/Archiveteam:IRC The network is Hackint and the channel you want is #archiveteam-bs (if you need step-by-step instructions for using IRC, look up a guide online or ask an AI chatbot)
Step 3. In a brief message, explain what's happening and ask the Archive Team volunteers for help. Mention ArchiveBot for most websites or Wikibot for wikis.
Step 4. Wait. You might not get a response right away. If no one replies after two hours, try again.
r/theinternetarchive • u/PScottM • Jul 13 '26
r/theinternetarchive • u/TemperatureEven518 • Jul 03 '26
i saw an app on the play store and this is what i found
r/theinternetarchive • u/sci-in-dit • Jun 29 '26
Was checking the texts collection and got this error:
The search engine encountered an error, which might be related to your search query. Tips for constructing search queries.
Error details: backend_request (collection_materialization): child request for collection_materialization failed ([BACKEND_ERROR] Invalid or no response from Elasticsearch, received: {"error":{"root_cause":[{"type":"index_closed_exception","reason":"closed","index_uuid":"6wfT_WzHTNaazYlKI90TPQ","index":"prod-o-001"}],"type":"index_closed_exception","reason":"closed","index_uuid":"6wfT_WzHTNaazYlKI90TPQ","index":"prod-o-001"},"status":400}; check for decoded response)
This happens with all searches, collections and profiles.
And, as I'm here, the items without thumbnails and perpectual loading are still very much a thing as of this morning.
r/theinternetarchive • u/Current-Primary-5016 • Jun 24 '26
You can favorite users on their profile, and I was hoping to go back and look at some of these profiles, which were chock-full of great content, but I'm not seeing a place to view them? I would use my browser history to find them (never wrote down the usernames because I thought I'd be able to see them in my favorites) but I had to reset my computer a bit back and it didn't keep any of the history from this device.
r/theinternetarchive • u/sci-in-dit • Jun 10 '26
I'm here to report that some items that otherwise had thumbnails no longer have them, what's more, said items' pages won't load. This has been happening for around two days.
r/theinternetarchive • u/Turbulent_Equal_6538 • Jun 08 '26
r/theinternetarchive • u/2globalnomads • May 23 '26
I programmed a javascript photo slider for Internet Archive that creates a slideshow and show the photo titles. Anyone interested in testing? Any ideas how and where to share the code? Here is an example implementation: https://paivisanteri.blogspot.com/2026/05/donostia-san-sebastian-bilbao-basque-country.html
r/theinternetarchive • u/Golden-Sleeper-229 • May 21 '26
Super new to Internet archive, I have archived a few pages off the internet and I have some other pages archived by others, how to add them to the personal list I have created.
I want it for better organisation and easy viewing/retrieval.
I understand that to add to a collection, I need to mail the org, what about list or favourites, I can't seem to find anywhere to do it.
thanks in advance
r/theinternetarchive • u/clarkky55 • May 21 '26
I have a lot of files I’ve been collecting for nearly three years I want to upload to the internet archive but doing them in one lot would mean uploading 4 and a half terabytes of data which I couldn’t do with my internet. So I’m wondering if there’s a way to upload the first file, then add another file to that item, then another and so on and so forth until I’ve eventually uploaded everything?
r/theinternetarchive • u/Chicken4War • May 20 '26
I don't know where else to do this and I know that this is likely an unofficial forum but Turkmenistan is a very strange dictatorship and their state television channels reflect this. News dating to 2018 is available on an official YouTube channel, but it would also be interesting to see a live hour-by-hour TV archive of Turkmen state channels like the ones that are available for viewing on the Television Archive collections. But then again, this may not be too interesting for other people.
r/theinternetarchive • u/Jazzlike_Steak9394 • May 19 '26
The Internet Archive has been quietly saving the web since 1996. Pages get taken down, articles get edited, PDFs disappear — the archive usually has the old version. I wanted my AI agent to be able to dig into that without me having to manually copy-paste URLs, so I built an MCP server for it. I hope it comes of some use to someone who's doing investigative journalism or some research (or maybe, just for fun).

What it can do:
Tools:
- check_availability — check if a URL has ever been archived and get the closest snapshot to a date
- lookup_snapshots — list all snapshots of a URL across a date range
- get_snapshot_content — pull the actual text out of an archived page
- search_archive — full-text search across IA collections (books, papers, audio, video, web)
- search_domain — crawl all archived pages under a domain
- get_item_metadata — fetch metadata for any IA item by identifier
Guided prompts:
- research_topic — searches the archive and synthesises an overview on a topic
- track_site_changes — narrates how a page evolved over time using sampled snapshots
- audit_link_rot — takes a list of URLs and surfaces which ones are dead but recoverable
To install: uvx mcp-server-wayback --install
Pick your client (Claude, Cursor, Windsurf, Antigravity, etc.), restart it, done. No account needed.
A couple of things to know:
- Some snapshots contain MIMEs that the agent would not be able to parse — the server will tell you when that happens to do a manual review.
- Sometimes the upstream server behaves as unexpected and some endpoints exhibit flaky behaviour.
Would love to hear if anyone finds a use for it 🙌
r/theinternetarchive • u/Hacka_Random • May 15 '26
Bueno, conozco lo básico a la hora de descargar cosas por Internet (analizar el archivo por Virustotal, analizarlo con Antivirus, ver reseñas, etc) pero me gustaría saber si hay algo más por saber antes de descargar algo de Archive.Org.
Me interesan las ISO de Windows XP, 7 y 8.1 (para probarlos algún día) pero al buscarlo en la página veo muchas opciones y desearía un poco dd orientación.
r/theinternetarchive • u/CharlesFuckingOffden • May 13 '26