Is it Time to Block the Internet Archive?

Yesterday, Reddit announced that it would begin “ramping up” restrictions on the Internet Archive, ultimately blocking access to most of its website.
According to Reddit spokesperson Tim Rathschmidt, the move aims to protect users. They say that the Internet Archive has been unable to “defend their site and comply with platform policies. This includes respecting user privacy by deleting removed content.”
However, the bigger point of contention appears to be AI training. In February 2024, Reddit struck a $60 million deal with Google to enable AI training. According to Rathschmidt, they became aware of instances where AI companies would scrape data from the Internet Archive’s servers and use that for training.
As a result, Reddit says it will begin severely restricting the Internet Archive’s access to the site. It will be limited to the home page, not any of the subreddits or individual posts.
It’s a significant change for a site whose slogan was once “The Front Page of the Internet.” However, it does raise a simple question: Is it time for other sites to follow suit? In particular, should sites that are wary of AI block the Internet Archive? It’s a difficult question.
When a Block Isn’t a Block
Last month, I discussed how ChatGPT completely ignored my site’s robots.txt and rehashed a previous 3 Count column when asked about copyright news.
Curious, I later asked ChatGPT about this issue. The AI provided me with five reasons, supposedly, why my site has been indexed without permission. However, it was the second reason that stood out.

To be clear, ChatGPT could not have pulled the data before the robots.txt was added. The block had been in place for over a year, and the article was just one day old at the time. However, the mention of archives and other partners seems more relevant.
The Internet Archive had indexed the post within hours of it going live. So, it’s at least theoretically possible that ChatGPT indexed it from there. Although I think it’s more likely that ChatGPT ignored my robots.txt (something even ChatGPT said was possible), the possibility remains open.
And that is a major problem. No matter what I or anyone else does on their site, if the content appears somewhere else, it can and likely will be used for AI training.
To be clear, that is not a new problem. It’s one that book authors have been wrestling with for years.
The Battle Over Books3
If AI companies have made one thing clear, it’s that they don’t care if they have legal access to the works they train on.
The epitome of this is the use of the Books3 dataset. Simply put, Books3 is a collection of nearly 200,000 pirated books. Weighing in at over 36 GB, the collection is in plain text, making it ideal for training.
Meta, OpenAI, Anthropic and many others have used that collection. Almost none of the books in there are with the permission of their authors or publishers.
Though the question of whether AI training using copyright-protected works is an infringement remains open, AI companies have regularly shown a lack of consideration for the human creators they build off of. If an AI companies knowingly use a dataset filled with pirated textbooks, why would they obey robots.txt, a voluntary system?
To be clear, OpenAI, Anthropic and others do say they will follow the protocol. But that promise is meaningless if they’re going to grab the text from anywhere else.
While I still think that using robots.txt is a worthwhile step since it takes so little energy, it’s clear that it cannot be relied upon.
However, Reddit never relied solely on robots.txt; it also imposed IP blocks, API restrictions and other methods to prevent AI bots. But even those methods weren’t enough, and now it’s taking another step, blocking access to the Internet Archive.
Unfortunately, that likely won’t be enough.
The Bigger Problem
Simply put, when it comes to AI training, an opt-out system can never work. At least not completely.
As long as there is one copy of your work available on a site, archive, database or other location that does not block AI bots, your work will be indexed and trained upon. It doesn’t matter if the other copies were unauthorized or even straight-up piracy.
The AI companies know this. They’ve never been shy about using pirated content. It’s also why they can afford to pay lip service to robots.txt. They know they will most likely get the content one way or another.
There’s no practical way to opt out of a system when others, outside of your control, can defeat your opt-out.
This, in turn, brings us back to the Internet Archive and the Wayback Machine. Reddit claims that this has been a vector for AI bots to crawl and train on their content. However, the statement doesn’t provide any proof or specify which companies are doing it. Still, it’s a claim that makes sense and, at the very least, seems plausible.
But does that mean webmasters should follow Reddit’s lead? It depends on how you feel about both AI training and the Internet Archive/Wayback Machine.
Blocking the Internet Archive closes a potential vector for AI bots. But it also prevents your work from being archived in the Wayback Machine. Though I have been very critical of much of the Internet Archive’s actions, I’ve always felt the Wayback Machine was, for the most part, a public good with little harm to creators.
Many don’t agree with me, and I can see why. But blocking the Internet Archive won’t prevent the issue. It’s just one vector. As ChatGPT itself said, there are other databases, archives and datasets out there.
For me, it doesn’t seem worth it. However, it does highlight the problem with the current system we have with AI crawlers.
Bottom Line
To be clear, I think that opting out is still important. If you don’t want your work to be used to train AI systems, making that clear is a valuable act, even if your wishes are ultimately ignored.
But the story shows just how broken the current system, or lack thereof, is.
Though Reddit’s reasons for blocking other AI companies are far from pure, they highlight the issues others have. There’s no practical way to prevent AI training on your content, regardless of what you create, where you publish it and how many ways you opt out.
The tragedy is that tools like the Wayback Machine may be sacrificed in this fight. As companies and individuals seek to reclaim control of their sites, pulling out of the Wayback Machine becomes a very real consideration.
Still, it’s unclear if this will make any real difference for Reddit. However, it will make a difference for those seeking archived content down the line, as there will be serious gaps in Reddit’s history.
There’s no easy fix for this. Until AI companies respect the wishes of creators on an opt-in basis, stories like this will just become more common. Today it’s the Internet Archive, tomorrow it will be someone else.
This is just another way AI is changing the internet.
Want to Reuse or Republish this Content?
If you want to feature this article in your site, classroom or elsewhere, just let us know! We usually grant permission within 24 hours.
