Does Robots.txt Matter Anymore?

Image of an Angry Robot

If you run a website, you’ve likely had to deal with the Robots Exclusion Protocol, or robots.txt, as it’s better known.

The protocol’s idea is relatively simple. Robots.txt is a small, formatted text document that lets bots, like search engine crawlers, know whether they are allowed to crawl the site and what they are allowed to crawl.

For most of the internet’s history, it has only had one target: search engines. In most cases, it wasn’t about restricting access but guiding search engines to the content web admins wanted indexed. 

Used in conjunction with sitemaps, it was a tool to guide search engines through public spaces and block off private ones, such as paywalled content. As a result, though it was and is an important standard, it never got much attention.

However, that changed with the rise of AI. Suddenly, there was a new interest in blocking certain robots. Suddenly, web admins were doing what I did and adding requests to block AI crawlers.

But that proved to be wildly ineffective. As I showed on this site, robots.txt does nothing to prevent your site from being ingested for AI training. In fact, there may not be any practical solution for opting out, robots.txt or otherwise.

A Brief History of Robots.txt

Martijn Koster first proposed the idea of robots.txt in February 1994. At the time, the primary concern was that bad-acting bots could overwhelm servers and either waste precious resources or cause the site to go down.

It quickly became a de facto standard, and the major search engines of the day all agreed to follow it. Google did the same when it launched in 1998. 

However, for most of its history, the protocol was never formalized. That began to change in 2019 when Google started pushing to make it an official internet standard. However, as of this writing, the standard is still only a proposal published in September 2022.

But even before Google began pushing to make robots.txt a standard, cracks and limitations had started to show. The most significant and most apparent was that there was no method for enforcement. The standard was, and still is, voluntary.

That came to a head in April 2017, when the Internet Archive announced it would no longer listen to or obey robots.txt instructions. Though they offered other ways to remove sites from the Archive, they said that following robots.txt instructions was counter to their goal of internet preservation.

Clearly, “bad” bots had ignored robots.txt since day one. However, the Internet Archive was, supposedly, one of the good bots. Still, the decision largely flew under the radar because the number of sites impacted was likely very low.

That, like everything else, changed in November 2022. That was when ChatGPT launched to the public, and users had broad access to generative AI tools for the first time. Though the technology isn’t new, it was the first time the average internet user could use it.

That brought the backlash about how AI was and is trained. However, as per a June 2024 report from Reuters, AI companies often ignored robots.txt. 

A May 2025 prepublication paper from Duke University backed up that finding, saying that many bots never even checked the robots.txt file. To make matters worse, compliance dropped as the rules became stricter. In short, a robots.txt that disallowed a bot was less likely to be followed than one that set rate limits.

However, this problem is not new and many companies have been creating solutions for “bad” bots for a long time. 

Fighting Back Against Bad Bots

Since robots.txt doesn’t have an enforcement system, many companies have been trying to create their own.

The most notable is Cloudflare, which launched in 2009. Even today, one of its core promises is blocking unwanted bots. In July 2025, they even launched a service to monetize AI crawlers

Other content delivery networks, such as Akami and AT&T, offer similar bot-blocking services. 

Where robots.txt is voluntary, this actively blocks unwanted bots and crawlers, denying them access to your site.

But even these tools have limitations. 

First, it will always be a game of cat and mouse with the bots. That fight has been ongoing for decades, and AI doesn’t move the needle on that issue.

But the bigger problem is that bots, including AI bots, don’t need access to your site to scrape your content. Most likely, that content exists elsewhere with or without your permission, on sites you don’t control (including, possibly, the Internet Archive). 

Your content is, almost certainly, being copied and republished on the internet. Even if it doesn’t show up in search engines, that doesn’t mean other bots can’t find it.

As long as the system is opt-out, there is no practical way for you, or anyone else, to prevent AI-oriented scraping.

If blocking bots directly can’t prevent scraping, there’s no way that robots.txt can. If active enforcement can’t solve the problem, there’s no hope that a voluntary system will.

Bottom Line

Robots.txt can’t prevent scraping. It was never meant to. It’s a voluntary standard that bad bots have long been free to ignore without repercussions. 

Instead, it was meant to be a cooperative effort between sites and search engines to prevent bot-caused issues while steering crawlers toward the correct pages. For legitimate search engines, it truly is a win-win.

That, in turn, is the only significance the standard now has. It’s a tool for search engine optimization. However, even that has been overshadowed by sitemaps and other more granular tools.

When it comes to blocking bots, it’s almost entirely useless. It always has been. 

However, one thing that has changed is the number of “good” bots that ignore it. When the Internet Archive decided not to honor it in 2017, it normalized ignoring the standard if it was against the operators’ goals.

Now, companies valued in the hundreds of billions ignore robots.txt. Would these AI companies feel comfortable doing so if the Internet Archive and others had not paved the way? Probably. They’ve been quietly grabbing this content for many years and, according to their legal arguments, they feel entitled to it. 

But I can’t pretend that the erosion of the standard didn’t make it easier. 

Previously, the sole determining factor between a good bot and a bad bot was whether it followed robots.txt. I’d argue that’s still the case. Consent still matters, even in an opt-out world.

It’s just a shame that the owners of the bots don’t feel the same way.

Want to Reuse or Republish this Content?

If you want to feature this article in your site, classroom or elsewhere, just let us know! We usually grant permission within 24 hours.

Click Here to Get Permission for Free