Skip to main content
Back to Insights

AI crawlers on your site: welcome them, meter them, or block them

Valentin Zsigmond
Valentin ZsigmondSep 23, 20265 min read

The web has never been only humans. For as long as there have been websites, there have been automated visitors: programs that fetch pages for some purpose of their own. The most familiar are search engine crawlers, and most site owners are glad to see them, because being crawled is how you get found. Automated traffic is simply a fact of running a site, and most of it is either useful or harmless.

Lately a new kind of automated visitor has become common, and it sits in an interesting middle ground. These are the crawlers run to gather content for artificial intelligence: some collect text to train models, others fetch pages in real time to answer a question someone has asked an assistant. They are not the search crawlers you have always wanted, and they are not the malicious scrapers you have always wanted to keep out. They are something new, and whether you welcome them is a genuine choice rather than an obvious yes or no.

Our view is that this is the site owner's decision to make, not ours and not the crawler's. So we treat it as a setting: each site can choose a stance, and we enforce that stance where it is most effective, at the network edge before the traffic ever reaches the site.

A new kind of visitor

It helps to be clear about what these crawlers are doing, because they are not all the same. Some are gathering large amounts of text to help train AI models. Others arrive at the moment a person asks an assistant something, fetch a page or two to find the answer, and may show or link your content in the response. The first is about building a system; the second is closer to a new way people might discover you.

That difference is why there is no single right answer. A crawler that might send interested readers your way is a different proposition from one that reads everything you have published to help build something you will never see a return from. Reasonable site owners look at the same two crawlers and come to opposite conclusions, and both can be right for their own situation.

Why the answer is not obvious

Consider the case for welcoming them. If assistants are becoming a way people find information, being readable by them is a little like being crawlable by search engines was twenty years ago. You may want your content represented, your business mentioned, your pages available as a source. For plenty of organizations, visibility is worth more than anything a crawler costs them.

Now consider the case against. Your content is your work, and you may not want it absorbed into a system with no return to you. Heavy crawling also has a cost: a determined crawler working through a large site consumes server resources that exist for your visitors. For some owners the calculation is simple, and the answer is no. Most sit somewhere in between, wanting to stay open without being consumed wholesale. That middle ground is exactly why a single blanket rule serves no one, and a choice does.

robots.txt and the conventions taking shape

The web already has a long-standing way for a site to state its wishes to automated visitors: a file called robots.txt at the root of the site, which lists who is welcome and which parts they may visit. Well-behaved crawlers read it and respect it, and it has quietly governed search crawling for decades.

The same mechanism is now being extended to AI crawlers. Many of them publish the names they identify themselves by, so a site can name them in robots.txt and say yes or no. Newer conventions are also emerging that let a site express more specific preferences about how its content may be used, and the picture is still settling. We keep track of these as they develop, so a site we look after states its preferences in the way the current conventions expect, rather than as a guess that no crawler recognizes.

Why polite requests need enforcement behind them

There is an important limit to all of this. robots.txt is a request, not a barrier. A well-behaved crawler reads it and complies, and the well-behaved ones are the majority. But a request only works on those willing to honor it, and there is nothing in the file itself that stops a crawler that chooses to ignore it.

So a stated preference needs something behind it for the cases where politeness runs out. If a site owner has decided to keep a certain kind of crawler out, that decision should hold whether or not the crawler cooperates. This is where a stated wish has to be backed by the ability to enforce it, and where the choice becomes real rather than advisory.

Three stances, and enforcing them at the edge

In practice we offer a site owner three broad stances. Welcome, which means these crawlers are treated much like search engines and left to do their work. Rate limit, which means they are allowed but held to a measured pace, so they can read the site without ever overwhelming it or crowding out real visitors. And block, which means they are turned away. A site can also mix these, welcoming one kind of crawler and blocking another, because the two are genuinely different.

Whatever the choice, we enforce it at the edge, in the network tier that sits in front of the site. This matters for a practical reason. Identifying and handling automated traffic there means the work of sorting it out never touches the servers that serve your visitors, and a person reading your site is never slowed down or challenged because of a policy aimed at machines. The same edge that keeps genuinely hostile traffic in check is where these preferences are applied, calmly and without collateral effect.

The owner decides

None of this is a reason for alarm. Automated traffic, including this new kind, is a normal part of being on the web, and it can be kept in its lane like anything else. What has changed is that there is a new decision to make, one that did not exist a few years ago, and that the right answer depends entirely on what a particular site is for.

What we provide is the ability to make that decision deliberately and to have it hold. Some of the sites we look after welcome AI crawlers, some meter them, some keep them out, and a few treat different crawlers differently. All of those are correct, because they reflect what each owner wants. The stance is yours to choose; making it stick is our part.

Share this article
Valentin Zsigmond
Written by
Valentin Zsigmond
Founder & Lead Engineer

Full-stack Drupal architect with almost two decades of experience leading large teams in enterprise environments. Founded Tilizy Digital to bring senior-level expertise directly to the organizations that need it. Mentor to dozens of Drupal developers and creator of SDX — a Drupal extension for building modern frontends with React and Vue.

Enjoyed this article?

If this resonated, imagine what we could do working together on your Drupal site.