You Can (Try to) Prevent Your Photos & Content from Being Used to Train AI Models With These Methods

The current options for preventing AI systems from training on our photos and text are inadequate. But there are methods available that are worth trying.

Tommy'. Photo by David Coleman - havecamerawilltravel.com
Text & Photos By David Coleman
Last Revised & Updated:
Filed Under: News

I MAY get commissions for purchases made through links in this post.

Have Camera Will Travel // Newsletter

Practical field notes, hands-on gear tests.

In July 2025, Cloudflare ramped up their efforts to control AI crawling on sites by make blocking automatic, allowing publishers and site owners to opt-in.

They’ve had optional AI crawler blocking available for a while, but it required site owners to proactively enable it (which I’ve been using for some time now; it seems to work well).

The core change is that it’s now enabled by default. There is also some provision for selected large sites to selectively allow crawlers to access a site in exchange for negotiated compensation. Cloudflare posted a more detailed explanation here.

Big tech corporations are training their AI systems on data that’s available on the web. They’re scraping content at huge scale, and they’re not asking permission. It’s gotten so bad that it’s sometimes knocking libraries, archives, and museums offline. Even Wikipedia is struggling to contain it.

To some, it’s fair use. For others, it’s flagrant theft of intellectual property. If you’re reading this post, odds are good that you lean more toward the latter than the former.

Some, like ChatGPT, have said that they’re developing a method for content owners to opt out of their content being used to train AI models.1

But that feels more PR stunt than actually useful, a bit like closing the barn doors after the horses have bolted. The reality is they’ve already vacuumed up huge swathes of content that is publicly accessible (and, as is increasingly obvious, more than a little that is behind paywalls). I have zero confidence that they’ll remove any content they’ve already included in their training corpus.

More importantly, as someone who makes their living from my intellectual property, I’d argue that this should all be opt in. In many cases, especially with photographers, our content is copyrighted (whether or not you’ve submitted your images to the U.S. Copyright Office or local copyright body).

But the reality is that that is not how AI companies are proceeding. They’re taking the content to train their systems without asking permission.2

The Washington Post has a helpful tool where you can look up to see what sites have been used in some of the training data. You can use it to see if your site is included. Several of my sites are in there.

Things to try

The reality is that our options as the creators and owners of intellectual property are currently woefully inadequate.

But if you’re more inclined to take active measures than howl at the wind, there are some things you can try.

None of this is anything approaching a satisfactory solution, but in the short term, it may be better than nothing (especially in combination).

Importantly, most of these are pretty easy to implement, so they don’t require a massive investment of time and energy.

That said, the reality is that the options are limited. Until there’s legislation with teeth and some real enforcement mechanisms, some of these might not amount to much more than wishful thinking. But, again, on the principle that trying something is better than not trying anything . . .

Finally, it’s a very fluid situation that’s changing rapidly. I’ll do my best to keep this updated with the latest changes, but things will change and what applies in one situation might not apply the same way in another.

Robots.txt content signals

This is a new method being developed and pushed by Cloudflare, among others. So far, it’s unproven, and it has the usual robots.txt limitation in that it’s basically an honor system and is only effective if the crawlers respect it.

The gist is that you can add some content signals to the beginning of your robots.txt to specify how your site’s content can be used. The three types of use specified are:

search: building a search index and providing search results (e.g., returning hyperlinks and short excerpts from your website’s contents). Search does not include providing AI-generated search summaries.

ai-input: inputting content into one or more AI models (e.g., retrieval augmented generation, grounding, or other real-time taking of content for generative AI search answers).

ai-train: training or fine-tuning AI models.

You can find more about it, including the full text you can copy and paste, here.

Robots.txt AI crawler user agents

The most obvious first step is to try to prevent AI systems from crawling and scraping your content in the first place. It’s also the area where there are explicit mechanisms in place intended to address this.

It takes advantage of what is known as the Robots Exclusion Protocol. It provides a simple but potentially effective way to keep bots from crawling all or parts of a website. It’s been around for years, is simple and can be effective, and implementing it for AI crawlers fits neatly with well-established practice.

Here’s the catch, though, and it’s a big one: whether the crawlers respect these rules is entirely voluntarily. While reputable actors should adhere to it, there’s no guarantee that everyone will, and there are no enforcement mechanisms. It’s a handshake deal, basically a situation of: “stop, or I’ll say stop again.” And in recent months there’s growing evidence that several AI companies are pretty much ignoring them or even being duplicitous by saying they adhere to it and then using measures to flout it.

I’m interested in what’s going on here with robots.txt. It’s a good-faith agreement that has held up for decades now, and it’s falling apart thanks to unscrupulous AI companies . . .

Elizabeth Lopatto, “Perplexity’s grand theft AI,” The Verge (27 June 2024)

But on the principle that at least some of the companies are respecting robots.txt at least some of the time, here’s the running list of user agents. You can find information below on the syntax for implementing them:

  • Google-Extended / documentation. Google-extended is a standalone crawler that Google has said explicitly “does not impact a site’s inclusion or ranking in Google Search. But it’s still worth treating such assurances with caution.
  • Apple-extended / documentation.
  • GPTBot / documentation. This also prevents ChatGPT-User, which is the part that works with OpenAI web search functionality. In other words, blocking GPTBot will also block a site from appearing in ChatGPT search results.
  • Anthropic-ai. Seems to have switched to ClaudeBot.
  • ClaudeBot / documentation.
  • PerplexityBot / (I haven’t found official documentation on this one, but it is out in the wild. NB: Wired has caught Perplexity ignoring the robots.txt.
  • CCbot / documentation. Common Crawl creates some of the underlying databases that are then used as part of the training data in AI companies.
  • PiplBot / documentation.
  • Bytespider / TikTok’s crawler. Reportedly ignores robots.txt, in which case adding it probably won’t do anything. Perhaps a more effective route is to use Cloudflare’s AI blocking tool (more on that below).
  • ImagesifBot / documentation.

It’s worth noting that this is only a very small portion of the bots out there, and that’s another reason that robots.txt is a deeply imperfect tool for this task. But this list does include the most active crawlers that appear to underpin many of the better-known AI company training sets. But, again, there’s nothing to stop those same companies from creating new, undeclared user agents or even deliberately misidentifying the user agents (spoofing the user agent).

So it’s unrealistic and misguided to expect your robots.txt to keep your content out of the hands of all AI crawlers.

  • OAI-SearchBot. This was announced when OpenAI announced that it was testing SearchGPT. Their explanation: “OAI-SearchBot is for search. OAI-SearchBot is used to link to and surface websites in search results in the SearchGPT prototype. It is not used to crawl content to train OpenAI’s generative AI foundation models.”

Example robots.txt code

That said, it can be an easy and effective first step. Below is an example of the syntax to use these user agents in your robots.txt file.

But first, a word of warning. You can ignore this and proceed if you already know your way around a robots.txt file. But if you don’t, it’s worth knowing that you can create problems with faulty instructions in this file. It won’t crash your site with a white screen of death, but there might be effects that you don’t see for a while. The biggest risk is that you accidentally instruct search engines to deindex your site and thereby render your site functionally invisible to search. So if you’re uncomfortable with working with the robots.txt file, it might be worth asking a professional developer or your website host to help you with it.

User-agent: CCBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: GPTBot
Disallow: /
User-agent: PiplBot
Disallow: /
User-agent: Apple-extended
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: PerplexityBot
Disallow: /
User-agent: Bytespider
Disallow: /
User-agent: ImagesiftBot
Disallow: /

Cloudflare option to block AI scrapers & crawlers

This is a brand new option that Cloudflare has made available. It’s a simple switch designed to block AI scrapers and crawlers.

It was initially launched quietly, but they’ve since added a much more detailed explanation of how it works. And, as you’d expect from Cloudflare, it sounds like it will be very effective and thorough. I’ve enabled it on my sites, but it’s too early to say what effect it’s having. It’s certainly not going to claw back content that’s already been pilfered, but I have quite a bit of confidence in Cloudflare’s ability to tackle this problem given not just their expertise but especially their unique position in the web infrastructure.

This feature is available for Cloudflare free and paid accounts, and you can find this in your Cloudflare account by going to a site and then:

Security > Bots > AI Scrapers and Crawlers

Dialing it up with stronger, targeted restrictions

The robots.txt method is widely adopted and well-established, but it’s not much stronger than a suggestion.

Sultan Ahmed Mosque Interior Hand Painted Tiles Calligraphy Istanbul Turkey. Photo by David Coleman - havecamerawilltravel.com

Have Camera Will Travel // Newsletter

Practical field notes, hands-on gear tests, and the fun of exploring the world with a camera in hand.

It’s also possible to build on this idea by prohibiting, restricting, or rate-limiting crawlers more directly.

But that involves much more technical chops than what I’m focusing on here. And it involves knowing and tracking the IP addresses that the crawlers are using at any given time, because some crawlers don’t identify themselves correctly. And they change all the time. So it’s a bit like swatting at flies.

UPDATE: A newer alternative, and a lot simpler, is to use Cloudflare’s AI blocking tool. It doesn’t give you any granular control, but it outsources the tricky technical part to a service perfectly positioned to deal with it.

But if you want to dial it up and block or rate limit crawlers, it’s possible, but it’s also tricky and beyond the scope of the measures I’m focusing on here.

Feeling mischievous? Serve the AI bots a 10GB file to chew on! :)

Terms of use to prevent AI scraping

This section focuses on a less technical and more legalistic approach. But I am not a lawyer, and nothing here is intended as legal advice.

A website’s terms of use can be a useful mechanism to include. It’s rarely preventative, but it might offer some extra teeth should it ever come to legal action.3

But there are no blanket rules for this–every situation and site is different–and you’ll need a lawyer to provide legal advice on this that’s specific to your content and situation.

That said, Raptive, one of the largest ad agencies working with independent content creators, has put together some template language as part of their guidance for asserting creative rights over their intellectual property.

Their example language, which I’ve adopted on my sites, is:

The owner of this website does not consent to the content on this website being used or downloaded by any third parties for the purposes of developing, training or operating artificial intelligence or other machine learning systems (“Artificial Intelligence Purposes”), except as authorized by the owner in writing (including written electronic communication). Absent such consent, users of this website, including any third parties accessing the website through automated systems, are prohibited from using any of the content on the website for Artificial Intelligence Purposes. Users or automated systems that fail to respect these choices will be considered to have breached these Terms of Service/this Agreement.

We have included on the pages of this website a robots meta tag with the “noai” or “noimageai” directive in the head section of the HTML page. Please note that even if such directives are not present on any web page or content file, this website still does not grant consent to use any content for Artificial Intelligence Purposes unless such consent is expressly contained.

If you decide to use this template, you’ll need to replace “these Terms of Service/this Agreement” with the title of your site’s Terms of Service or similar page that contains this material. And if you created a standalone “No AI” page, omit the last sentence.

You can find more about the robots meta tag mentioned in the last paragraph further down this page.

Real-world examples of anti-AI scraping terms of use

In addition to that, here are some real-world examples of terms of service / terms of use that I’ve come across being used as at the time of writing. I’m focusing here on some of the big hitters that seem relevant to news and photography.

Again, these might change, but I’ve included the direct links so you can check the current language.

New York Times

The New York Times is one of the most prominent outlets that have pushed back against AI system training on its data and is suing OpenAI and Microsoft over their use of NY Times content.

PROHIBITED USE OF THE SERVICES

(3) use the Content for the development of any software program, model, algorithm, or generative AI tool, including, but not limited to, training or using the Content in connection with the development or operation of a machine learning or artificial intelligence (AI) system (including any use of the Content for training, fine tuning, or grounding the machine learning or AI system or as part of retrieval-augmented generation).

Source: https://help.nytimes.com/hc/en-us/articles/115014893428-Terms-of-Service#4

Getty Images

Getty has a number of different license agreements and terms of use for different aspects of its business. And it’s a little muddied because they have their own custom AI Generator service.

But this is what they have in their content license agreement:

No Machine Learning, AI, or Biometric Technology Use. Unless explicitly authorized in a Getty Images invoice, sales order confirmation or license agreement, you may not use content (including any caption information, keywords or other metadata associated with content) for any machine learning and/or artificial intelligence purposes, or for any technologies designed or intended for the identification of natural persons. Additionally, Getty Images does not represent or warrant that consent has been obtained for such uses with respect to model-released content.

Source: https://www.gettyimages.com/eula

Associated Press

As of June 2024, I have not seen any language in their terms of service that explicitly mentions AI systems or LLMs, but it is presumably covered under this:

You may not (i) deep link or employ software or any automatic device, technology or algorithm, to “crawl”, “scrape”, search or monitor this Site and/or retrieve or copy Content or related information;

Source: https://www.ap.org/terms-and-conditions/

Reuters

c.Without our prior written consent, you may not: . . .

use the Service or Content for any purpose relating to the development and training of any machine learning (“ML”) and artificial intelligence (“AI”) activities and technologies, including, but not limited to using the Service or Content to build, train, enhance or tune any AI or ML technologies. You may not, at any time, directly or indirectly, use the Service or Content for any purpose relating to the development and training of any machine learning (“ML”) and artificial intelligence (“AI”) activities and technologies, including, but not limited to using the Service or Content to build, create, train, retrain, enhance or tune any AI or ML technologies (whether belonging to you or a third party).

Source: https://www.reuters.com/info-pages/terms-of-use/

Alamy

I’ve yet to find a specific mention of generative AI or LLMs in Alamy’s terms, but they are very much aware of the problem and have been taking action.

Specifically for images

There are different types of generative AI models currently in use. Some focus more on text; others more on imagery.

Here are a couple of methods that focus more on the imagery side of things.

Robots meta tag

You can also add a short line of code to the header of your HTML. It is being pushed by DevitantArt, one of the oldest and largest hubs for creator-driven graphic art.

But this is not anywhere near a standard yet, and there’s no guarantee that crawlers will respect it, and no real way of knowing if they do.

The specific tag they’re recommending is:

<meta name=”robots” content=”noai, noimageai”>

This is something included Raptive’s guidance. Put it in the header of your HTML.

Photo watermarks

Look, watermarks on photos are one of those things that photographers argue about endlessly.

I’m not going to rehash the pros and cons here, but my general view is that they’re more useful for brand building and pretty much useless for preventing someone from stealing your images.

But AI changes the balance slightly for me. On the one hand, AI imaging tools can remove watermarks from images really effectively, so it might seem that they’d be especially ineffective. But on the other hand — and this is the crucial point here — AI companies aren’t bothering to do that. It would require too much compute time for them to do it for every image; it would mean massive costs. Instead, they’re ingesting and even regurgitating images with the watermark intact (or almost intact).

The upshot is that there are some real-world examples already where watermarks have survived whatever transformations the AI image generation has done and providing red-handed evidence of precisely which source images were used (or “stolen,” depending on where you sit on this). That makes it pretty easy to draw a line directly from the generated image back to the source.

Here’s a good example that was included in Getty Images’ lawsuit against Stable Diffusion.

In short, this has added to my incentive to add watermarks to my images. Not because I think they’re a silver bullet — they’re not — but because it’s another arrow in my quiver if lawyers ever become involved.

How effective are these?

I want to be clear here. This doesn’t fix it. And it doesn’t make it right. But if you’re inclined to try something, these are good places to start.

Will any of these measures actually make a difference in the long run? Who knows. AI models have insatiable appetites for content, and right now the march to AI looks awfully inexorable. And the current combination of laws and enforcement is completely inadequate for the task. You’d think existing copyright laws should cover this situation, but in the real world, that’s just not working.

By themselves, the methods I’ve outlined above don’t amount to much. But they’re what we currently have to work with. And they they might start to matter more if governments start legislating in this area in meaningful ways. Europe has been leading the way for this type of thing in recent years, and it wouldn’t surprise me if we start seeing some early steps sometime soon.

Notes & References:

  1. Others, from Microsoft and ChatGPT, have expressed a much more entitled view, going so far as to say that they see content on the web as “freeware” available for training AI and that maybe some creative jobs shouldn’t have been there in the first place. [↩︎]
  2. In some cases, they’re signing deals to access the content, but that doesn’t apply to most of us. [↩︎]
  3. AI companies are also tweaking their terms of service in a different way to make it easier for them to use the data on their services for AI training. [↩︎]
Found this helpful? Personalize your search results.
Add as Preferred Source
Profile photo of David Coleman | Have Camera Will Travel | Washington DC-based Professional Photographer

David Coleman

I take photos for a living. Now based in Washington, D.C., I have spent the last 30+ years shooting across seven continents — from underwater environments to mountain peaks. My images and time-lapses have appeared in major newspapers, magazines, museums, professional sports stadiums, and even on massive architectural scrims covering world-famous buildings.

I started this site in 2009 to field-test gear, share problem-solving solutions, and share what I've learned in shooting around the world. I only review gear and services I have personally used. No armchair opinions — just real-world experience. Because, for me, the fun is in the making of the photo.

You can see my travel photography here, license my images or buy prints here, or sign up for my Substack.

Leave a Comment