State of the Bots

2026 Q1 & Q2, The Bad Bots

Foreword

By Dr. Jonathan Roberts, Chief Innovation Officer at People Inc.:

"A lot has changed in a year. Last summer the questions were:

  • Does content have value in an AI economy?
  • Will blocking work?
  • Will there be a market for information?

In July 2026, we know that good AI needs good inputs, and needs those inputs in real time. Blocking makes products worse, and great inputs make it better. You can’t just throw more chips at the problem.

  • Data centers are the engine.
  • Content is the fuel.

In the last 12 months stealing content has become big business:

  • There are now more bots online than humans.
  • Multiple AI Search startups have raised capital at valuations of over $2bn, selling access to content that they take for free.
  • Demand is growing exponentially.

At People Inc. we provide fast, direct, licensed access to our premium content for our partners, and have gone deep to block those who are wholesale stealing our content. In February we moved to block all bots from companies we don’t have a partnership with. We allow a short list of named crawlers from trusted partners, and we block tens of thousands of unique bad bots every day, and block tens of millions of daily attempts to take our content.

The most aggressive actors follow a standard pattern of behavior:

  1. They send a named crawler - we block it.
  2. They send an anonymous crawler - we block it.
  3. They send a crawler that spoofs googlebot - we block it.
  4. They send a crawler that attempts to look human, scrolling the page, and executing code - we provide a ‘are you human’ challenge. It fails and is blocked.
  5. They send multiple crawlers through residential internet connections and mobile devices. Because these home proxy networks use legitimate IP addresses they are extremely difficult to identify as being compromised by bad actors and therefore block. We are able to block some of this activity, but not all.

It’s now clear that:

  • Content has a lot of value in an AI economy.
  • Blocking works.
  • There’s a market being built in real time.

In the last year we’ve seen the emergence of around 30 “Napsters of content”, but no Spotify. How do we build a premium information economy that lets people who need information pay the people who create it?

  • Step 1: stop people taking it for free.
  • Step 2: flip the market from stealing to paying.

If we get this wrong - a few intermediaries sell the world’s information for their own gain, removing any reason to invest in new knowledge, killing the digital economy, and AI gets worse as inputs degrade.

If we get this right, the exploding demand for knowledge funds investment in new information creation and AI products get much better.

I know which future I want to live in, and we’re working across publishing, CDNs, platforms, legislation and regulation to bring it into reality."

Executive Summary

1

Bots are shockingly effective at evading cybersecurity protections and paywalls; we observed them impersonate Google and others, rotate IP addresses, and appear to use residential proxies, hammering publishers with bot traffic. On UK sites, we saw similar trends to those we observed on US sites: paywalled content was not necessarily safe from web scrapers, and 95% of the top UK sites could be scraped by at least one of the 14 scrapers we tested. Scrapers' disguises are sometimes good enough to fool ad systems — we observed scrapers triggering Google Ads bots, meaning the page treated a bot like a human and rendered ads to it, a concern for advertisers paying to reach humans. In one instance, we observed a scraper hit a publisher's site over 40 times in 3 seconds, creating bandwidth cost the publisher ends up paying.

2

The economics of online publishers are eroding in real time, and European publishers are being hit hardest. Data from publishers on TollBit shows European sites get scraped more, receive fewer referrals in return, and have their robots.txt rules bypassed more often. Median AI scrapes per site were 4× higher on European sites compared to North American sites, and AI scrapes on European sites were nearly 20% higher in June than in January 2026, with local news and sports sites seeing the highest quarter-over-quarter growth. Note: these numbers reflect only identified bots, so the true scale of AI scraping is likely even higher.

3

Publishers aren't receiving anything back from the scrapes, and the exchange is getting worse. It takes 179 AI bot visits to get a single human visitor referral in return from AI applications for European publishers — a rate over 3x worse than North American sites. This imbalance worsened over H1; the scrape-to-referral ratio went from 150:1 in Q1 2026 to 227:1 in Q2 2026. AI bot traffic accounted for only 0.05% of total human referrals on European sites, which also carry roughly 3x the AI load per human visitor as their North American peers, seeing about 1 AI bot scrape for every 33 human visits.

4

Robots.txt instructions do very little to stop these bots. The median European site's robots.txt instructions to not scrape were ignored 2.8x more often than North American sites'.

5

The methods these scrapers use to evade detection are the same ones used to access infrastructure or manipulate social media by nation-state hackers and sophisticated cyber-criminals. Now, these tools are used by venture-backed Silicon Valley companies serving major corporate clients — without intervention, the volume will compound exponentially as the companies behind it scale. A majority of these scraping companies advertise using methods such as residential proxies, the same functionality used by bad actors in cyberattacks, to gain access to sites and evade their cybersecurity measures.

6

Bots should not be allowed to mimic humans or other traffic on the web; they should be required to self-identify. As these disguise techniques grow more sophisticated, bots become harder to detect, placing an undue burden and expense on the sites. Self-identification would let publishers and sites decide how to respond to bot visitors.

Read the Full Report

Fill out the form below to read the our full report.

We're a small startup tackling big problems. May we (very) occasionally email you updates like this one to help push the industry forward?

TollBit State of the BotsQ1 & Q2 2026new