By Dr. Jonathan Roberts, Chief Innovation Officer at People Inc.:
"A lot has changed in a year. Last summer the questions were:
Does content have value in an AI economy?
Will blocking work?
Will there be a market for information?
In July 2026, we know that good AI needs good inputs, and needs those inputs in real time. Blocking makes products worse, and great inputs make it better. You can’t just throw more chips at the problem.
Data centers are the engine.
Content is the fuel.
In the last 12 months stealing content has become big business:
There are now more bots online than humans.
Multiple AI Search startups have raisedcapital at valuations of over $2bn, selling access to content that they take for free.
Demand is growing exponentially.
At People Inc. we provide fast, direct, licensed access to our premium content for our partners, and have gone deep to block those who are wholesale stealing our content. In February we moved to block all bots from companies we don’t have a partnership with. We allow a short list of named crawlers from trusted partners, and we block tens of thousands of unique bad bots every day, and block tens of millions of daily attempts to take our content.
The most aggressive actors follow a standard pattern of behavior:
They send a named crawler - we block it.
They send an anonymous crawler - we block it.
They send a crawler that spoofs googlebot - we block it.
They send a crawler that attempts to look human, scrolling the page, and executing code - we provide a ‘are you human’ challenge. It fails and is blocked.
They send multiple crawlers through residential internet connections and mobile devices. Because these home proxy networks use legitimate IP addresses they are extremely difficult to identify as being compromised by bad actors and therefore block. We are able to block some of this activity, but not all.
It’s now clear that:
Content has a lot of value in an AI economy.
Blocking works.
There’s a market being built in real time.
In the last year we’ve seen the emergence of around 30 “Napsters of content”, but no Spotify. How do we build a premium information economy that lets people who need information pay the people who create it?
Step 1: stop people taking it for free.
Step 2: flip the market from stealing to paying.
If we get this wrong - a few intermediaries sell the world’s information for their own gain, removing any reason to invest in new knowledge, killing the digital economy, and AI gets worse as inputs degrade.
If we get this right, the exploding demand for knowledge funds investment in new information creation and AI products get much better.
I know which future I want to live in, and we’re working across publishing, CDNs, platforms, legislation and regulation to bring it into reality."
Executive Summary
1
Bots are shockingly effective at evading cybersecurity protections and
paywalls; we observed them impersonate Google and others, rotate IP addresses,
and appear to use residential proxies, hammering publishers with bot
traffic. On UK sites, we saw similar trends to those we observed on US
sites: paywalled content was not necessarily safe from web scrapers, and 95%
of the top UK sites could be scraped by at least one of the 14 scrapers we
tested. Scrapers' disguises are sometimes good enough to fool ad systems — we
observed scrapers triggering Google Ads bots, meaning the page treated a bot
like a human and rendered ads to it, a concern for advertisers paying to reach
humans. In one instance, we observed a scraper hit a publisher's site over 40
times in 3 seconds, creating bandwidth cost the publisher ends up paying.
2
The economics of online publishers are eroding in real time, and European
publishers are being hit hardest. Data from publishers on TollBit shows
European sites get scraped more, receive fewer referrals in return, and have
their robots.txt rules bypassed more often. Median AI scrapes per site were 4×
higher on European sites compared to North American sites, and AI scrapes on
European sites were nearly 20% higher in June than in January 2026, with local
news and sports sites seeing the highest quarter-over-quarter growth. Note:
these numbers reflect only identified bots, so the true scale of AI scraping
is likely even higher.
3
Publishers aren't receiving anything back from the scrapes, and the exchange
is getting worse. It takes 179 AI bot visits to get a single human visitor
referral in return from AI applications for European publishers — a rate over
3x worse than North American sites. This imbalance worsened over H1; the
scrape-to-referral ratio went from 150:1 in Q1 2026 to 227:1 in Q2 2026. AI
bot traffic accounted for only 0.05% of total human referrals on European
sites, which also carry roughly 3x the AI load per human visitor as their
North American peers, seeing about 1 AI bot scrape for every 33 human visits.
4
Robots.txt instructions do very little to stop these bots. The median
European site's robots.txt instructions to not scrape were ignored 2.8x more
often than North American sites'.
5
The methods these scrapers use to evade detection are the same ones used to
access infrastructure or manipulate social media by nation-state hackers and
sophisticated cyber-criminals. Now, these tools are used by venture-backed
Silicon Valley companies serving major corporate clients — without
intervention, the volume will compound exponentially as the companies behind
it scale. A majority of these scraping companies advertise using methods such
as residential proxies, the same functionality used by bad actors in
cyberattacks, to gain access to sites and evade their cybersecurity measures.
6
Bots should not be allowed to mimic humans or other traffic on the web; they
should be required to self-identify. As these disguise techniques grow more
sophisticated, bots become harder to detect, placing an undue burden and
expense on the sites. Self-identification would let publishers and sites
decide how to respond to bot visitors.
Section 1: The Scraper Ecosystem Powering AI Applications & How Proxy Networks Work
Key Insights
The web scrapers that power AI tools are using sophisticated methods to
conceal the identities of the bots they use for scraping using many
techniques. Using our Scraper Audit tool, we can observe how effective
scrapers are in getting access to content, and to an extent, the techniques
they use to gain access:
In testing on sites in the UK, we saw similar trends to those we observed
on sites in the US: paywalled content was not necessarily safe from web
scrapers, and 19 out of the 20 top UK sites we tested could be scraped by
at least one of the 14 scrapers we tested.
In testing on various publisher sites, we observed a number of techniques
that scrapers used to get access to publisher sites, including rotating
IP addresses and masquerading as other user agents, including Googlebot.
We also noticed Google Ads bots could be triggered on sites, a concern for
advertisers who are paying for their ads to be seen by humans, and not
scraper bots.
One of the most concerning methods scrapers are using is residential
proxies, the same technique bad actors use to mask cyberattacks. Residential
proxies are a potential national security concern as they make it virtually
impossible to know who is on the other end, operating a bot. A majority of
the web scrapers in our Scraper Index
openly advertise that they use residential proxies.
AI developers are combining multiple data acquisition processes and
technologies, many offered by specialist web scraping vendors, into tech
stacks that obtain access to the Internet. Fueled by the AI boom, hundreds of
these third-party scrapers are in use, many by Fortune 1000 companies.
These intermediaries scrape publisher content to resell it as data or search
infrastructure, rather than to directly power their own chatbots or AI
applications. Many of these scrapers employ specialized cybersecurity evasion
tools and do not comply with robots.txt by default. This places an undue
burden on sites to attempt to detect these increasingly hard-to-detect
scrapers.
The data acquisition stack has existed for quite some time, but its effects
have become more visible recently. With AI, the need to access content & data
has exploded. Funded by the AI investment boom, developers are combining
multiple data acquisition processes and technologies, with more advanced or
clandestine approaches deployed where cybersecurity barriers are encountered.
These increasingly complicated tech stacks are intended to secure "just
enough" access to content to ensure the applications they serve can handle
the range of queries needed to compete. These stacks are comprised of some or
all of the following:
Direct crawling and scraping via first-party bots, either declared or
disguised.
Secondary source crawling / scraping, targeting surfaces on which the
content itself may be visible on the open web (i.e., Google Search results
page).
Residential IP proxies, presenting as ordinary user traffic — often with
location and device characteristics/IP addresses that match the human
audience.
Circumvention services that evade bot detection / IP protection measures.
Cloud-based headless browsers that load web content in a browser running in
the cloud, rather than on a user's device. These are often used with a
residential IP proxy: the proxy makes the request appear to come from a
normal human user.
Third-party scraping services for outsourcing the end-to-end process of
scraping a page.
For more detailed information, please visit our
Scraper Index.
For obvious reasons, this environment is one in which it is exceedingly
difficult for digital publishers to exercise control over the use of their
IP. They are in a structurally disadvantaged position, having to respond to
new threats as they emerge, whereas the web scrapers just need to evolve
beyond the latest protection. Most websites deploy only one cybersecurity
tool, whereas AI companies can choose from dozens of scraping tools
interchangeably.
Two notable additions here to our prior reporting on the AI Data Acquisition
stack:
First is the increase in attention we've seen to residential proxies, along
with the potential national security risk associated with them. A
residential proxy routes traffic through everyday consumer devices so that
the request appears to come from a human using an ordinary home internet
connection. This technique is utilized by many web scrapers — and this is the
same exact technique that bad actors use to conceal who they are in
cyberattacks and large-scale cognitive warfare, such as astroturfing
campaigns. Recently, there has been a lot of attention to this in the Wall
Street Journal reporting, with an
article here
on this topic, and a recent separate
piece here
with a view of an FBI director stating "we all want to sort of find and
defeat these Residential Proxy networks," because of their association with
crime. Those efforts are already underway. In July 2026, Google, working with
the FBI and others,
disrupted NetNut,
one of the largest residential proxy networks, which it estimated ran on more
than 2 million hijacked consumer devices such as smart TVs and streaming
boxes. Google reported that in a single week, it observed hundreds of
distinct threat groups, including cybercriminal and espionage actors, using
the network to hide the origin of their traffic. The Google blog post about
this also linked out to a
study by a cybersecurity firm
that noted many of the largest proxy providers' websites highlight their
utility for AI platforms and tools, "implying it is a primary use case for
their residential proxies."
In an analysis of nearly 40 scraping vendors, we find that many explicitly
advertise cybersecurity detection evasion techniques, and many do not
default to abiding by robots.txt. A majority of the web scraping companies in
our Web Scraper Index openly advertise
that they use residential proxies and have pools of hundreds of millions of
IP addresses. Using residential proxies appears to be a fairly standard
practice for scrapers today. Residential proxies make it even more difficult
for websites to detect, and therefore block, allow, or monetize bot/scraper
traffic. In a site's logs, a bot visitor using a residential proxy will look
just like a human visitor.
Another addition to our prior reporting is some data on third-party
scrapers/data resellers across the TollBit network, as we have started to
notice some of their traffic actually starting to stand out in our log data
across publisher sites. (Important to note, many of these third parties do
not self-identify their user agents, but a few do, and so we can analyze
their data). Across all sites on TollBit in H1 2026, we observed that Common
Crawl, the open corpus behind much LLM training, reached the widest share of
any identifiable third-party scraper/crawler. We observed Common Crawler
scrapes on ~74% of sites on TollBit: it roughly doubled its reach
quarter-over-quarter (from Q1 to Q2). Because a single Common Crawl scrape
can support many models and AI applications downstream, its footprint across
sites matters far more than its per-site volume: one scrape from Common Crawl
may then become data that then gets reused to power a vast AI ecosystem.
Two other notable players that scrape content and resell it as a search or
data API: Parallel, a web-search/data API, which had scrapes on 74% of sites
on TollBit, nearly doubled its per-site scraping (+94%), and Diffbot (which
sells a structured "knowledge graph" of the web) roughly tripled per-site
(+195%) and has scrapes on over 33% of sites on TollBit. Because a reseller's
crawl can be resold to many buyers, these figures understate true exposure —
one scrape resurfaces across multiple AI products, strengthening the case for
licensing terms that follow the data downstream.
1.1 Scraper Audit Tool & Results on UK Sites
In order to examine the sophistication of web scrapers against websites'
cybersecurity tooling, TollBit developed the Scraper Audit, a tool for
website owners to test 14 popular web scrapers, many of which are Europe-based
companies, against their own sites. The Scraper Audit allows website owners
to initiate scraping on any page to evaluate how well scrapers evade
defenses, access content, and retrieval completeness.
In selecting web scrapers, we picked a few industry leaders and a couple of
newer scrapers offering advanced technology, despite their company size and
time in market, and have plans to add more to the tool. For each of these web
scrapers, whenever available, we purchased subscriptions of their web
scraping APIs to leverage their full suite of AI-driven parsing, advanced
anti-bot bypassing, and premium proxy networks.
To demonstrate the utility of this tool in auditing scrapers, we conducted an
evaluation of this tool across the top 20 sites based in the UK as defined by
the
Press Gazette.
Our test of these sites resulted in similar outcomes to the tests run for the
previous edition of the report:
Even paywalled content is not necessarily safe. In our testing of 14 of
these scraping tools, we observed that some were able to scrape various
paywalled articles in full.
In fact, non-paywalled sites offered no discernible increase in protection
compared to their paywalled counterparts, as scrapers successfully
retrieved full content from the vast majority of target pages regardless of
subscription barriers.
In the rare instance where all scrapers failed to fully extract data —
occurring in only one of the twenty sites — the failure was primarily
attributed to server-side rendering implementations of the paywall. On
these specific sites, the full article was not present in the initial page
and required server-side authorization before the content was served.
1.2 Scraper Audit: Other Observations
Since our last bot report, we've added functionality to our Scraper Audit
Tool to allow publishers to observe the suspected chain of events that occurs
as a web scraper attempts to access their sites. In the tool, publishers can
enter a specific URL in the Scraper Audit Tool and observe some of the
strategies the web scrapers utilize in an attempt to scrape the page. In sum,
we observed that scrapers often deploy various tactics to obtain access to
sites, some hammering sites dozens of times with different user agents,
sometimes masquerading as other bots (including Googlebot, Chrome, etc.), and
others rotating IP addresses in order to attempt to trick the site's
cybersecurity and allow them to gain access to scrape the page. We also
observed that some of the web scrapers would trigger Google Ads bot on sites
— concerning for advertisers who are paying for their ads to be seen by
humans, and not bots/web scrapers.
A related note: In our last bot report, we noted that we did not observe any
scrapers checking robots.txt files before attempting to access sites.
The strongest signal (Canary verification) we have is that the Scraper Audit
adds a unique ID to the request of each URL we tell the webscraper to scrape,
and we then look for that unique ID in the respective site's logs so that we
can identify which scraper visited the URL. When we find it, we can be highly
confident the request was triggered by the scraper audit test we ran.
The second signal, which we treat as lower confidence, is fingerprint
matching. Some requests that hit the server in response to our scrape don't
carry the unique ID. For those, we build a fingerprint from the request's
attributes (its timestamp, IP address, geo, and user agent — all of which we
can obtain from the respective site's logs) and match it against the scrape
we initiated. A request that lines up across all of these — arriving at the
same moment, from the same IP and geo, with the same user agent — points
strongly to the scraper, even without the ID to confirm it. We are relying on
very close timing to make this assumption (typically sub-second to a few
seconds after invocation), which we assume is related to the scraper test we
run. We know when the scraper was triggered and what page it was triggered
to, because we control when the scrapers are triggered and to which pages,
and we have the publishers' site logs data that includes information like IP
address, timestamp of the scrape to the site, etc.
Here are some real examples of what we observed with this new functionality
on publisher sites. Each of the examples below shows a single scraper
instance attempting to access a publisher's URL:
1.2.1 Example 1: Masquerading as Google & IP Address Rotation
The scraper in this example appeared to have tried more than one disguise to
get in. It first used User-Agent Spoofing and posed as Googlebot to gain
trusted access. When that was blocked, it appeared to switch to a Chrome
browser, and its apparent location shifted from Los Angeles, US, to San
Sebastián de los Reyes, ES, meaning it also changed its IP address. Once it
changed to a browser, it was first served 301 redirects, then kept trying
until it got a 200 and pulled the content. This showcases that when one
identity is refused, a scraper may just change who it claims to be and where
it appears to come from and try again.
Scraper masquerading as Google, request timeline
Scraper masquerading as Google, evidence detail
1.2.2 Example 2: Multiple Browser Impersonations & IP Address Rotation
This attempt to access a publisher's URL appeared to trigger 43 requests from
a single web scraper, routed through many different IPs and locations,
cycling between Chrome and Firefox identities throughout the attempt to
access the site. The locations of the scrapes heavily varied: alongside US
cities (San Jose, Ashburn, Clifton, Pennington), the scrapes arrived from
multiple Hungarian locations (Sopron, Sümeg, Nyíregyháza) and Bahrain (Bu
Kuwarah), all within seconds.
Multiple browser impersonations, request timeline
Multiple browser impersonations, evidence detail
1.2.3 Example 3: Multiple Browsers & IP Rotation
This scrape request appeared to result in 32 requests hitting the URL at
essentially the same time, from many different IPs across scattered US
cities and cycling between Chrome and Firefox identities throughout. One
scrape is arriving from a dozen-plus origins at once, and being disguised on
both the network and identity side. The status column is a real mix of 200s
(served), 403s (blocked), and 301s (redirected), so the publisher's defenses
were firing and worked some of the time, but the blocks didn't matter because
each request came from a different IP. A 403 on one address did nothing to
stop the next from a different IP, and the scraper kept rotating until enough
200s landed to assemble the content.
Multiple browsers and IP rotation, request timeline
Multiple browsers and IP rotation, evidence detail
1.2.4 Example 4: Scraper Triggers Google Ad Bots
The two Google Ads Bot requests are worth taking a look at here. They appear
to have fired because the scraper rendered the page like a real browser,
loading its ads, which is what set them off. The scrape looks human enough to
trigger the ad stack; it could be taken as evidence that the scraper's
disguise is convincing and successful. This presents an ad fraud concern for
advertisers who are paying to see their ads viewed by humans, and not bots.
(This is a topic we mentioned previously in our
Q2 2025 Bot report, Section 1).
Scraper triggering Google Ad bots
TollBit has made the Scraper Audit tool available on the TollBit platform to
help websites test and understand their vulnerability to IP threats from
major web scrapers. If you are interested in seeing how your site is being
scraped by these third parties, please contact
team@tollbit.com.
1.3 The Inefficiency and Cost of Web Scraping
This arms race between scrapers and website operators is not only
inefficient for website owners, who are increasingly spending on
cybersecurity to detect bots that are only getting harder to detect, but it
is also deeply inefficient for AI developers. As mentioned in our prior bot
report, we have found advanced scraping services charging up to $22.50 for
1,000 pages. Given the volume of data needed to serve consumer AI
applications, the data acquisition costs for the most popular chatbots and
specialized agents are likely to run into the tens of millions of dollars a
year. And this is before considering the legal fees needed to fight lawsuits.
We have also observed that scrapers take 16 seconds on average to access
publisher sites, and some of them may fail entirely in securing access to the
page requested, whereas accessing content via a sanctioned route (i.e. via a
platform like TollBit) typically takes less than 0.7 seconds.
There's a better way. Instead of spending huge sums on technical workarounds
to bypass IP controls, AI developers can simply pay for the content they use.
This reduces latency and increases reliability. By relying on licensing
rather than scraping, journalism is funded — not the scrapers — and the media
ecosystem AI depends on is sustained.
1.4 The Limits of Cybersecurity and the Need for Regulation
We've seen publishers increasingly take matters into their own hands to
combat the surge in AI bots, increasingly investing in cybersecurity to
actively block this traffic from their sites or even take action by filing
lawsuits. As mentioned in our
Q3 & Q4 2025 report, in
late 2025, both Reddit and Google turned to the courts over unauthorized
scraping. Google alleged that scrapers had circumvented SearchGuard, the
in-house protection system it built at a cost of millions. If a company with
Google's resources and engineering talent cannot keep determined scrapers
out, an individual publisher has no realistic chance of doing so alone.
Cybersecurity is important, but it is not a foolproof solution. These tools
are costly to run, and many of our publishers, even those with the most
sophisticated defenses in place, have seen scrapers and AI bots get through
anyway.
This brings us to an important limitation of the data in the subsequent
sections of this report: we rely on AI companies self-identifying through
their user agents. Developers sometimes use third-party vendors to scrape on
their behalf; they may masquerade as other user agents, rotate their IP
addresses, use headless browsers that mimic humans, use residential proxies,
etc. This activity is difficult, and often impossible, for sites to trace
back to a specific AI company. The true level of scraping, and of
unpermitted access to European content, is, therefore, likely greater than
this report shows.
As AI traffic becomes harder to distinguish from human traffic, it is
important that regulators take a position that, at the bare minimum, AI bots
should not be allowed to mimic humans on the Internet. Ideally, they would be
required to self-identify and state their intended use, which is what
New York's Stealth Crawler Prohibition Act
would require. That bill, passed by the New York State Legislature in June
2026 and awaiting the Governor's signature, would require crawlers accessing
a news site to disclose their identity and purpose through an accurate
user-agent string before they access the content.
Self-identification enables a site owner to decide how to respond: allow,
block, or charge for access via infrastructure controls like the ones
TollBit provides. It would help ensure that the undue burden of paying to
serve, detect, and block unwanted bots no longer falls on the websites
themselves.
Read the Full Report
Fill out the form below to read the our full report.