By Dr. Jonathan Roberts, Chief Innovation Officer at People Inc.:
"A lot has changed in a year. Last summer the questions were:
Does content have value in an AI economy?
Will blocking work?
Will there be a market for information?
In July 2026, we know that good AI needs good inputs, and needs those inputs in real time. Blocking makes products worse, and great inputs make it better. You can’t just throw more chips at the problem.
Data centers are the engine.
Content is the fuel.
In the last 12 months stealing content has become big business:
There are now more bots online than humans.
Multiple AI Search startups have raisedcapital at valuations of over $2bn, selling access to content that they take for free.
Demand is growing exponentially.
At People Inc. we provide fast, direct, licensed access to our premium content for our partners, and have gone deep to block those who are wholesale stealing our content. In February we moved to block all bots from companies we don’t have a partnership with. We allow a short list of named crawlers from trusted partners, and we block tens of thousands of unique bad bots every day, and block tens of millions of daily attempts to take our content.
The most aggressive actors follow a standard pattern of behavior:
They send a named crawler - we block it.
They send an anonymous crawler - we block it.
They send a crawler that spoofs googlebot - we block it.
They send a crawler that attempts to look human, scrolling the page, and executing code - we provide a ‘are you human’ challenge. It fails and is blocked.
They send multiple crawlers through residential internet connections and mobile devices. Because these home proxy networks use legitimate IP addresses they are extremely difficult to identify as being compromised by bad actors and therefore block. We are able to block some of this activity, but not all.
It’s now clear that:
Content has a lot of value in an AI economy.
Blocking works.
There’s a market being built in real time.
In the last year we’ve seen the emergence of around 30 “Napsters of content”, but no Spotify. How do we build a premium information economy that lets people who need information pay the people who create it?
Step 1: stop people taking it for free.
Step 2: flip the market from stealing to paying.
If we get this wrong - a few intermediaries sell the world’s information for their own gain, removing any reason to invest in new knowledge, killing the digital economy, and AI gets worse as inputs degrade.
If we get this right, the exploding demand for knowledge funds investment in new information creation and AI products get much better.
I know which future I want to live in, and we’re working across publishing, CDNs, platforms, legislation and regulation to bring it into reality."
Executive Summary
1
Bots are shockingly effective at evading cybersecurity protections and paywalls; we observed them impersonate Google and others, rotate IP addresses, and appear to use residential proxies, hammering publishers with bot traffic. On UK sites, we saw similar trends to those we observed on US sites: paywalled content was not necessarily safe from web scrapers, and 95% of the top UK sites could be scraped by at least one of the 14 scrapers we tested. Scrapers' disguises are sometimes good enough to fool ad systems — we observed scrapers triggering Google Ads bots, meaning the page treated a bot like a human and rendered ads to it, a concern for advertisers paying to reach humans. In one instance, we observed a scraper hit a publisher's site over 40 times in 3 seconds, creating bandwidth cost the publisher ends up paying.
2
The economics of online publishers are eroding in real time, and European publishers are being hit hardest. Data from publishers on TollBit shows European sites get scraped more, receive fewer referrals in return, and have their robots.txt rules bypassed more often. Median AI scrapes per site were 4× higher on European sites compared to North American sites, and AI scrapes on European sites were nearly 20% higher in June than in January 2026, with local news and sports sites seeing the highest quarter-over-quarter growth. Note: these numbers reflect only identified bots, so the true scale of AI scraping is likely even higher.
3
Publishers aren't receiving anything back from the scrapes, and the exchange is getting worse. It takes 179 AI bot visits to get a single human visitor referral in return from AI applications for European publishers — a rate over 3x worse than North American sites. This imbalance worsened over H1; the scrape-to-referral ratio went from 150:1 in Q1 2026 to 227:1 in Q2 2026. AI bot traffic accounted for only 0.05% of total human referrals on European sites, which also carry roughly 3x the AI load per human visitor as their North American peers, seeing about 1 AI bot scrape for every 33 human visits.
4
Robots.txt instructions do very little to stop these bots. The median European site's robots.txt instructions to not scrape were ignored 2.8x more often than North American sites'.
5
The methods these scrapers use to evade detection are the same ones used to access infrastructure or manipulate social media by nation-state hackers and sophisticated cyber-criminals. Now, these tools are used by venture-backed Silicon Valley companies serving major corporate clients — without intervention, the volume will compound exponentially as the companies behind it scale. A majority of these scraping companies advertise using methods such as residential proxies, the same functionality used by bad actors in cyberattacks, to gain access to sites and evade their cybersecurity measures.
6
Bots should not be allowed to mimic humans or other traffic on the web; they should be required to self-identify. As these disguise techniques grow more sophisticated, bots become harder to detect, placing an undue burden and expense on the sites. Self-identification would let publishers and sites decide how to respond to bot visitors.
Section 1: The Scraper Ecosystem Powering AI Applications & How Proxy Networks Work
Key Insights
The web scrapers that power AI tools are using sophisticated methods to conceal the identities of the bots they use for scraping using many techniques. Using our Scraper Audit tool, we can observe how effective scrapers are in getting access to content, and to an extent, the techniques they use to gain access:
In testing on sites in the UK, we saw similar trends to those we observed on sites in the US: paywalled content was not necessarily safe from web scrapers, and 19 out of the 20 top UK sites we tested could be scraped by at least one of the 14 scrapers we tested.
In testing on various publisher sites, we observed a number of techniques that scrapers used to get access to publisher sites, including rotating IP addresses and masquerading as other user agents, including Googlebot.
We also noticed Google Ads bots could be triggered on sites, a concern for advertisers who are paying for their ads to be seen by humans, and not scraper bots.
One of the most concerning methods scrapers are using is residential proxies, the same technique bad actors use to mask cyberattacks. Residential proxies are a potential national security concern as they make it virtually impossible to know who is on the other end, operating a bot. A majority of the web scrapers in our Scraper Index openly advertise that they use residential proxies.
AI developers are combining multiple data acquisition processes and technologies, many offered by specialist web scraping vendors, into tech stacks that obtain access to the Internet. Fueled by the AI boom, hundreds of these third-party scrapers are in use, many by Fortune 1000 companies.
These intermediaries scrape publisher content to resell it as data or search infrastructure, rather than to directly power their own chatbots or AI applications. Many of these scrapers employ specialized cybersecurity evasion tools and do not comply with robots.txt by default. This places an undue burden on sites to attempt to detect these increasingly hard-to-detect scrapers.
The data acquisition stack has existed for quite some time, but its effects have become more visible recently. With AI, the need to access content & data has exploded. Funded by the AI investment boom, developers are combining multiple data acquisition processes and technologies, with more advanced or clandestine approaches deployed where cybersecurity barriers are encountered. These increasingly complicated tech stacks are intended to secure "just enough" access to content to ensure the applications they serve can handle the range of queries needed to compete. These stacks are comprised of some or all of the following:
Direct crawling and scraping via first-party bots, either declared or disguised.
Secondary source crawling / scraping, targeting surfaces on which the content itself may be visible on the open web (i.e., Google Search results page).
Residential IP proxies, presenting as ordinary user traffic — often with location and device characteristics/IP addresses that match the human audience.
Circumvention services that evade bot detection / IP protection measures.
Cloud-based headless browsers that load web content in a browser running in the cloud, rather than on a user's device. These are often used with a residential IP proxy: the proxy makes the request appear to come from a normal human user.
Third-party scraping services for outsourcing the end-to-end process of scraping a page.
For more detailed information, please visit our Scraper Index.
For obvious reasons, this environment is one in which it is exceedingly difficult for digital publishers to exercise control over the use of their IP. They are in a structurally disadvantaged position, having to respond to new threats as they emerge, whereas the web scrapers just need to evolve beyond the latest protection. Most websites deploy only one cybersecurity tool, whereas AI companies can choose from dozens of scraping tools interchangeably.
Two notable additions here to our prior reporting on the AI Data Acquisition stack:
First is the increase in attention we've seen to residential proxies, along with the potential national security risk associated with them. A residential proxy routes traffic through everyday consumer devices so that the request appears to come from a human using an ordinary home internet connection. This technique is utilized by many web scrapers — and this is the same exact technique that bad actors use to conceal who they are in cyberattacks and large-scale cognitive warfare, such as astroturfing campaigns. Recently, there has been a lot of attention to this in the Wall Street Journal reporting, with an article here on this topic, and a recent separate piece here with a view of an FBI director stating "we all want to sort of find and defeat these Residential Proxy networks," because of their association with crime. Those efforts are already underway. In July 2026, Google, working with the FBI and others, disrupted NetNut, one of the largest residential proxy networks, which it estimated ran on more than 2 million hijacked consumer devices such as smart TVs and streaming boxes. Google reported that in a single week, it observed hundreds of distinct threat groups, including cybercriminal and espionage actors, using the network to hide the origin of their traffic. The Google blog post about this also linked out to a study by a cybersecurity firm that noted many of the largest proxy providers' websites highlight their utility for AI platforms and tools, "implying it is a primary use case for their residential proxies."
In an analysis of nearly 40 scraping vendors, we find that many explicitly advertise cybersecurity detection evasion techniques, and many do not default to abiding by robots.txt. A majority of the web scraping companies in our Web Scraper Index openly advertise that they use residential proxies and have pools of hundreds of millions of IP addresses. Using residential proxies appears to be a fairly standard practice for scrapers today. Residential proxies make it even more difficult for websites to detect, and therefore block, allow, or monetize bot/scraper traffic. In a site's logs, a bot visitor using a residential proxy will look just like a human visitor.
Another addition to our prior reporting is some data on third-party scrapers/data resellers across the TollBit network, as we have started to notice some of their traffic actually starting to stand out in our log data across publisher sites. (Important to note, many of these third parties do not self-identify their user agents, but a few do, and so we can analyze their data). Across all sites on TollBit in H1 2026, we observed that Common Crawl, the open corpus behind much LLM training, reached the widest share of any identifiable third-party scraper/crawler. We observed Common Crawler scrapes on ~74% of sites on TollBit: it roughly doubled its reach quarter-over-quarter (from Q1 to Q2). Because a single Common Crawl scrape can support many models and AI applications downstream, its footprint across sites matters far more than its per-site volume: one scrape from Common Crawl may then become data that then gets reused to power a vast AI ecosystem.
Two other notable players that scrape content and resell it as a search or data API: Parallel, a web-search/data API, which had scrapes on 74% of sites on TollBit, nearly doubled its per-site scraping (+94%), and Diffbot (which sells a structured "knowledge graph" of the web) roughly tripled per-site (+195%) and has scrapes on over 33% of sites on TollBit. Because a reseller's crawl can be resold to many buyers, these figures understate true exposure — one scrape resurfaces across multiple AI products, strengthening the case for licensing terms that follow the data downstream.
1.1 Scraper Audit Tool & Results on UK Sites
In order to examine the sophistication of web scrapers against websites' cybersecurity tooling, TollBit developed the Scraper Audit, a tool for website owners to test 14 popular web scrapers, many of which are Europe-based companies, against their own sites. The Scraper Audit allows website owners to initiate scraping on any page to evaluate how well scrapers evade defenses, access content, and retrieval completeness.
In selecting web scrapers, we picked a few industry leaders and a couple of newer scrapers offering advanced technology, despite their company size and time in market, and have plans to add more to the tool. For each of these web scrapers, whenever available, we purchased subscriptions of their web scraping APIs to leverage their full suite of AI-driven parsing, advanced anti-bot bypassing, and premium proxy networks.
To demonstrate the utility of this tool in auditing scrapers, we conducted an evaluation of this tool across the top 20 sites based in the UK as defined by the Press Gazette.
Our test of these sites resulted in similar outcomes to the tests run for the previous edition of the report:
Even paywalled content is not necessarily safe. In our testing of 14 of these scraping tools, we observed that some were able to scrape various paywalled articles in full.
In fact, non-paywalled sites offered no discernible increase in protection compared to their paywalled counterparts, as scrapers successfully retrieved full content from the vast majority of target pages regardless of subscription barriers.
In the rare instance where all scrapers failed to fully extract data — occurring in only one of the twenty sites — the failure was primarily attributed to server-side rendering implementations of the paywall. On these specific sites, the full article was not present in the initial page and required server-side authorization before the content was served.
1.2 Scraper Audit: Other Observations
Since our last bot report, we've added functionality to our Scraper Audit Tool to allow publishers to observe the suspected chain of events that occurs as a web scraper attempts to access their sites. In the tool, publishers can enter a specific URL in the Scraper Audit Tool and observe some of the strategies the web scrapers utilize in an attempt to scrape the page. In sum, we observed that scrapers often deploy various tactics to obtain access to sites, some hammering sites dozens of times with different user agents, sometimes masquerading as other bots (including Googlebot, Chrome, etc.), and others rotating IP addresses in order to attempt to trick the site's cybersecurity and allow them to gain access to scrape the page. We also observed that some of the web scrapers would trigger Google Ads bot on sites — concerning for advertisers who are paying for their ads to be seen by humans, and not bots/web scrapers.
A related note: In our last bot report, we noted that we did not observe any scrapers checking robots.txt files before attempting to access sites.
The strongest signal (Canary verification) we have is that the Scraper Audit adds a unique ID to the request of each URL we tell the webscraper to scrape, and we then look for that unique ID in the respective site's logs so that we can identify which scraper visited the URL. When we find it, we can be highly confident the request was triggered by the scraper audit test we ran.
The second signal, which we treat as lower confidence, is fingerprint matching. Some requests that hit the server in response to our scrape don't carry the unique ID. For those, we build a fingerprint from the request's attributes (its timestamp, IP address, geo, and user agent — all of which we can obtain from the respective site's logs) and match it against the scrape we initiated. A request that lines up across all of these — arriving at the same moment, from the same IP and geo, with the same user agent — points strongly to the scraper, even without the ID to confirm it. We are relying on very close timing to make this assumption (typically sub-second to a few seconds after invocation), which we assume is related to the scraper test we run. We know when the scraper was triggered and what page it was triggered to, because we control when the scrapers are triggered and to which pages, and we have the publishers' site logs data that includes information like IP address, timestamp of the scrape to the site, etc.
Here are some real examples of what we observed with this new functionality on publisher sites. Each of the examples below shows a single scraper instance attempting to access a publisher's URL:
1.2.1 Example 1: Masquerading as Google & IP Address Rotation
The scraper in this example appeared to have tried more than one disguise to get in. It first used User-Agent Spoofing and posed as Googlebot to gain trusted access. When that was blocked, it appeared to switch to a Chrome browser, and its apparent location shifted from Los Angeles, US, to San Sebastián de los Reyes, ES, meaning it also changed its IP address. Once it changed to a browser, it was first served 301 redirects, then kept trying until it got a 200 and pulled the content. This showcases that when one identity is refused, a scraper may just change who it claims to be and where it appears to come from and try again.
Scraper masquerading as Google, request timeline
Scraper masquerading as Google, evidence detail
1.2.2 Example 2: Multiple Browser Impersonations & IP Address Rotation
This attempt to access a publisher's URL appeared to trigger 43 requests from a single web scraper, routed through many different IPs and locations, cycling between Chrome and Firefox identities throughout the attempt to access the site. The locations of the scrapes heavily varied: alongside US cities (San Jose, Ashburn, Clifton, Pennington), the scrapes arrived from multiple Hungarian locations (Sopron, Sümeg, Nyíregyháza) and Bahrain (Bu Kuwarah), all within seconds.
Multiple browser impersonations, request timeline
Multiple browser impersonations, evidence detail
1.2.3 Example 3: Multiple Browsers & IP Rotation
This scrape request appeared to result in 32 requests hitting the URL at essentially the same time, from many different IPs across scattered US cities and cycling between Chrome and Firefox identities throughout. One scrape is arriving from a dozen-plus origins at once, and being disguised on both the network and identity side. The status column is a real mix of 200s (served), 403s (blocked), and 301s (redirected), so the publisher's defenses were firing and worked some of the time, but the blocks didn't matter because each request came from a different IP. A 403 on one address did nothing to stop the next from a different IP, and the scraper kept rotating until enough 200s landed to assemble the content.
Multiple browsers and IP rotation, request timeline
Multiple browsers and IP rotation, evidence detail
1.2.4 Example 4: Scraper Triggers Google Ad Bots
The two Google Ads Bot requests are worth taking a look at here. They appear to have fired because the scraper rendered the page like a real browser, loading its ads, which is what set them off. The scrape looks human enough to trigger the ad stack; it could be taken as evidence that the scraper's disguise is convincing and successful. This presents an ad fraud concern for advertisers who are paying to see their ads viewed by humans, and not bots. (This is a topic we mentioned previously in our Q2 2025 Bot report, Section 1).
Scraper triggering Google Ad bots
TollBit has made the Scraper Audit tool available on the TollBit platform to help websites test and understand their vulnerability to IP threats from major web scrapers. If you are interested in seeing how your site is being scraped by these third parties, please contact team@tollbit.com.
1.3 The Inefficiency and Cost of Web Scraping
This arms race between scrapers and website operators is not only inefficient for website owners, who are increasingly spending on cybersecurity to detect bots that are only getting harder to detect, but it is also deeply inefficient for AI developers. As mentioned in our prior bot report, we have found advanced scraping services charging up to $22.50 for 1,000 pages. Given the volume of data needed to serve consumer AI applications, the data acquisition costs for the most popular chatbots and specialized agents are likely to run into the tens of millions of dollars a year. And this is before considering the legal fees needed to fight lawsuits.
We have also observed that scrapers take 16 seconds on average to access publisher sites, and some of them may fail entirely in securing access to the page requested, whereas accessing content via a sanctioned route (i.e. via a platform like TollBit) typically takes less than 0.7 seconds.
There's a better way. Instead of spending huge sums on technical workarounds to bypass IP controls, AI developers can simply pay for the content they use. This reduces latency and increases reliability. By relying on licensing rather than scraping, journalism is funded — not the scrapers — and the media ecosystem AI depends on is sustained.
1.4 The Limits of Cybersecurity and the Need for Regulation
We've seen publishers increasingly take matters into their own hands to combat the surge in AI bots, increasingly investing in cybersecurity to actively block this traffic from their sites or even take action by filing lawsuits. As mentioned in our Q3 & Q4 2025 report, in late 2025, both Reddit and Google turned to the courts over unauthorized scraping. Google alleged that scrapers had circumvented SearchGuard, the in-house protection system it built at a cost of millions. If a company with Google's resources and engineering talent cannot keep determined scrapers out, an individual publisher has no realistic chance of doing so alone.
Cybersecurity is important, but it is not a foolproof solution. These tools are costly to run, and many of our publishers, even those with the most sophisticated defenses in place, have seen scrapers and AI bots get through anyway.
This brings us to an important limitation of the data in the subsequent sections of this report: we rely on AI companies self-identifying through their user agents. Developers sometimes use third-party vendors to scrape on their behalf; they may masquerade as other user agents, rotate their IP addresses, use headless browsers that mimic humans, use residential proxies, etc. This activity is difficult, and often impossible, for sites to trace back to a specific AI company. The true level of scraping, and of unpermitted access to European content, is, therefore, likely greater than this report shows.
As AI traffic becomes harder to distinguish from human traffic, it is important that regulators take a position that, at the bare minimum, AI bots should not be allowed to mimic humans on the Internet. Ideally, they would be required to self-identify and state their intended use, which is what New York's Stealth Crawler Prohibition Act would require. That bill, passed by the New York State Legislature in June 2026 and awaiting the Governor's signature, would require crawlers accessing a news site to disclose their identity and purpose through an accurate user-agent string before they access the content.
Self-identification enables a site owner to decide how to respond: allow, block, or charge for access via infrastructure controls like the ones TollBit provides. It would help ensure that the undue burden of paying to serve, detect, and block unwanted bots no longer falls on the websites themselves.
Read the Full Report
Fill out the form below to read the our full report.