Reading 23

Web Crawler Restrictions, AI Training Datasets & Political Biases

Bouchaud and Ramaciotti, arXiv, 2025

Reflowable text extracted from PDF. Each page links to an embedded original page image for diagrams, equations, and layout. Text extraction may affect reading order or mathematical symbols.

PDF Page 1

Web Crawler Restrictions, AI Training Datasets

& Political Biases

Paul Bouchaud1,2, Pedro Ramaciotti1,2,3

1Complex Systems Institute of Paris Ile-de-France CNRS, Paris, France. 2médialab, Sciences Po, Paris, France.

3Learning Planet Institute, CY Cergy Paris University, Paris, France.

Abstract

Large language models rely on web-scraped text for training; concurrently, content creators are increasingly blocking AI crawlers to retain control over their data. We analyze crawler restrictions across the top one million most-visited websites since 2023 and examine their potential downstream effects on training data composition. Our analysis reveals growing restrictions, with blocking patterns varying by website popularity and content type. A quarter of the top thousand websites restrict AI crawlers, decreasing to one-tenth across the broader top million. Content type matters significantly: 34.2% of news outlets disallow OpenAI’s GPTBot, rising to 55% for outlets with high factual reporting. Additionally, outlets with neutral political positions impose the strongest restrictions (58%), whereas hyperpartisan websites and those with low factual reporting impose fewer restrictions —only 4.1% of right-leaning outlets block access to OpenAI. Our findings suggest that heterogeneous blocking patterns may skew training datasets toward low-quality or polarized content, potentially affecting the capabilities of models served by prominent AI-as-a-Service providers.

CCS Concepts

• Computing methodologies →Natural language processing; • Information systems →Web crawling; Web mining; • Applied computing →Law, social and behavioral sciences.

Keywords

large language models, web scraping, dataset bias, news media, political bias, CommonCrawl

1 Introduction

The training of large language models that powered the recent surge in AI chatbots crucially relied on massive text corpora [5, 37] —amounting to trillions of words scraped from billions of web pages

[8]. This large-scale data collection, and its subsequent commercial exploitation, has raised fundamental copyright concerns and prompted content creators to resist having their work used without compensation or attribution [13, 19, 34]. Following the training of their flagship model, OpenAI announced in August 2023 support for the Robots Exclusion Protocol [24], a standard initiated in the late 1990s that allows webmasters to specify which web crawlers can access their content by placing directives in a robots.txt file at the root of their websites. Following this development, several organizations and content creators called for blocking OpenAI’s and other AI providers’ crawlers [16, 17, 35] as a means to exert greater control over the use of their data.

Meanwhile, since the emergence of natural language processing, a wealth of work has addressed how biases in training corpora

might translate to biased outcomes. Gender and racial biases [6, 21, 31] in word embedding models such as word2vec [22] and GloVe [28] and subsequent sentence-level embeddings such as BERT [9], were traceable to biases in the underlying training corpus [4, 12]. However, the increased size of modern model representations and datasets, coupled with opacity regarding training data composition, impedes such investigations. Nonetheless, one resource remains central to most AI training efforts [1, 5, 37]: CommonCrawl [8].

CommonCrawl is a petabyte-large archive of web crawl data collected since 2008 [1, 8]. Due to its size and coverage, it is massively used in the training of AI generative models, in particular through derivative datasets. Indeed, CommonCrawl collectors have taken a stated aim not to curate or polish the content of the collection, capturing a fraction of the web as it is [1]. Consequently, large language model builders rely on filtered versions of CommonCrawl such as C4 [29], RefineWeb [27], or FineWeb [26] that typically remove duplicates, HTML boilerplate, and subsample certain languages [38]. Some curation rules apply normative judgments by filtering out content deemed “undesirable” [1, 20, 26] and have been shown to disproportionately exclude text from and about minorities [11], potentially skewing the representation in models trained on such filtered datasets.

In addition to content filtering, another potential source of bias may arise directly from the collection process itself, as AI crawlers face growing restrictions [10, 16]. If AI model builders do not properly address the skew induced by the protection measures content creators adopt against exploitative practices, the trained AI models may inherit systematic biases that may compromise their fairness and representativeness. In this work, we quantify and characterize the adoption of AI crawler restrictions since 2023 across the one million most visited websites worldwide, and evaluate the downstream effects on CommonCrawl and its derivative training datasets.

Our analysis reveals growing AI crawler restrictions since 2023, with blocking patterns varying by website popularity and content type. Over 25% of the top thousand most visited websites restrict AI crawlers, but this proportion decreases to 10% when considering the broader top million websites. Content type also plays a significant role: while only 4% of shopping websites disallow OpenAI’s GPTBot, 34.2% of news outlets do, rising to 55% for outlets with high factual reporting. Outlets with neutral political positions impose the strongest restrictions, with 58% disallowing OpenAI’s crawler, whereas hyperpartisan websites and those with low factual reporting impose fewer restrictions—only 4.1% of right-leaning outlets block access.

Additionally, we observe that the likelihood for a website to impose AI crawler restrictions decreases as its audience ideology diverges from the political center, with websites on both the farleft and far-right ends of the ideological spectrum imposing fewer

arXiv:2510.09031v1 [cs.SI] 10 Oct 2025

PDF Page 2

Bouchaud and Ramaciotti

AI crawler restrictions than moderate sources. Paralleling the increased restrictions on AI crawlers imposed by moderate sources, we observe their relative decline within training datasets in favor of hyperpartisan material. Through textual analysis of political content, we characterize the potential biases this shift may introduce, revealing over-representation of politically charged terminology in available training data.

These findings call for acute consideration in the curation of future training datasets to account for skewed restrictions in raw data collection, as the systematic withdrawal of high-quality news content, coupled with the relative over-representation of politically hyperpartisan sources, demands active intervention by model developers to prevent embedding polarization into AI systems.

2 Datasets

In this article, we examine restrictions on AI crawlers deployed by the most popular websites worldwide. To this end, we rely on Google Chrome’s analytics to identify popular websites and on Cloudflare Radar to categorize webpage content. Additionally, we combine a large-scale collection of robots.txt directives left by webmasters as of September 2025 and historical data since 2023 from CommonCrawl snapshots.

Websites Dataset (CrUX). As a window into the web, we first consider in this study the set of the one million most visited websites worldwide accounting, in August 2025, for 95% of global traffic on Google Chrome in terms of page loads [32]. This dataset stems from the Chrome User Experience Report (CrUX) and represents aggregate analytics from hundreds of millions of Chrome users who have usage statistic reporting enabled, opted in to history syncing, and have no history sync passphrase set. Ruth et al. [33] performed comparative analysis of top lists, including the (deprecated) Alexa list and Cisco’s Umbrella, and showcased the high coverage of the CrUX dataset. In addition to the mere list, we will leverage the popularity buckets disclosed in CrUX (top 1K, 10K, 100K, 1M). Unless stated otherwise, all analyses are performed over this set of websites or subsets thereof.

Content Categorization. We enrich this list of one million websites with a characterization of their content. To this end, we leverage, akin to Ruth et al [32], Cloudflare’s Radar APIs [7]. Cloudflare operates, beyond DDoS protection, a DNS and parental control service whose filters allow parents or schools to block domains classified as “Adult Themes” or more generally falling under CIPA filters. To enable such filtering, Cloudflare characterized websites content —thus extending beyond Cloudflare-protected websites—, we then rely on this categorization. For simplicity, we aggregate Cloudflare’s granular content categories into “supercategories” — we will drop the prefix for discussion fluidity— relying on the taxonomy, reported in the Appendix, curated, and assessed in [32].

To validate Cloudflare’s categorization, we extracted HTML meta keywords from 661k successfully crawled websites and identified overrepresented terms per category using chi-squared statistics. For instance, pages categorized as Shopping & Auctions tend to have keywords such as: silver, watches, jewellery, while those categorized as Adult Themes have: porn, telegram, sex; we display the top words of each category in the Appendix.

Robots.txt Collection. Additionally, we focus on the directives set by webmasters to crawlers scanning the web, indicating to them which portions of the website, if any, they are permitted to access. In September 2025, for each of the top one million websites we fetched the robots.txt file deployed, if any, at the root of their (sub)domains. Additionally, to go beyond a mere snapshot, we relied on CommonCrawl’s periodic releases [8] of very-large-scale crawls of billions of webpages. In particular, this corpus contains timestamped robots.txt files, offering unparalleled chronological insights. However, rather than being a “copy of the web”, only a fraction of it is crawled [1, 8]. Within the top one million websites, on average over releases, 409k valid robots.txt files were collected.

3 Overview 3.1 Content Category

We first aim to provide an overview of the web cross two dimensions: topic categories and languages, updating the description made by Ruth et al. [32] in early 2022. We observe that half of the pages are written primarily in English (50.5%), followed by Japanese (6.2%) and Spanish (6.0%), see figure in Appendix. Second, we observe in Figure 1 that Shopping & Auctions websites account for the largest fraction at 18.6%. Business & Economy, Education, Entertainment and Technology follow, representing between 7.6% and 10.5%. Aligned with [32], we observe that Adult content shows significant variation across popularity tiers. Within the top thousand most visited websites worldwide, 20.3% are Adult websites, compared to 15.0% in the top 10k and 3.7% in the top 1M.

Shopping & Auctions

Business & Economy

Education

Entertainment

Technology

Other

News & Politics

Society & Lifestyle

Travel Government & Politics

Adult Themes

Health

Sports

Vehicles

Gambling

Real Estate

Internet Communication

Job Search & Careers Questionable Content Religion Violence

Weather 1M Most Visited Websites by Content

Figure 1: Distribution of content categories within the one million most visited websites worldwide (CrUX).

3.2 Crawling Restrictions

To prevent the content of their websites from being automatically indexed, scraped, and otherwise programmatically used without the website owner’s agreement, multiple mechanisms have been developed to prevent such crawling.

Illustration from source page 2
Illustration from source page 2.

PDF Page 3

Web Crawler Restrictions, AI Training Datasets & Political Biases

In the top one million websites, we observe that 55.6% of them deployed so-called robots.txt at the root of their (sub)domains.

Alternative proposals to prevent the crawling to feed AI model such as adding noai in the “robots" meta tag of the page appeared on less than 0.2% of analyzed websites, similarly [10] observed low adoption of the “TDM Reservation Protocol”. In addition to such passive voluntary systems, Liu et al. [16] explored active measures whereby websites restrict the activity of some public crawlers, for instance by blocking traffic coming from IP addresses known to be operated by crawlers. Nevertheless, robots.txt is the de facto standard to specify directives to web crawlers and as such will be further examined in the following section.

4 Robots Exclusion Protocol

With the advent of generative artificial intelligence, the datasets on which these models are trained have garnered attention and warranted renewed inspection of the Robots Exclusion Protocol [10, 14, 16]. We first consider a large-scale characterization of directives left by webmasters for crawlers before focusing specifically on those used by AI model providers.

4.1 Overview

The robots.txt can specify rules applying uniformly to all crawlers using the catch-all directive (Allow/Disallow for User-Agent: *) or by specifying targeted rules for particular crawlers; for instance, allowing Google Search’s crawler to index certain pages while blocking all other robots. Within our collected sample, 64.1% of robots.txt files used only the catch-all rules.

Among crawlers named in robots.txt files of websites specifying granular directives, we can delineate the following categories:

• SEO & Marketing: MJ12bot (36.5%), AhrefsBot (36.0%), AhrefsSiteAudit (23.5%), and SemrushBot (13.9%) • Search Engines: Nutch (25.1%), GoogleBot (18.1%), CCBot (14.7%), Yandex (10.8%), and bingbot (6.8%) • Social Media: Pinterest (23.6%), FacebookBot (5.5%), and facebookexternalhit (5.0%) • Generative AI Training: GPTBot (19.9%), ClaudeBot (14.7%), Google-Extended (14.4%), Applebot-Extended (10.2%), and meta-externalagent (8.8%) • AI Search & Chatbots: ChatGPT-User (7.5%), PerplexityBot (6.5%), Claude-Web (5.2%), OAI-SearchBot (4.9%)

Alongside crawlers indexing the web for analytics, search engines, or generating social media previews, we observe AI-related ones. Granular distinctions emerge between crawlers collecting data for model training, e.g., GPTBot, ClaudeBot, those powering AI search engines, e.g., OAI-SearchBot, PerplexityBot, and those responding to individual user queries, e.g., ChatGPT-User, Claude-Web.

Importantly, the Robot Exclusion Protocol relies on good faith and voluntary compliance by the organizations operating crawlers. In 2025, Liu et al. [16] and Kim et al. [15] showed that AI crawlers operated by OpenAI, Anthropic, Meta, and CommonCrawl respected websites’ robots.txt directives as stated. Yet, it should be noticed that the scope of these directives can be ambiguous. For instance, if a chatbot provider trains on user discussions without further curation, content from a website that blocked GPTBot but allowed

(or, due to tacit consent, did not block) ChatGPT-User could still be used for training as it may appear in user chats.

4.2 AI Crawlers

2025 Snapshot. Focusing on the most popular AI chatbots [36], we observe that, as of August/September 2025 and among the top one million websites, OpenAI’s GPTBot is fully blocked on 10.6% of websites, Anthropic’s ClaudeBot on 9.1%, and Google’s Google-Extended on 8.9%. Additionally, 9.5% disallow CCBot, the crawler used by CommonCrawl, whose public datasets are extensively leveraged in training AI models [1, 5, 37].

Alternatively, AI-powered search engine crawlers are more frequently permitted, with only 6.0% of websites disallowing Perplexity’s crawler, and 6.5% and 6.1% blocking OpenAI’s and Anthropic’s user-initiated crawlers (ChatGPT-User & Claude-Web). As a baseline, 4.0% of examined websites blocked Google Search’s indexing crawler GoogleBot at their root.

We observed instances of websites, particularly news outlets, that disallowed popular AI crawlers except one due to licensing agreements [23, 25]. Examples include the Wall Street Journal, New York Post, The Times, Bild, Washington Post, and Le Monde, all allowing OpenAI’s crawlers while explicity disallowing others.

Jan 23 Apr 23 Jul 23 Oct 23 Jan 24 Apr 24 Jul 24 Oct 24 Jan 25 Apr 25

2%

3%

4%

5%

6%

7%

Disallowed by robots.txt

Googlebot

GPTBot

CCBot Claudebot Google-Extended ChatGPT-User

Claude-Web PerplexityBot

Figure 2: Fraction of webpages disallowing specific web crawlers on their domain over time, underlying data from CommonCrawl releases.

Chronology. Relying on CommonCrawl’s periodic releases [8], we can reconstruct the evolution of robots.txt since January 2023. We observe in Figure 2 that while the fraction of webpages blocking Google Search’s indexing crawler GoogleBot remains stable over the period 2023-2025, the fraction blocking OpenAI’s GPTBot rises after its announcement in summer 2023 [24]. Following their public launch, other AI crawlers such as ClaudeBot and PerplexityBot experienced growing disallowance rates in robots.txt files throughout our observation window.

Restrictions per category. Finally, we examine how robots exclusion directives are used across different categories of website content and popularity. Relying on Cloudflare categorization, Figure 3 reveals that disallowance of OpenAI’s GPTBot varies by topic across websites, with rates exceeding 14% for Entertainment and News & Politics, declining to 9.0% for Education, and reaching only 4.0% for Shopping & Auctions. Additionally, we observe that more popular websites are more likely to disallow OpenAI’s GPTBot: while 13.4% of websites in the top 100k block it, this proportion nearly doubles to 25.2% among the top thousand most visited websites worldwide.

Illustration from source page 3
Illustration from source page 3.

PDF Page 4

Bouchaud and Ramaciotti

Entertainment News & Politics Education Shopping & Auctions Website Category

0%

2%

4%

6%

8%

10%

12%

14%

Disallowed by robots.txt

1.2±0.1%

2.6±0.1%

5.3±0.2%

1.0±0.1%

14.4±0.3% 14.1±0.3%

11.1±0.3%

3.9±0.1%

6.2±0.2%

12.3±0.3%

9.0±0.2%

4.0±0.1%

2.9±0.1%

7.9±0.2%

6.9±0.2%

3.1±0.1%

Googlebot GPTBot CCBot PerplexityBot

Top1K Top5K Top10K Top50K Top100K Top500K Top1M Website Popularity

0%

5%

10%

15%

20%

25%

Disallowed by robots.txt

2.8%

25.2%

2.2%

19.2%

2.8%

17.1%

3.5%

15.0%

3.8%

13.4%

4.2%

11.1%

3.9%

10.6%

GPTBot Googlebot

Figure 3: Fraction of webpages disallowing specific web crawlers on their domain, segmented by: (A) Content category per Cloudflare’s classification, (B) CrUX popularity bucket. Error bars represent standard deviations over 100 bootstrap with replacement over websites.

5 CommonCrawl & Derivatives

Since the observed directives that restrict crawlers from collecting training data for generative AI models are not uniformly distributed across the web, and because biases in training datasets can lead to discriminatory outputs in AI models [2, 3, 12], we evaluate how the growing restriction of online resources impacts CommonCrawl and derivative datasets.

Indeed, refined datasets of CommonCrawl have emerged as fundamental in the development of recent large language models, representing over 80% of training corpora of Llama [37] and GPT-3 [5]. We will in particular consider two subsets of CommonCrawl: FineWeb and FineWeb-Edu [26]. FineWeb is a cleaned and deduplicated English subset of CommonCrawl, while FineWeb-Edu is a subset of FineWeb containing English text with high “educational quality” [26]. Specifically, FineWeb-Edu was created by first scoring half a million FineWeb text snippets using a large language model, then by training a regression head over text embedding representations to extend this scoring to hundreds of millions of text snippets. Large language models trained on FineWeb-Edu were shown by HuggingFace in June 2024 to “outperform all openly accessible webbased datasets on a number of reasoning- and knowledge-intensive benchmarks such as MMLU” [26].

5.1 CommonCrawl

First of all, to provide an overview of CommonCrawl’s content, we consider a random sample of one billion tokens from its latest dump (CC-MAIN-2025-38) and analyze the language and content of the underlying crawled pages, relying on Cloudflare categorization [7]. Overall, aside from an under-representation of Adult content compared to CrUX pages and an over-representation of pages in German, CommonCrawl appears analogous to the web in terms of language and content category; see the distribution in the Appendix.

5.2 News Outlets

Beyond this overall characterization, we now focus on news outlets. We relied on political skew and factual reporting assessments from Media Bias/Fact Check1. We collected the robots.txt files of 3 668 annotated news outlets. Overall, 34.2% of media outlets disallowed OpenAI’s GPTBot, with strong disparities across political leaning and factual reporting. As displayed in Figure 4, 58.0%

1https://mediabiasfactcheck.com

of neutral outlets block OpenAI’s GPTBot, compared to 19.6% of left-leaning outlets and 4.1% of right-leaning outlets. Regarding factual reporting, 55.4% of outlets with high factual reporting block OpenAI crawlers, compared to 8.4% of those with mixed factual reporting and 3.7% of those with low factual reporting. Additionally, we observe that across political leaning and factual reporting categories, the fraction of websites blocking training-related crawlers (GPTBot) is higher than those used to answer specific user requests (ChatGPT-User), as displayed in Figure 4.

Left Left-Center Neutral Right-Center Right Political Leaning

0%

20%

40%

60%

Disallowed by robots.txt

Low Mixed High Factual Reporting

GPTBot ChatGPT-User

Media Bias / Fact-Check

Figure 4: Fraction of websites disallowing OpenAI’s GPTBot and ChatGPT-User, as a function of their political skew and factual reporting assessment by MBFC.

In parallel, comparing the first Fineweb release from 2020 with that from 2025 (both larger than 170 billion tokens), the fraction of tokens stemming from sources evaluated as “high factual reporting” by MBFC decreased from 0.46% to 0.26% (a 41.3% relative decrease). Similarly, the fraction of hyperpartisan right content increased from 19.3% to 24.8% (a 28.8% relative increase), and of hyperpartisan left content rose from 18.9% to 22.4% (an 18.6% relative increase).

5.3 Educational Content

Similarly, for educational resources, we examine FineWeb-Edu, a subset of FineWeb curated for its “educational quality” [26]. Importantly, each snippet in FineWeb-Edu is associated with the exact URL from which the content was extracted and its collection date by CommonCrawl’s CCBot. This URL-level metadata enables us to assess robots.txt directives at the specific URL level rather than merely at the broader domain level.

For each release since 2015, we randomly sampled 25k URLs from domains in the top one million. We observe that over 20% of the tokens in the FineWeb-Edu dataset collected prior to 2024 originated from domains that, as of August 2025, block CommonCrawl bots from accessing those very resources; see Figure in Appendix.

We preemptively seek to address the hypothesis that such restrictions on high-quality educational content may lead to overall decay in training dataset educational quality. For every FineWeb release since 2015, we randomly sample 10k text snippets and assess their “educational quality” using the language model used by HuggingFace to curate FineWeb-Edu [26]. The output is a score between 0 (not educational) and 5 (highly educational). We do not observe quality decay. Instead, the fraction of FineWeb (in tokens) scoring at least 3 (threshold to be included in FineWeb-Edu) grew from 8.3% on average prior to 2022 to 9.7% since 2022; see Figure in Appendix.

Illustration from source page 4
Illustration from source page 4.
Illustration from source page 4
Illustration from source page 4.

PDF Page 5

Web Crawler Restrictions, AI Training Datasets & Political Biases

5.4 Political Content

Finally, we explore the directives left by webmasters to crawlers as a function of websites’ political leaning. To assess such leaning, we rely on the large-scale audience characterization performed by Robertson et al. [30]. Specifically, they linked half a million Twitter accounts to US voter registration records and collected the URLs these accounts shared on the platform. The leaning of each domain was then assessed as the fraction of registered Democrats versus Republicans who had shared a URL from that domain, resulting in a continuous scale ranging from -1 to 1. A score of -1 (respectively +1) indicates that the domain was shared exclusively by Democrats

(respectively Republicans), while a domain receives a score of 0 if and only if it was shared by equal proportions of Democrats and Republicans. This scale has been shown to align with other estimates derived from Facebook data, expert raters, and community assessments [30].

1.00 0.75 0.50 0.25 0.00 0.25 0.50 0.75 1.00 Political Leaning

0%

5%

10%

15%

20%

Fraction of websites

disallowed by robots.txt

GoogleBot

GPTBot

Left leaning Right leaning

Figure 5: Fraction of websites disallowing GPTBot as a function of their audience ideological leaning, as characterized by Robertson et al [30]. Error bars represent Clopper-Pearson 95% confidence intervals. The solid curve shows the fitted quadratic logistic regression with 95% confidence band (shaded region). The fraction disallowing GoogleBot is shown as a baseline.

Robot Exclusion. We collected the robots.txt files of 15.2k domains ideologically scaled by Robertson et al. [30] and display in Figure 5 the fraction of websites that disallow OpenAI’s GPTBot as a function of their audience ideological leaning, with GoogleBot as a baseline.

We observe that domains shared in similar proportions by Democrats and Republicans —those with low absolute scores— disallow OpenAI’s GPTBot significantly more frequently than ideologically skewed domains. While 18.5% of balanced domains (absolute score2 below 0.5) block OpenAI’s crawler, this fraction plummets to 7.3% and 7.0% for hyperpartisan left (≤−0.5) and hyperpartisan right (≥0.5) domains, respectively. Similar results holds for CommonCrawl’s CCBot. In contrast, in September 2022, the fraction of websites disallowing CommonCrawl’s CCBot was equivalent between outlets with balanced and skewed audiences, at 2.6% and 3.0% respectively.

2Threshold of ±0.5 set to align with AdFontes’s labels.

Representation in FineWeb. Without inferring causation, we observe that the fraction of tokens extracted from hyperpartisan domains in FineWeb [26] increased from 27.7% to 31.0% between 2020 and 2025 (a 12.1% relative increase).

To characterize potential biases that may arise from increased hyperpartisan content representation in FineWeb, we analyze word co-occurrence patterns in text authored by those websites. Our focus on co-occurrence rather than individual word frequencies aligns with large language model attention mechanisms, which are designed to capture contextual relationships between words.

Specifically, from FineWeb latest release (CC-MAIN-2025-26) we filter text snippets from domains with characterized political leanings [30]. We split them into sentences and randomly sample five million sentences for each group: hyperpartisan left, hyperpartisan right, and balanced domains. Subsequently, we compute the co-occurrence within the same sentence of each word pair and apply chi-square statistics to identify word pairs that co-occur in the same sentence significantly more often in hyperpartisan text than in balanced content.

Among hyperpartisan left-leaning text, the following word pairs are among the most overrepresented compared to balanced text (arbitrary word order within pairs): (Climate; Change), (Human; Rights), (Grants; NIH), (Union; Workers), (Student; Yale), (Indigenous; Communities), (Women; Rights), (Women; Abortion), (People; Color), (Jewish; Community).

Similarly, among hyperpartisan right-leaning text: (President; Trump), (President; Biden), (AFA; Family), (Jesus; God), (Donald; Trump), (Christ; Jesus), (Fox; News), (Pro; Life), (Faith; God), (Supreme; Court).

For comparison purposes, we identify words that co-occur more frequently in hyperpartisan text alongside terms referring to potentially discriminated population:

• Young: (Left) People, Women, Black; (Right) God, Church, Men • Old: (Left) Black, Scholar, Jewish; (Right) Testament, God, Biden • Woman: (Left) Black, White, Trans; (Right) God, Abortion, Man • Man: (Left) Black, White, Trump; (Right) God, Jesus, Trump • Transgender: (Left) People, LGBTQ, Rights; (Right) Children, Women, School • LGBTQ: (Left) Tennessee, Rhetoric, Safe; (Right) Children, Agenda, Christian • Immigrant: (Left) Republicans, Border, Detention; (Right) Illegal, Border, Trump • Muslim: (Left) Hindu, United, Americans; (Right) Islamic, World, Brotherhood

6 Discussion

The web’s scale and diversity have established it as a foundational resource for training the AI models that power the development of foundation models. However, our analysis of the one million most visited websites worldwide reveals substantial and growing adoption of AI crawler restrictions, with over 25% of the top thousand most visited websites blocking OpenAI’s web crawler. Importantly, these restrictions are not uniformly distributed across the web;

Illustration from source page 5
Illustration from source page 5.

PDF Page 6

Bouchaud and Ramaciotti

instead, they vary significantly by website popularity and content type, leading to potential skewing in subsequent web-crawled datasets. High-quality factual news outlets exhibit particularly pronounced restrictions, with over 55% blocking of OpenAI’s GPTBot.

Furthermore, we observe that outlets with high factual reporting and neutral political positions impose the strongest restrictions on OpenAI’s crawlers, whereas hyperpartisan websites and outlets with low factual reporting impose fewer restrictions. Quantitatively, while 58% of neutral outlets restrict OpenAI’s GPTBot only 4.1% of right-leaning outlets do so.

This pattern extends beyond news media to websites shared online more generally. Using Robertson et al.’s characterization of over 15k websites shared on Twitter [30], we find that hyperpartisan websites—both left- and right-leaning—impose significantly fewer restrictions than politically balanced sources. While 18.5% of balanced domains block OpenAI’s GPTBot, only 7.3% and 7.0% of hyperpartisan left and right domains do so, respectively. Correspondingly, the representation of hyperpartisan domains in CommonCrawl-derived training datasets increased by 54.6% (left) and 43.5% (right) between pre-2023 and post-2023 periods, while balanced domains decreased by 11.3%. Textual analysis reveals overrepresentation of politically charged, civic, and religious terminology in hyperpartisan content, with distinct semantic associations emerging for terms referring to marginalized groups.

These patterns raise concerns about systematic biases emerging in training datasets used to develop large language models. The withdrawal of high-quality news content and moderate political sources, combined with a relative increased representation of hyperpartisan material, may affect model outputs in ways that remain difficult to audit given the scale and opacity of training data compositions.

Importantly, our analysis focuses exclusively on robots.txt directives, the de facto standard for programmatically communicating crawler restrictions. However, this represents only one form of access control. Longpre et al. [18] demonstrated that when considering platform Terms of Service —written in natural language rather than machine-readable format— the fraction of C4, another large-scale CommonCrawl derivative, restricted from collection rose to 45%, compared to only 5% based on robots.txt alone. This suggests our findings may even underestimate the true extent of restricted content and the magnitude of potential biases introduced through selective data availability. Additionally, our focus on two English datasets, FineWeb & FineWeb-Edu, while reflecting the emphasis given to English in AI model training [5, 37], may overlook some biases, ignoring a potential education quality decay for nonEnglish text for instance, and should be further investigated in the future.

The heterogeneous adoption of crawler restrictions, designed to protect content creators from exploitative practices, demands careful attention from AI model builders, who should undertake active curation strategies beyond mere boilerplate and heuristicbased content removal. Our work points to one potential approach: by establishing the distribution of crawling allowances, sampling strategies may be adjusted to preserve the representativeness of training datasets. Future research should examine how these compositional shifts affect downstream model performance, particularly

regarding fairness, factual accuracy, and political neutrality in AI systems.

References

[1] Stefan Baack. 2024. A Critical Analysis of the Largest Source for Generative

AI Training Data: Common Crawl. In The 2024 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’24). ACM, 2199–2208. doi:10.1145/ 3630106.3659033 [2] Christine Basta, Marta R. Costa-jussà, and Noe Casas. 2019. Evaluating the

Underlying Gender Bias in Contextualized Word Embeddings. In Proceedings of the First Workshop on Gender Bias in Natural Language Processing, Marta R. Costajussà, Christian Hardmeier, Will Radford, and Kellie Webster (Eds.). Association for Computational Linguistics, Florence, Italy, 33–39. doi:10.18653/v1/W19-3805 [3] Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret

Shmitchell. 2021. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (Virtual Event, Canada) (FAccT ’21). Association for Computing Machinery, New York, NY, USA, 610–623. doi:10.1145/3442188.3445922 [4] Tolga Bolukbasi, Kai-Wei Chang, James Zou, Venkatesh Saligrama, and Adam

Kalai. 2016. Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings. arXiv:1607.06520 [cs.CL] https://arxiv.org/abs/ 1607.06520 [5] Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan,

Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language Models are Few-Shot Learners. arXiv:2005.14165 [cs.CL] https://arxiv.org/abs/2005.14165 [6] Aylin Caliskan, Joanna J. Bryson, and Arvind Narayanan. 2017. Semantics derived

automatically from language corpora contain human-like biases. Science 356, 6334 (April 2017), 183–186. doi:10.1126/science.aal4230 [7] Cloudflare. 2025. Domain categories. https://developers.cloudflare.com/ cloudflare-one/policies/gateway/domain-categories/ [8] Common Crawl Foundation. 2025. Common Crawl Dataset. https:// commoncrawl.org/. [9] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT:

Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805 [cs.CL] https://arxiv.org/abs/1810.04805 [10] Michael Dinzinger, Florian Heß, and Michael Granitzer. 2024. A Survey of Web

Content Control for Generative AI. arXiv:2404.02309 [cs.IR] https://arxiv.org/ abs/2404.02309 [11] Jesse Dodge, Maarten Sap, Ana Marasovic, William Agnew, Gabriel Ilharco, Dirk

Groeneveld, Margaret Mitchell, and Matt Gardner. 2021. Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus.. In EMNLP (1). Association for Computational Linguistics, 1286–1305. [12] Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov. 2023. From

Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Computational Linguistics, Toronto, Canada, 11737–11762. doi:10.18653/v1/2023. acl-long.656 [13] Michael M. Grynbaum and Ryan Mac. 2023. The Times Sues OpenAI and Mi-

crosoft Over A.I. Use of Copyrighted Work. The New York Times (27 December 2023). https://www.nytimes.com/2023/12/27/business/media/new-york-timesopen-ai-microsoft-lawsuit.html [14] Audrey Hingle and Mallory Knodel. 2025. Robots.txt Is Having a Moment: Here’s

Why We Should Care. TechPolicy.Press (April 3 2025). https://www.techpolicy. press/robotstxt-is-having-a-moment-heres-why-we-should-care/ [15] Taein Kim, Karstan Bock, Claire Luo, Amanda Liswood, and Emily Wenger. 2025.

Scrapers selectively respect robots.txt directives: evidence from a large-scale empirical study. arXiv:2505.21733 [cs.NI] https://arxiv.org/abs/2505.21733 [16] Enze Liu, Elisa Luo, Shawn Shan, Geoffrey M. Voelker, Ben Y. Zhao, and Ste-

fan Savage. 2024. Somesite I Used To Crawl: Awareness, Agency and Efficacy in Protecting Content Creators From AI Crawlers. arXiv e-prints, Article arXiv:2411.15091 (Nov. 2024), arXiv:2411.15091 pages. arXiv:2411.15091 [cs.HC] doi:10.48550/arXiv.2411.15091 [17] Shayne Longpre, Robert Mahari, Ariel Lee, Campbell Lund, Hamidah Oderinwale,

William Brannon, Nayan Saxena, Naana Obeng-Marnu, Tobin South, Cole Hunter, Kevin Klyman, Christopher Klamm, Hailey Schoelkopf, Nikhil Singh, Manuel Cherep, Ahmad Mustafa Anis, An Dinh, Caroline Chitongo, Da Yin, Damien Sileo, Deividas Mataciunas, Diganta Misra, Emad Alghamdi, Enrico Shippole, Jianguo Zhang, Joanna Materzynska, Kun Qian, Kush Tiwary, Lester Miranda, Manan Dey, Minnie Liang, Mohammed Hamdy, Niklas Muennighoff,

PDF Page 7

Web Crawler Restrictions, AI Training Datasets & Political Biases

Seonghyeon Ye, Seungone Kim, Shrestha Mohanty, Vipul Gupta, Vivek Sharma, Vu Minh Chien, Xuhui Zhou, Yizhi Li, Caiming Xiong, Luis Villa, Stella Biderman, Hanlin Li, Daphne Ippolito, Sara Hooker, Jad Kabbara, and Sandy Pentland. 2024. Consent in Crisis: The Rapid Decline of the AI Data Commons. In Advances in Neural Information Processing Systems, A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (Eds.), Vol. 37. Curran Associates, Inc., 108042–108087. https://proceedings.neurips. cc/paper%5Ffiles/paper/2024/file/c3738949a80306cc48a8ea8ba0560f9d-PaperDatasets%5Fand%5FBenchmarks%5FTrack.pdf [18] Shayne Longpre, Robert Mahari, Ariel Lee, Campbell Lund, Hamidah Oderinwale,

William Brannon, Nayan Saxena, Naana Obeng-Marnu, Tobin South, Cole Hunter, Kevin Klyman, Christopher Klamm, Hailey Schoelkopf, Nikhil Singh, Manuel Cherep, Ahmad Mustafa Anis, An Dinh, Caroline Chitongo, Da Yin, Damien Sileo, Deividas Mataciunas, Diganta Misra, Emad Alghamdi, Enrico Shippole, Jianguo Zhang, Joanna Materzynska, Kun Qian, Kush Tiwary, Lester Miranda, Manan Dey, Minnie Liang, Mohammed Hamdy, Niklas Muennighoff, Seonghyeon Ye, Seungone Kim, Shrestha Mohanty, Vipul Gupta, Vivek Sharma, Vu Minh Chien, Xuhui Zhou, Yizhi Li, Caiming Xiong, Luis Villa, Stella Biderman, Hanlin Li, Daphne Ippolito, Sara Hooker, Jad Kabbara, Sandy Pentland, and Data Provenance Initiative. 2025. Consent in crisis: the rapid decline of the AI data commons. In Proceedings of the 38th International Conference on Neural Information Processing Systems (Vancouver, BC, Canada) (NIPS ’24). Curran Associates Inc., Red Hook, NY, USA, Article 3431, 46 pages. [19] Nicola Lucchi. 2023. ChatGPT: A Case Study on Copyright Challenges for

Generative Artificial Intelligence Systems. European Journal of Risk Regulation 15, 3 (Aug. 2023), 602–624. doi:10.1017/err.2023.59 [20] Alexandra Luccioni and Joseph Viviano. 2021. What’s in the Box? An Analysis

of Undesirable Content in the Common Crawl Corpus. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (Eds.). Association for Computational Linguistics, Online, 182–189. doi:10.18653/v1/ 2021.acl-short.24 [21] Chandler May, Alex Wang, Shikha Bordia, Samuel R. Bowman, and Rachel

Rudinger. 2019. On Measuring Social Biases in Sentence Encoders. In Proceedings of the 2019 Conference of the North. Association for Computational Linguistics. doi:10.18653/v1/n19-1063 [22] Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient

Estimation of Word Representations in Vector Space. arXiv:1301.3781 [cs.CL] https://arxiv.org/abs/1301.3781 [23] OpenAI. 2023. OpenAI and Axel Springer Partnership to Enhance AI in Journal-

ism. https://openai.com/index/axel-springer-partnership/. Accessed: 2025-09-26. [24] OpenAI. 2023. Overview of OpenAI Crawlers. https://platform.openai.com/docs/

bots. Accessed: 2025-09-27. [25] OpenAI. 2024. A Landmark Multi-Year Global Partnership with News Corp. https://openai.com/index/news-corp-and-openai-sign-landmark-multiyear-global-partnership/. Accessed: 2025-09-26. [26] Guilherme Penedo, Hynek Kydlíček, Loubna Ben allal, Anton Lozhkov, Mar-

garet Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf. 2024. The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale. arXiv:2406.17557 [cs.CL] https://arxiv.org/abs/2406.17557 [27] Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru,

Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023. The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only. arXiv:2306.01116 [cs.CL] https://arxiv.org/abs/2306.01116 [28] Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. GloVe:

Global Vectors for Word Representation. In Empirical Methods in Natural Language Processing (EMNLP). 1532–1543. http://www.aclweb.org/anthology/D141162 [29] Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang,

Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21, 1, Article 140 (Jan. 2020), 67 pages. [30] Ronald E. Robertson, Shan Jiang, Kenneth Joseph, Lisa Friedland, David Lazer,

and Christo Wilson. 2018. Auditing Partisan Audience Bias within Google Search. Proceedings of the ACM on Human-Computer Interaction 2, CSCW (Nov. 2018), 1–22. doi:10.1145/3274417 [31] David Rozado. 2020. Wide range screening of algorithmic bias in word embedding

models using large sentiment lexicons reveals underreported bias types. PLOS ONE 15, 4 (April 2020), e0231189. doi:10.1371/journal.pone.0231189 [32] Kimberly Ruth, Aurore Fass, Jonathan Azose, Mark Pearson, Emma Thomas,

Caitlin Sadowski, and Zakir Durumeric. 2022. A world wide view of browsing the world wide web. In Proceedings of the 22nd ACM Internet Measurement Conference (IMC ’22). ACM, 317–336. doi:10.1145/3517745.3561418 [33] Kimberly Ruth, Deepak Kumar, Brandon Wang, Luke Valenta, and Zakir Du-

rumeric. 2022. Toppling top lists: evaluating the accuracy of popular website lists. In Proceedings of the 22nd ACM Internet Measurement Conference (IMC ’22).

Shopping & Auctions

Business & Economy

Entertainment

Education

News & Politics

Technology

Other

Society & Lifestyle

Health

Government & Politics

Travel

Sports CommonCrawl

Figure 6: Distribution of topics in CommonCrawl, assessed through Cloudflare domain categorization. Categories less prevalent than 2% are aggregated in “other”.

English

Russian

German

Japanese

Chinese

French

Spanish

Portuguese

Italian

Dutch

Polish

Persian

0%

10%

20%

30%

40%

50%

Language Fraction

44.3%

50.5%

6.1% 3.9% 6.0% 3.8% 5.2%6.2% 5.2%

1.4% 4.5%3.5% 4.4%6.0%

2.1%3.4% 2.0%1.9% 1.9%1.3% 1.7%1.8% 0.8%1.8%

CommonCrawl Web - Top 1M CrUX

Figure 7: Distribution of languages of pages included in CommonCrawl and in the top one million in CrUX analytics. Importantly, because of a limited Google Chrome usage in People’s Republic of China, the fraction of Chinese pages is abnormally low in CrUX analytics.

ACM, 374–387. doi:10.1145/3517745.3561444 [34] Alain Strowel. 2023. ChatGPT and Generative AI Tools: Theft of Intellectual

Labor? IIC - International Review of Intellectual Property and Competition Law 54, 4 (April 2023), 491–494. doi:10.1007/s40319-023-01321-y [35] Tony Stubblebine. 2023. Default No to AI Training on Your Stories. https://blog.

medium.com/default-no-to-ai-training-on-your-stories-abb5b4589c8 Medium announced it would block AI crawlers, stating "AI companies have leached value from writers in order to spam internet readers". [36] João Tomé, Jorge Pacheco, and Carlos Azevedo. 2025. From Googlebot to gptbot:

Who’s crawling your site in 2025. https://blog.cloudflare.com/from-googlebotto-gptbot-whos-crawling-your-site-in-2025/ [37] Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne

Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. LLaMA: Open and Efficient Foundation Language Models. arXiv:2302.13971 [cs.CL] https://arxiv.org/abs/2302.13971 [38] Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaud-

hary, Francisco Guzmán, Armand Joulin, and Edouard Grave. 2020. CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data. In Proceedings of the Twelfth Language Resources and Evaluation Conference, Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Hélène Mazo, Asuncion Moreno, Jan Odijk, and Stelios Piperidis (Eds.). European Language Resources Association, Marseille, France, 4003–4012. https://aclanthology.org/2020.lrec-1.494/

A Appendix

Received 7 October 2025

Illustration from source page 7
Illustration from source page 7.
Illustration from source page 7
Illustration from source page 7.

PDF Page 8

Bouchaud and Ramaciotti

Supercategory Granular Categories Top keywords Adult Themes Pornography, Adult Themes, Nudity porn, telegram, sex, free, videos Business & Economy Business, Economy & Finance, Brokerage & Investing, Cryptocurrency, Professional Networking

smm, online, instagram, jobs, followers

Education Educational Institutions, Education, Science, Space & Astronomy, School Cheating

universitas, online, moodle, kampus, universit

News & Politics News & Media, Government & Politics news, online, newspaper, latest, breaking Government & Politics Politics, Advocacy, and Government-Related online, government, development, portal, records Entertainment Audio Streaming, Music, Magazines, Cartoons & Anime, Movies & Home Video, Arts, Entertainment, Gaming, Video Streaming, Television, Comic Books, Paranormal

manga, anime, monster, sasaki, chapter

Gambling Gambling online, slot, result, casino, betting Health Health & Fitness, Sex Education health, medical, cosmetics, healthcare, care Internet Communication Forums, Webmail, Chat & Messaging forum, video, email, online, indian, free Job Search & Careers Job Search & Careers jobs, academy, work, vacancies, career Other Redirect, Unknown Questionable Content Drugs, Questionable Content, Questionable Activities, Alcohol, Hacking, Profanity

movies, free, watch, download, streaming

Real Estate Real Estate estate, casa, property, muebles, venta Religion Religion suresi, prayer, bible, quran, church Shopping & Auctions Ecommerce, Auctions & Marketplaces, Coupons silver, watches, jewellery, gold, necklace Society & Lifestyle Lifestyle, Clothing and Fashion, Food & Drink, Hobbies & Interests, Home & Garden, Pets, Parenting, Photography, Astrology, Dating & Relationships, Arts & Crafts, Sexuality, Tobacco, Body Art, Digital Postcards

dating, food, photo, baby, horoscope

Sports Sports football, live, soccer, sports, bike Technology Technology, File Sharing, Artificial Intelligence free, download, apk, video, spotify Travel Travel flight, travel, bus, booking, hotel Vehicles Vehicles duramax, car, diesel, auto, tuner Violence Weapons, Violence holsters, gun, airsoft, firearms, reloading Weather Weather weather, meteo, forecast, previsioni, italia Table 1: Taxonomy of Cloudflare’s website content categorization. For each "super-category" we report some of the most overrepresented words in the meta tags of webpages categorized as such.

2015 2017 2019 2021 2023 2025 0%

5%

10%

15%

20%

Fraction of FineWeb-Edu

disallowed by robots.txt

A) GoogleBot CCBot

2015 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 0%

20%

40%

60%

80%

100%

Fraction of FineWeb

4% 4% 5% 5% 5% 5% 4% 5% 4% 4% 4%

62% 62% 62% 62% 60% 60% 60% 59% 58% 58% 58%

25% 25% 25% 25% 27% 27% 27% 27% 28% 28% 27%

7% 7% 7% 7% 7% 7% 8% 8% 8% 8% 9%

B)

5 4 3 2 1 0

Figure 8: (A) Fraction (in tokens) of FineWeb-Edu sourced from websites that disallow CCBot as of August/September 2025, by year of collection. (B) Distribution of “educational quality” score in FineWeb over time (text snippet score weighted by token count).

Illustration from source page 8
Illustration from source page 8.
↑ Top