• 4 mins read
  • Published

AI Bots Ramp Up Content Scraping, Publishers See Less Traffic

Ken Doctor media analyst FAYFO Media

by Ken Doctor

AI Bots Ramp Up Content Scraping, Publishers See Less Traffic FAYFO Media © fayfo.com
AI Bots Ramp Up Content Scraping, Publishers See Less Traffic © fayfo.com

Automated AI bots are extracting more content from publisher sites. New data shows European publishers face higher scraping rates. Referral traffic from AI systems remains minimal.

Media companies are facing a surge in automated content scraping by AI bots, with significant implications for traffic, monetization, and editorial strategy. According to TollBit’s latest “State of the Bots” report, the first half of 2026 saw over 22 billion visits to websites by AI bots, out of a total dataset of more than 987 billion site visits. Of particular concern, around 1.9 billion of these scraping actions reportedly bypassed or ignored robots.txt directives set by publishers.

European publishers are being hit especially hard. TollBit’s analysis found that, on average, sites in Europe are scraped by AI bots about four times more often than those in North America. Between January and June, scraping activity on European sites rose by nearly 20 percent. Robots.txt instructions were disregarded almost three times as frequently in Europe as in North America, though the reasons for this disparity remain unclear. TollBit suggested that multilingual content may be especially attractive to AI companies, but also noted its own customer base is not fully representative of the entire web.

The economic challenge for publishers goes beyond increased bot traffic. Traditional search engines have historically sent visitors back to publisher sites in exchange for indexing their content. With AI systems, this exchange is breaking down. TollBit reported that referrals from AI chatbots generate 96 percent fewer click-throughs than classic Google searches. While Google claims that AI-driven traffic converts better, the Reuters Institute for the Study of Journalism found that organic Google traffic to over 2,500 news sites worldwide dropped by 33 percent between November 2024 and November 2025, based on Chartbeat data. Media executives surveyed expect search traffic to decline by another 43 percent over the next three years. Although platforms like ChatGPT and Perplexity are sending more visitors, their share of total referral traffic remains small-Google Search still delivers about 500 times more referrals than ChatGPT, according to Reuters Institute.

Publishers are responding with new technical barriers. Cloudflare now blocks AI training crawlers by default on new domains unless site owners opt in. Starting September 2026, Cloudflare plans to further differentiate between bots used for search, training, and AI agents. The company is also developing models to allow paid access to content for AI systems. On the user side, the Reuters Institute’s Digital News Report 2026 shows that while AI chatbots are not yet a dominant news source, their use is growing. Users who access news via chatbots are less likely to click through to original sources compared to those using traditional search engines, but when they do, verification and source credibility play a larger role.

German publishers are taking varied approaches. Outlets like the “Süddeutsche” allow bots to access their content, hoping for a competitive edge. Others, such as the “Spiegel,” block bots until fair compensation models are established. Smaller publishers and niche media may have little choice; if they block bots, it is unlikely that Google, OpenAI, or Perplexity will classify their content as essential.

TollBit’s business model is closely tied to this issue, as the company helps publishers detect, block, and monetize AI-driven access. TollBit offers a bot paywall, aiming to convert unauthorized scraping into paid access. While its report highlights a real and growing problem, the company’s commercial interest means its data should not be seen as a neutral snapshot of the entire internet. However, the scale of the challenge is supported by independent data from the Reuters Institute and Cloudflare. The longstanding deal of the open web-content in exchange for reach-is under pressure from AI search and autonomous agents. For publishers, the key question is shifting from whether AI systems can use their content to under what terms and at what price.

Recent moves by major tech companies reflect the growing tension between AI development and publisher interests. For example, Microsoft’s introduction of AI token budgets for engineers signals a broader industry focus on balancing AI usage with business impact and resource allocation.

Related articles