The Dark Side of AI: How Large Language Models are Overwhelming the Internet

As the use of large language models (LLMs) has increased, so has their hunger for data. These models are trained on vast amounts of text from the internet, but their data-gathering practices are transparent and respectful of website owners' wishes. Many AI companies claim to respect the robots.txt file, which allows website owners to specify which bots can crawl their sites. However, even when their intentions are genuine, AI companies benefit from loopholes in robots.txt, making it difficult for website owners to block their bots. The impact of these bots is significant, with websites like Wikipedia seeing a 50% increase in multimedia downloads due to AI company bots, but also facing increased costs and a 22.7% drop in overall traffic.

Key Takeaways:

  • Large language models are trained on vast amounts of text from the internet, but their data-gathering practices are often opaque and disrespectful of website owners' wishes.
  • AI companies claim to respect robots.txt files, but often exploit loopholes to continue crawling websites despite owners' requests to block them.
  • Wikipedia saw a 50% increase in multimedia downloads due to AI company bots, but also faced a 22.7% drop in overall traffic, highlighting the conflicting interests of AI companies and website owners.
  • Companies like Cloudflare are taking steps to treat AI bots as hackers, emphasizing the need for a more transparent and respectful approach to data collection.
  • The competitive pressures faced by AI companies make them less likely to honor the trust-based mechanism of robots.txt, leading to a vicious cycle that threatens the very data AI needs to improve further.
  • An international agreement, similar to the Montreal Protocol, could encourage countries to co-ordinate legislative efforts to compel companies to honor robots.txt instructions and level the playing field.
  • Without intervention, this could lead to a devastating outcome for the internet, as AI may drive websites to shut down, destroying the data necessary for AI to continue improving.

Statistics:

  • 10 trillion words: the amount of data equivalent to that used to train a previous version of ChatGPT.
  • 13 billion: the number of Globe and Mail op-eds equivalent to the amount of data used to train ChatGPT.
  • 36.5 million years: the time it would take a columnist to generate sufficient data to train a similar model.
  • 50%: the increase in multimedia downloads on Wikipedia between January 2024 and April 2025, due to AI company bots.
  • 22.7%: the drop in Wikipedia's overall traffic between 2022 and 2025.
  • 1994: the year the robots.txt file standard was introduced.

Sources:

  • Wikipedia's multimedia downloads increased by 50% between January 2024 and April 2025 due to AI company bots.
  • Wikipedia saw a 22.7% drop in overall traffic between 2022 and 2025.
  • Cloudflare alleged that Perplexity is actively developing new ways to hide its crawling activities to circumvent existing cybersecurity barriers.
  • The Montreal Protocol: a treaty that bound countries to co-ordinate laws phasing out substances that eroded the Earth's ozone layer.
  • Canada's Online News Act: an attempt to compel social-media companies to compensate news organizations for lost revenue.