All news

AI is writing the web, and it can't train anymore!

The whole internet is now filled with AI-written text. Humans are not putting any effort into writing text with AI, and almost a third of text is now written by AI. This has led to worthless training of AI, and it is not gaining any more knowledge now.

By Pravesh Srivastava2 min read
Bar chart: the share of web text labelled AI-generated rose from 10.1% in June 2024 to about 16% in June 2025, 27.5% in June 2026 and 31.1% in August 2026.
Image: Sahi Padhai, using data from Russell et al. (2026)

Let's look at some data and understand what is going on with the internet now, as static content is losing its value every day, and how it has affected the technology itself when humans stopped using their creativity in writing. If we look at web pages from Common Crawl, which are cleaned with the FineWeb filters that AI builders usually use, we find that AI-generated text was around 10.1% of the total share of text in June 2024. In June 2025, it became around 16%, and in August 2026 it was found to be 31.1%. This is forecast to be 51% by the end of 2028.

Researchers at Pangram Labs trained around 800 small language models and saw that, for models trained on the usual 20 tokens per parameter, training was fine if they mixed in up to about 8 AI tokens per 100 human tokens. With more AI tokens than that, the training got hurt. One of the most surprising things is that the popular cleaning filters for world wide web data favour AI text. Another angle to look at is that the increase in data influences training costs directly. The more garbage the data is, the more iterations have to be done, so overall a lot of compute is wasted.

Even though this study shows some worry, this research is still an arXiv preprint, but so was the "Attention Is All You Need" paper, which revolutionized the NLP field.

Why it matters to our readers: they must understand that people relied largely on the web for information that was not accessible from other parts of the world. However, with AI-generated content dominating at a rapid pace, the motivation to search for and write original information is going down, and data is being contaminated by AI-generated content, which many techies also call "AI SLOP" in slang.

Moreover, students of AI need to understand that just having more data is not better, which is clearly proven when researchers trained AI models with AI data itself. Finally, it connects to the wider worry about AI learning from its own output, implying it doesn't learn anything new.

Following are some of the concepts that are frequently used when dealing with AI model training, explained here in simple language.

  • Token: a word or piece of a word that a model reads and predicts. Lesson: Tokenization
  • Loss: a score for how wrong the model's guesses are. Lower is better.
  • Parameter: one of the adjustable numbers inside a model, set during training.
  • Scaling law (Chinchilla): a formula predicting how much a model improves with more data and computing; the rule of thumb is about 20 tokens of text per parameter. Lesson: What a language model is

Learn the idea

The paper

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell et al.

arXiv papers are preprints: other scientists have not yet checked them (peer review).

Sources

  1. arXiv: the paper
  2. WildAI: data, models and code
  3. Common Crawl