Let's look at some data and understand what is going on with the internet now, as static content is losing its value every day, and how it has affected the technology itself when humans stopped using their creativity in writing. If we look at web pages from Common Crawl, which are cleaned with the FineWeb filters that AI builders usually use, we find that AI-generated text was around 10.1% of the total share of text in June 2024. In June 2025, it became around 16%, and in August 2026 it was found to be 31.1%. This is forecast to be 51% by the end of 2028.
Researchers at Pangram Labs trained around 800 small language models and saw that, for models trained on the usual 20 tokens per parameter, training was fine if they mixed in up to about 8 AI tokens per 100 human tokens. With more AI tokens than that, the training got hurt. One of the most surprising things is that the popular cleaning filters for world wide web data favour AI text. Another angle to look at is that the increase in data influences training costs directly. The more garbage the data is, the more iterations have to be done, so overall a lot of compute is wasted.
Even though this study shows some worry, this research is still an arXiv preprint, but so was the "Attention Is All You Need" paper, which revolutionized the NLP field.
Why it matters to our readers: they must understand that people relied largely on the web for information that was not accessible from other parts of the world. However, with AI-generated content dominating at a rapid pace, the motivation to search for and write original information is going down, and data is being contaminated by AI-generated content, which many techies also call "AI SLOP" in slang.
Moreover, students of AI need to understand that just having more data is not better, which is clearly proven when researchers trained AI models with AI data itself. Finally, it connects to the wider worry about AI learning from its own output, implying it doesn't learn anything new.
Following are some of the concepts that are frequently used when dealing with AI model training, explained here in simple language.
- Token: a word or piece of a word that a model reads and predicts. Lesson: Tokenization
- Loss: a score for how wrong the model's guesses are. Lower is better.
- Parameter: one of the adjustable numbers inside a model, set during training.
- Scaling law (Chinchilla): a formula predicting how much a model improves with more data and computing; the rule of thumb is about 20 tokens of text per parameter. Lesson: What a language model is