Ground Truth.
AI, checked against the source.

News · 2026-08-21

Pew finds one in ten English web pages shows signs of AI authorship

Pew Research Center published a measurement of how much of the English-language web now shows signs of machine authorship, and the headline number is 10%. Pew sampled 490,000 pages from Common Crawl across 49 crawls spanning January 2021 to July 2026, 10,000 pages per crawl, and found that 10% of the July 2026 sample showed significant signs of AI authorship. Among pages in that crawl carrying a publication date after ChatGPT's release, the figure was 35%.

Key facts

The number everyone will repeat is 35%. The number that is actually defensible is 10%, and the difference between them is the whole methodological story.

Start with what Common Crawl is, since the study rests on it. Common Crawl is a nonprofit that has been scraping and archiving the public web at regular intervals for over a decade, releasing the snapshots freely. It is the raw material behind a great deal of language-model training data, which makes it an unusually apt place to ask how much of the web is now machine-written: the corpus that trained the models is measurably filling up with the models' own output.

Pew's detector choice is the load-bearing methodological decision, and the organization documented it carefully in a separate methodology page. It used Pangram's open-weight editlens_Llama-3.2-3B model, which scores each page between 0 and 1, and treated a score of 0.2 or higher as showing meaningful signs of AI authorship or editing. That threshold is deliberately inclusive, it catches pages a person wrote and a model polished, not only pages a model produced whole.

Detection is probabilistic, and Pew is admirably direct about this rather than reporting a single accuracy figure. It compared the open detector against a commercial one across 62,370 pages, reporting 96% raw agreement and a Cohen's kappa of 0.61, a statistic that adjusts for how often two raters would agree by chance and which lands in the moderate range. Pew also flags that the open model appears more prone to false positives on pages written before ChatGPT existed, which is exactly the kind of caveat that usually gets sanded off in secondhand coverage.

An analogy for the 10-versus-35 gap: imagine measuring what fraction of cars on the road are electric. Count every car and you get a low number, because the road is full of vehicles built decades ago. Count only cars registered in the last three years and the number jumps. Both are true. Only one answers the question "what is being built now," and only the other answers "what is out there." The 35% figure is the second kind, restricted further to pages that happen to publish a machine-readable date, which is itself a non-random slice of the web.

Why this matters is not mainly about slop. It is about the feedback loop. Models trained on the web are now training on text other models wrote, and a companion paper this week gives that loop a measurable shape. arXiv 2608.19437, published August 19, studied 68 models from 12 providers between March 2023 and July 2026 using open-ended creativity tasks, embedding responses and tracking semantic distance over release date. Its conclusion is not that models became more creative. It is that model outputs have become more similar to each other over time. Less spread, not more.

Put those two findings side by side and you get the interesting claim: the pool of text is increasingly machine-written, and the machines writing it are increasingly writing the same way. That is a convergence pressure on the whole written commons, and it is the strongest reason to care about provenance and watermarking, which is under pressure from a separate direction, as a tool that strips SynthID and C2PA marks passing 4,900 stars illustrates.

The honest caveat, and it is a big one: these are classifier outputs, not ground truth. Nobody asked the authors. A detector with a moderate agreement statistic and a known bias toward false positives on older text is a measuring instrument with a documented error bar, and the right way to read 10% is "roughly a tenth, by this instrument, at this threshold." Pew reported it that way. Most coverage will not.


Primary source, verified: read the paper →

Key questions

How did Pew decide a page was AI-written?

It ran each page through an open-weight detector, Pangram's editlens model built on Llama 3.2 3B, which scores a page from 0 to 1, and counted anything scoring 0.2 or above as showing meaningful signs of AI authorship or editing. That is a probabilistic judgment, not a determination.

Why is the 35% figure higher than the 10% figure?

The 35% applies only to the subset of pages in the July 2026 crawl that carried a detectable publication date after ChatGPT's release, which is a much newer and smaller slice of the web. The 10% figure covers the whole July 2026 sample, including old pages that predate generative AI entirely.

How reliable is the detector Pew used?

Pew compared its open detector against a commercial one across 62,370 pages and reported 96% agreement with a Cohen's kappa of 0.61, which indicates moderate agreement beyond chance. Pew also notes the open model appears to have a higher false-positive rate on pages written before ChatGPT existed.
Cite this

APA

Ground Truth. (2026, August 21). Pew finds one in ten English web pages shows signs of AI authorship. Ground Truth. https://groundtruth.day/news/pew-finds-one-in-ten-english-web-pages-shows-signs-of-ai-authorship.html

BibTeX

@misc{groundtruth:pew-finds-one-in-ten-english-web-pages-shows-signs-of-ai-authorship,
  title  = {Pew finds one in ten English web pages shows signs of AI authorship},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {aug},
  url    = {https://groundtruth.day/news/pew-finds-one-in-ten-english-web-pages-shows-signs-of-ai-authorship.html}
}

Topics: research · society · ai-detection · web · content-provenance