Pew Research analysed about 490,000 English-language web pages drawn from the Common Crawl archive, covering January 2021 to July 2026, using the Open Pangram detection model. One page in ten across the whole sample shows significant signs of having been written by a machine. Among pages published after ChatGPT's launch in November 2022, the share is 35%.
The distribution by domain is the more interesting part. Commercial sites are furthest along, at 9.35% of .com pages as of January 2026, up from roughly 1% five years earlier. Non-profits sit at 4.6%. Universities and government at about 1% each.
Read the caveat before the headline
Pew is explicit that a detection signal is not proof: a flagged page may have been AI-assisted rather than machine-written outright.
That distinction covers an enormous amount of ground. A journalist who drafts unaided and runs a grammar tool over the result, a small business owner who asks a model to tidy their About page, and a content farm generating a thousand pages a night all leave traces that overlap in the statistics.
Detection classifiers are also imperfect in a way that is not symmetric. They tend to flag writing that is fluent, regular and structurally conventional, which describes a great deal of competent human prose, particularly by writers working in a second language or to a house style. Any figure of this kind should be read as an upper bound on a fuzzy quantity, not a census.
The tells
The linguistic markers Pew tracked are worth listing because they are becoming folk knowledge: Oxford commas up 63%, em dashes roughly doubled, the words "delve", "interplay" and "testament" more than doubled, and constructions built on negative parallelism, of the form "it is not x, it is y", up nearly threefold since 2023.
None of these is evidence on its own. Plenty of careful writers have always used an Oxford comma. What the aggregate shows is a distribution shifting: the average page on the commercial web now reads more like the average model output than it did four years ago, because a growing share of it is.
Why this matters commercially
Three reasons, in ascending order of importance.
The first is search. Google's ranking systems and its AI summaries are both trained on and served from this corpus, and a web in which a third of new commercial pages are machine-assembled is a harder place to distinguish a useful answer from a plausible one. Boursel reported earlier this week that Google is offering publishers new tools to address AI-driven traffic losses, which is the same problem seen from the publisher's side.
The second is the economics of publishing. If the marginal cost of a page approaches zero, the volume of pages rises until attention, not production, is the constraint. That is already visible in the commercial web's numbers, and it depresses the value of the advertising inventory those pages carry.
The third is model training, and it is the one with the longest tail. Frontier models are trained on scrapes of this same web. As the share of machine-written text in each successive scrape rises, models are increasingly trained on the output of earlier models. Researchers have shown that recursive training on synthetic data degrades a model's grasp of the tails of a distribution, the rare and unusual cases, in a process sometimes called model collapse. Nobody knows where the threshold is, or whether careful data curation avoids it entirely, but the input is measurably changing.
The number to watch
Not the headline 35%, which is a detection estimate over a single archive. The trend in the .com figure, which went from about 1% to 9.35% in five years, and the gap between it and the .edu and .gov figures that have stayed near 1%.
That gap says something useful: the shift is concentrated where publishing is a business rather than a record. Where there is an incentive to produce volume, volume is being produced. Where the incentive is to produce a document that has to be correct, much less has changed.



