News : On August 20, 2026, the Pew Research Center published an analysis of 490,000 English-language webpages drawn from the Common Crawl archive between January 2021 and July 2026: 10% of pages in the July 2026 sample show significant signs of AI authorship or editing, against roughly 1% in January 2021 (Pew Research Center, August 20, 2026).
The share is not what surprises. The speed at which it was reached is. In January 2021, about one page in a hundred carried the statistical markers of machine-assisted writing. By January 2026, it was close to one in ten for .com domains. Pew put a number on something every publisher sensed but could not measure, and the method matters as much as the result.
What Pew measured, and how
The Pew Research Center's Data Labs team collected nearly half a million English-language pages from Common Crawl, the public web archive, across five years. Each text was run through a detector called Open Pangram, a model that picks up regularities in vocabulary and sentence construction that appear more often in language models than in human authors.
The headline figure comes from a random sample of 10,000 pages collected in July 2026: 10% show significant signs of AI authorship. Pew immediately adds the qualification that matters. That basket mixes old and recent pages, many of them predating ChatGPT and therefore impossible to attribute to it. Keeping only pages published after ChatGPT's release in late November 2022, the share rises above one third.
It is the second figure that describes today's editorial output. On the recent English-language web, more than one page in three carries traces of a language model.
The split across domains is wide
The distribution is very uneven, and it tracks the commercial pressure bearing on each type of site fairly closely. Here are the six-month averages Pew published for the January 2026 point.
| Domain | January 2021 | January 2024 | January 2026 |
|---|---|---|---|
| .com | 1.09% | 3.76% | 9.35% |
| .org | 0.83% | 2.00% | 4.59% |
| .edu | 0.57% | 1.70% | 1.03% |
| .gov | 0.41% | 1.42% | 0.76% |
Commercial domains concentrate the phenomenon. A .com site is roughly twice as likely to carry these markers as a non-profit site, and ten times as likely as a university or government one. At the start of the series in 2021, all four families sat within a whisker of each other, around 1%. The gap widened as content production became a cost line to optimise.
The models' writing tics are now measurable
The most useful part of the study for a publisher is not the headline percentage, it is the tally of markers. Pew counted the frequency of several writing traits on pages published after November 2022, in occurrences per 10,000 words.
| Marker (per 10,000 words) | January 2023 | January 2026 | Change |
|---|---|---|---|
| Em dash | 5.79 | 11.19 | about 2x |
| Oxford comma | 34.04 | 55.51 | +63% |
| AI-typical vocabulary | 11.94 | 26.02 | more than double |
| Negative parallelism | 0.87 | 2.36 | nearly triple |
"Negative parallelism" is the "it's not just X, it's Y" construction, now a recognisable signature. The vocabulary Pew tracked includes an explicit list of over-represented terms: delve, tapestry, testament, underscore, pivotal, showcase, meticulous, intricate, landscape.
Pew sets out the limit itself, and it is an important one: taken alone, none of these signs proves anything. Human authors use em dashes and Oxford commas. The markers only become readable across very large corpora, never on a single document.
What this changes for your visibility
The production method is not the problem. This bears repeating, because the study will be misread. Google does not sanction the use of AI. Its spam policies target scaled content abuse, defined as generating many pages "for the primary purpose of manipulating search rankings and not helping users". The policy does mention generative AI tools, but always with the same condition attached: without adding value for users. We covered that distinction in Google doesn't penalise AI content, it penalises bad content.
The real risk is sameness. If more than a third of recent pages carry the same stylistic signature, then reading like the average becomes a competitive handicap. An article that is correct, clean and perfectly interchangeable gives nobody a reason to rank it above another. That is the mechanism we observed across sites publishing at volume, described in our analysis of content production at scale, and in the deindexing waves that followed.
For AI citation, the stakes are even more direct. An answer engine picks a handful of sources from thousands of pages saying the same thing. What separates them is what your page holds that the others do not: a number you measured, a client case, a method, a reasoned disagreement. A page that only restates the consensus offers no handhold. We showed, using another Pew survey on AI-generated summaries, how tight the citation space has become.
What to do now
1. Revisit your pages published since 2023. Not to hunt for AI, which would be pointless. To check, page by page, that each one asserts at least one thing only you could write. If the answer is no, the page is another duplicate.
2. Have a practitioner read it. It is the only filter that catches what a detector cannot see: the example that rings false, the missing nuance, the objection any practitioner would have raised.
3. Run your drafts against the markers. We run an automated check on every article before publication, including this one, and it blocks the release if the density of typical constructions crosses a threshold. The markers Pew reports overlap almost exactly with our own list, which is reassuring about the method and fairly uncomfortable about how ordinary the pattern is.
4. Name your sources. "Studies show that" is the phrase models produce most and the one readers can verify least. Citing Pew, with its date, its sample and its limits, costs three minutes and changes the nature of the text.
5. Do not confuse detection with penalty. No detector is wired into Google's ranking. What degrades with volume production is perceived usefulness, not a machine score.
The limits of this study
It should be read for what it is. Pew acknowledges that "AI detection models aren't perfect" and that they sometimes misclassify human-written text as showing signs of AI authorship, and the reverse. The findings hold in aggregate, not at the level of a single document.
The sample is English-language only. No data is published here for other language webs, and nothing guarantees the magnitudes transfer. Common Crawl, for its part, samples the web without covering it exhaustively, with its own collection biases. Finally, the study measures the presence of stylistic markers, not content quality, and still less search performance. It says nothing about ranking.
Frequently asked questions
Does Google penalise AI-written content?
Not on the basis of how it was produced. Google's spam policies target scaled content abuse, defined as generating many pages primarily to manipulate rankings rather than to help users. The policy explicitly mentions using generative AI tools to produce many pages without adding value. The test is intent and value, not the tool.
10% or a third: which figure should you use?
They measure different things. 10% is the share of the whole July 2026 sample, which mixes old and recent pages. More than a third is the share once you keep only pages published after ChatGPT's release in late November 2022. To judge current editorial output, the second figure is the relevant one.
Does the study cover non-English pages?
No. Pew states the sample covers 490,000 English-language pages from Common Crawl. No equivalent data is published for other languages in this study, so the magnitudes cannot be transferred directly to non-English webs.
Does an em dash mean my text will be flagged as AI?
No. Pew says so explicitly: on their own, an em dash or an Oxford comma prove nothing, since human authors use them too. These markers only become meaningful statistically, across very large corpora. A detector applied to a single document errs in both directions, which Pew acknowledges for its own tool.
The Cicero take
The news is not that the web is filling up with machine-produced text, which everyone suspected. It is that there is now a public, dated and methodologically legible measure of it, and that this measure describes a shared stylistic signature. When a third of recent output uses the same constructions, the resemblance is the problem.
Our reading is straightforward. The question to ask before publishing is no longer "was this written by an AI", which interests neither Google nor your readers. It is "does this page contain something nobody else could have written". On that test, most editorial calendars do not survive an honest audit.
Sources
- → Pew Research Center, "How much of the internet is written with AI?" (primary source), August 20, 2026: analysis of 490,000 English-language pages from Common Crawl, Open Pangram detector, 10,000-page sample for July 2026, series by domain and by stylistic marker.
- → Google Search Central, spam policies (primary source): definition of scaled content abuse and explicit mention of generative AI tools used without adding value.
- → Common Crawl Foundation (primary source): the public web archive Pew used as its sampling base.
- → Search Engine Journal, August 24, 2026: coverage of the study and its implications for the search industry.
Growth and SEO & GEO content strategist, I founded Cicéro to help businesses build lasting organic visibility : on Google and in AI-generated answers alike. Every piece of content we produce is designed to convert, not just to exist.
LinkedIn