Roughly one in ten web pages now shows signs of being AI-written
Pew ran 490,000 web pages through an AI detector. Every figure in its 20 August 2026 report, one entry each, and what each number cannot tell you.
Pew Research Center published an essay on 20 August 2026 under a blunt title: "How Much of the Internet Is Written With AI?" It ran 490,000 pages from 49 Common Crawl snapshots through a detection model, and qualified its estimates carefully.
So here is every figure Pew published, one entry each: what Pew measured, the number, what it does and does not support, then a verdict. All of it comes from the essay and its methodology page, both fetched on 12 September 2026, and where Pew's wording carries the caveat we quote it.
One page in ten, in a single July 2026 snapshot
What Pew measured. A random sample of 10,000 English-language pages collected in July 2026 from Common Crawl, a nonprofit web archive that snapshots the publicly reachable web roughly once a month. Pew scored each page's body text, stripped of HTML and media, with a detection model.
The number. In Pew's words: "Of all the pages in this sample, 10% show significant signs of AI authorship."
What it does and does not mean. One month of one archive, in one language. It counts pages rather than words, so a page scores the same whether a model wrote every line or tidied a few. The sample also holds the web's back catalogue: Pew notes that many pages in it could not have been written by AI, because they predate the tools. And a page must be publicly reachable to enter a Common Crawl snapshot, so sites "that have paywalls or require a login to access content are likely underrepresented".
Verdict: a sound headline figure, provided July 2026, English, and pages rather than words travel with it.
Filter to pages published since ChatGPT and the share hits 35%
What Pew measured. The same July 2026 crawl, narrowed to pages whose HTML carries a publication date, then to dates after ChatGPT's public release on 30 November 2022.
The number. Over one-third. The methodology page is exact: "35% of pages in the July 2026 crawl with post-ChatGPT publication dates show signs of AI authorship."
What it does and does not mean. Pew flags the weakness before anyone else can. Only about 10% to 15% of pages in a crawl sample carry a publication date field at all, and that subset "is not a random subset of the web, so our post-ChatGPT estimate reflects the prevalence of AI authorship among dated content rather than the web as a whole".
Verdict: the louder of the two numbers, and the one with the larger asterisk. It describes dated pages, so never shorten it to 35% of the web.
Where is AI-authored text concentrated?
What Pew measured. The same detection run across all 49 crawls, split by top-level domain: .com, .org, .edu and .gov.
The number. In samples from 2026, around one in ten .com pages show signs of AI authorship, "about double the share on .org domains (4.6%), and 10 times the rate on .edu or .gov domains (both around 1%)". The data behind Pew's chart puts the final point at .com 9.35%, .org 4.59%, .edu 1.03% and .gov 0.76%, each point a six-month average. When ChatGPT was released, Pew says, the patterns appeared at similar rates across all four.
What it does and does not mean. The gap tracks who publishes and why. It says nothing about which suffix deserves your trust: the .gov rate of 0.76% tells you nothing about whether a given .gov page is accurate, and most .com pages in 2026 were written by people. The .edu and .gov figures also sit near the false-positive floor described below.
Verdict: the most useful finding here for a reader deciding where to slow down and check.
Four tells got measurably more common
What Pew measured. How often four features of language appear per 10,000 words, across pages published after 30 November 2022, comparing the 2026 web against a 2023 snapshot. This part is plain text counting, not detector output.
The number. Em dashes appear about twice as often. Oxford commas are up 63%. Words AI models reach for have more than doubled; Pew's examples are "delve", "interplay" and "testament", from a 27-item list that also carries "pivotal" and "tapestry". The fourth is "negative parallelism", Pew's name for structuring a comparison as "it's not just X, it's Y", which has nearly tripled while staying "still fairly rare overall". The series, in uses per 10,000 words from January 2023 to the 2026 point: em dash 5.79 to 11.19, Oxford comma 34.04 to 55.51, AI vocabulary 11.94 to 26.02, negative parallelism 0.87 to 2.36.
What it does and does not mean. These are averages over a very large body of text. Pew is explicit that such marks alone "don't necessarily mean a particular piece of writing was produced using AI", and that "humans use these in their writing too". Nor does the detector work by counting them: Pew describes it as learning "more subtle statistical patterns in word choice and sentence structure".
Verdict: the most quotable part of the study and the least applicable to any single document.
About the tool: Open Pangram, and a threshold of 0.2
What Pew measured. Here the instrument is the finding. Pew fed each page's body text to editlens_Llama-3.2-3B, an open-weight model published by Pangram and known as Open Pangram. It returns a score from 0 to 1, where 0 is fully human-written and 1 fully AI-generated, letting it describe pages that mix the two. Pew counted any page at 0.2 or above as carrying "meaningful" signs of AI authorship or editing.
The number. 0.2, and Pew says where it came from: the threshold "was chosen to closely calibrate the results from Open Pangram to those from Pangram 3.3", Pangram's commercial model. Pew then put 62,370 pages from seven crawls through Pangram 3.3 as a check. The two agreed in 96% of cases, with a Cohen's kappa of 0.61.
What it does and does not mean. Two detectors agreeing tells you they behave alike. Neither is ground truth for who wrote a page, so that cannot establish accuracy. A kappa of 0.61 is moderate once you discount what chance produces, and Pew draws the honest conclusion: "AI detection models are probabilistic tools, and their classifications of an individual page should not necessarily be taken as definitive verdicts about that page's authorship."
Verdict: the threshold defines the headline. At 0.2 the bucket includes pages a person wrote and a model edited, which is why Pew keeps saying signs of.
Treat the 2021 and 2022 figures as approximate
What Pew measured. The same pipeline on crawls from before ChatGPT existed, a de facto control.
The number. Around 1%. Open Pangram put roughly that share of pre-ChatGPT pages in the AI-signs bucket, while Pangram 3.3 returned estimates Pew calls "very low" on the same material.
What it does and does not mean. Pew reads that as the open model having "a higher false-positive rate on pages created before AI use was common", so those shares "should be treated as approximate and likely include some human-written content that was misclassified by Open Pangram". A 2021 page flagged as AI-written was almost certainly nothing of the kind.
Verdict: a study that publishes its own error floor has earned trust on the rest. It also means a slice of every later figure is the tool being wrong.
If you already wrote like that
What Pew measured. Frequencies across a population, and nothing about the authorship of your paragraph.
The number. The one worth noticing is the figure the study does not print: no published rate for how often an individual human-written page gets flagged in 2026. The pre-ChatGPT control hints at the scale and goes no further.
What it does and does not mean. If you have used em dashes and Oxford commas since long before any of this, the web drifting towards your habits does not make your writing machine-made, and nothing in Pew's essay says so. What changed is how much those features tell anyone: a mark appearing twice as often as in 2023 carries half the signal, in either direction. Editing your punctuation out to look human optimises for a classifier at a reader's expense, and Pew's plainest line, that humans use these in their writing too, is the better guide.
Verdict: keep your habits. The shift Pew measured happened in the corpus. It did not happen in your sentence.
What can you conclude from one detector score?
What Pew measured. A number between 0 and 1 for each page, then a cut at 0.2. That is the entire apparatus.
The number. 0.2 again, set to line up with another model rather than to mark where authorship changes hands.
What it does and does not mean. A score above the line says a page's text resembles, statistically, writing produced or substantially edited with AI. It cannot tell you how much AI was involved, or at which stage: a human draft polished by a model and a machine draft rewritten by a person land in the same place. Pew's line on individual pages is the one to carry away, that a classification of one page should not be taken as a definitive verdict about its authorship. How these tools hold up when a school or an employer points one at a named person's work is a separate question with its own evidence, and none of the figures above settle it.
Verdict: detector scores are built for populations. For one page, read the page.
One boundary is worth drawing. A detector asks who or what wrote a page, a question about style. Whether the page is correct runs on a different mechanism, which we wrote up in AI hallucinations explained.
Where we sit in this
We have an interest to state: MultiChats sells access to 60 models from 18 providers, which makes us one of the places pages like these come from. Whether a given model suits your writing is something you settle on your own material, and no study of 490,000 pages will answer it for you.
The version worth repeating
10% of a random sample of 10,000 English-language pages from July 2026. 35% of the pages in that crawl carrying a post-ChatGPT publication date. About 9.35% on .com against roughly 1% on .edu and .gov. Measured with an open-weight detector at a 0.2 threshold, on the part of the web that does not ask you to log in. With those conditions attached the finding survives a sceptic. Without them, somebody turns it into "a third of the internet is fake" by the weekend.