On August 20, 2026, the Pew Research Center published a data essay titled "How Much of the Internet Is Written With AI?" — one of the most ambitious quantitative attempts to answer a question that haunts publishers, writers and everyone who makes a living from content: what share of the web actually bears the fingerprints of machine writing? The team analyzed roughly 490,000 English-language webpages from the Common Crawl archive spanning five full years — January 2021 through July 2026, covering two years before ChatGPT's launch and nearly four years after it — using an open-weight AI detection model called Open Pangram. The numbers are precise, revealing and unevenly distributed. This in-depth analysis walks through them figure by figure, decodes the linguistic "tells" that AI models leave in generated text, and spells out what the findings mean in practice for professional writers, publishers and the health of the web itself.

Source: Pew Research Center (pewresearch.org)
The headline number: from 1% to roughly 10% of .com pages
Before ChatGPT, signs of AI authorship were rare and strikingly evenly spread across the web's main top-level domains: around 1% or less across .com, .org, .edu and .gov throughout 2021 and 2022. Then, after November 2022, the curve bent and kept climbing. In samples from early 2026, around 9.35% of .com pages showed signs of substantial AI writing or editing — roughly one in every ten commercial pages, an increase of nearly nine-fold over the pre-ChatGPT baseline.
The other domains tell a radically different story: about 4.6% of .org pages (largely organizations) and only around 1% of .edu educational and .gov government pages. Commercial content is inflating with machine generation at roughly ten times the rate of educational and governmental institutions. That gap is the whole story in miniature: AI generation has flowed to wherever speed, volume and profit incentives live, and stayed nearly absent where institutional review is slower and stricter.

Source: Pew Research Center — analysis of 490,000 webpages from Common Crawl using Open Pangram
The full timeline: read the curve exactly as Pew published it
The table below is the raw data as it appears in the study itself (shares of pages with significant signs of AI authorship or editing, six-month averages), and it rewards a line-by-line read:
| Crawl date | .com | .org | .edu | .gov |
|---|---|---|---|---|
| Jan 2021 | 1.09% | 0.83% | 0.57% | 0.41% |
| Jul 2021 | 1.04% | 1.02% | 0.34% | 0.23% |
| Jan 2022 | 1.08% | 0.68% | 0.51% | 0.39% |
| Jul 2022 | 1.07% | 0.63% | 0.64% | 0.25% |
| Jan 2023 | 1.63% | 0.90% | 0.56% | 0.34% |
| Jul 2023 | 2.76% | 1.89% | 0.28% | 1.00% |
| Jan 2024 | 3.76% | 2.00% | 1.70% | 1.42% |
| Jul 2024 | 4.70% | 2.10% | 0.57% | 1.72% |
| Jan 2025 | 5.48% | 2.89% | 0.95% | 1.62% |
| Jul 2025 | 6.64% | 2.85% | 1.02% | 1.40% |
| Jan 2026 | 9.35% | 4.59% | 1.03% | 0.76% |
Three observations a careful reader extracts from this table. First, the real .com explosion happened between July 2025 and January 2026 (from 6.64% to 9.35% in a single half-year) — the acceleration is ongoing, and nothing in the data suggests saturation. Second, .org doubled too, but later and more slowly (from 2.85% to 4.59% over the same window), as if the nonprofit sector is following at a lag. Third, .edu and .gov have stayed roughly flat around 1% for five full years despite small transient bumps — meaning the barrier is not technical, since models can write anything, but organizational: strict institutional review blocks pure machine publication.
For anyone building content businesses, the distribution deserves a closer look than the topline. The commercial domain is the race track, and it is also where the risk compounds: as generated content thickens in a category, its marginal return in search drops, and platforms push harder toward demonstrated human expertise — a shift we track directly in how pages rank and what actually reaches readers.
The linguistic tells: how AI-generated text gives itself away
The most instructive part of the Pew essay — and the most practically useful for writers — is its dissection of the stylistic features models reach for more often than humans do. Comparing today's internet to a 2023 snapshot, the researchers measured:
- Em dashes (—): now appear about twice as frequently as before, because models are trained on journalistic and academic writing that leans on them heavily.
- Oxford commas: a 63% increase in usage.
- The AI vocabulary: words like delve, testament, tapestry, landscape, pivotal and underscore have more than doubled in frequency; Pew published a full list of more than 25 model-favorite words.
- Negative parallelism: the "it's not just X, it's Y" comparison structure has nearly tripled, though it remains fairly rare overall.
In absolute per-10,000-word figures the study published: em dashes rose from 5.79 in early 2023 to 11.19 in early 2026, Oxford commas from 34.04 to 55.51, AI-typical vocabulary from 11.94 to 26.02, and negative parallelism from 0.87 to 2.36. The full list Pew classified as "AI-typical vocabulary" also includes: additionally, align with, boasts, bolstered, crucial, emphasizing, enduring, enhance, essential, fostering, garner, highlight, interplay, intricate, key, landscape, meticulous, perfectly, pivotal, showcase, significant, tapestry, testament, underscore, valuable and vibrant.
Better still, the study itself quoted a sample "written by AI" paragraph that stacks every tell into two sentences — reproduced here exactly so you can see it with your own eyes: "The internet has evolved into something far beyond a simple network of connected pages — it's a living, shifting ecosystem shaped by algorithms, creators, and audiences alike, bolstered by waves of AI-generated content that blur the lines between human and machine expression. As users keep delving in to endless streams of posts, videos, and synthetic voices, a pivotal question emerges about authenticity, ownership, and trust—because it's not just information, it's influence." Count the fingerprints: the em dash, the closing negative parallelism, the celebratory vocabulary — a complete signature in a single paragraph.

Source: Pew Research Center — uses per 10,000 words
The practical lesson is not simply "avoid these words." It is that your audience has started reading these patterns as a signal. Text that opens every paragraph with a polished parallel structure and piles on invigorating adjectives (pivotal, crucial, essential) loses reader trust before it loses rankings. Good human writing — with its uneven sentences, specific examples and lived expertise — has become an explicit competitive advantage, not just a stylistic preference.
An honesty check: Pew's number is not "one third"
Some fast-moving press coverage seized the study with headlines claiming "a third of the web" is now AI-written. That is not Pew's number. The figures published in the study itself say roughly 10% of .com pages in 2026 samples show significant signs of AI authorship or editing, with much lower rates elsewhere. Even that figure is an estimate: detection tools err in both directions — classifying human text as generated and vice versa — and their value emerges when reading millions of pages together, not when judging a single page. We flag this because reading numbers exactly as published, without inflation, is the only way to make good decisions from them.
Why this matters if you write in Arabic (or any non-English language)
It might seem that an English-language study has little to say to Arabic-content creators. The opposite is true, for three reasons:
- The same wave is coming. What happened to the English web between 2023 and 2026 is happening now in Arabic as writing tools spread. Being later to the curve is the one advantage: we already know its shape. It starts with merchants filling stores with generated product descriptions and ends with collapsing reader trust and search visibility for anyone who failed to differentiate.
- The stylistic patterns translate. The literal word list is English, but its spirit carries over: the polished balanced sentences, the ceremonial openings, the worn metaphors — all of it reads as "generated tone" to an Arabic reader's ear even without a name for it.
- The opportunity runs the other way. The thicker the cheap generated content gets, the higher the value of demonstrated expertise: real hands-on tool testing, verified primary-source numbers, honest opinions about a product's flaws. That is precisely what we commit to in our AI tools reviews: actual use and official sources before publication.
What this means for you (publishers, writers, working students)
- If you run a website: do not bet on pure generation to fill pages. The data shows the English market has started punishing it — through trust erosion and search demotion — and Arabic will follow. Let AI assist with drafting and editing, but keep the expertise, examples and judgment human.
- If you are a professional writer: you now compete with a machine that is excellent at polish. Your edge is not polish; it is what the machine lacks — you used the thing yourself, made mistakes, and can describe them. Audit your drafts and delete every sentence any model could write about any topic.
- If you are a student or researcher: the Pew essay is a model of method: a declared sample (490,000 pages), a named tool (Open Pangram), a declared data source (Common Crawl), and clear time windows. When the next AI study crosses your feed, ask: what sample? what instrument? who funded it?
- If you worry about detectors flagging your writing: remember that detection is unreliable at the individual-document level, and the fix is not "trick the detector" but adding genuine human value. Deep manual editing and personal examples change a text substantively, not cosmetically.
Quick comparison: Pew's figures versus the popular narrative
| Indicator | What Pew's 2026 study says | What the popular narrative says |
|---|---|---|
| Share of .com pages with AI signs | ~9–10% in 2026 samples | "A third" or "most content" in rushed headlines |
| .edu and .gov domains | ~1% only | Rarely mentioned |
| Accuracy of individual detection | Errs both ways; value is statistical | Sometimes presented as final verdicts on a text |
| Language studied | English only | Results generalized to all languages |
The study's limitations, stated plainly
The research is methodologically solid but explicitly bounded. The sample is English-only, so we do not yet know AI-content rates in Arabic or whether the curve will look the same. Pangram, like any detector, is prone to misclassification even if it performs better in aggregate. Common Crawl is a sample of the web, not the web, and it skews toward more-crawled pages. And the study measures "significant signs" of AI writing or editing — it does not separate fully generated text from human text re-phrased by a model, a distinction that has become nearly impossible technically and arguably irrelevant practically.
Frequently asked questions
What percentage of internet content is written by AI according to Pew?
In early-2026 samples, about 9.35% of .com pages showed signs of substantial AI writing or editing, versus about 4.6% of .org pages and around 1% of .edu and .gov pages. The study covered 490,000 English-language webpages from 2021 to 2026.
Is it true that a third of web pages are AI-written?
No. That figure does not appear in Pew's study; some press coverage inflated the findings. The published number is roughly 10% of the commercial .com domain, with other domains much lower.
Which words reveal AI-generated text?
The study's English highlights include delve, testament, tapestry, landscape, pivotal, underscore, fostering and intricate, along with heavy use of em dashes, Oxford commas and the "it's not just X, it's Y" structure. The list is English, but the stylistic patterns have Arabic equivalents.
Do the findings apply to Arabic content?
Pew did not study Arabic; its sample is entirely English. But the same economic logic — speed and volume incentives — makes the trend predictable, and the opportunity is for Arabic publishers to stay ahead of the wave with documented human expertise instead of pure generation.
Are AI-text detection tools reliable?
Not at the individual level; the study itself acknowledges errors in both directions. Their value appears when analyzing very large collections of text statistically, not as a final verdict on any single document.
The bottom line
The most important thing in Pew's study is not the number but the direction and the distribution: the commercial web is saturating with generated content at roughly ten times the rate of institutional domains, and the gap between domains is a gap of incentives, not technology. For content makers, the message is anticipatory: readers and search engines are learning fast to recognize "generated tone," and value is migrating toward demonstrated, documented expertise. Whoever starts today building content no model could produce — with lived experience, verified numbers and primary sources — will collect the trust others are losing. That is exactly the standard we hold ourselves to in our AI tools reviews on Truescho.
Sources
- Original study: How Much of the Internet Is Written With AI? — Pew Research Center
- Corroborating coverage: TechCrunch — A third of webpages published since ChatGPT's launch show signs of AI authorship
- Community discussion: Hacker News thread #49386851