
Half the Web Is AI-Generated. Almost None of It Ranks.
A 64,000-URL study found only 14% of AI content gets indexed by Google because search engines reward signal over volume.
I pulled the traffic data for a B2B SaaS company’s blog last month. They’d published 240 posts in 2025, up from 60 in 2024. The quadrupled output came almost entirely from AI. Their content team of two was using Claude and GPT to produce four posts per week instead of one. By every volume metric, the content program was a success.
The traffic data told a different story. Of the 240 posts, 31 were receiving any organic traffic at all. The other 209 were invisible. Not ranking on page two or three. Not ranking at all. Google hadn’t indexed 40% of them, and the ones that were indexed sat beyond position 100. The company had spent a year producing content that almost nobody would ever see.
This isn’t an outlier. A study of 64,000 URLs found that only 14% of AI-generated content appears in Google’s index. The same study checked LLM citation rates and found that only 18% of AI content gets cited by ChatGPT in its responses. The content flood is real. The content is being ignored.
Something fundamental has shifted in how search engines and LLMs evaluate content. Understanding what changed is the difference between a content program that compounds and one that produces expensive noise.
The signal problem

signal problem reframed as system design.
The instinct is to frame this as a quality problem. AI content is generic, so it doesn’t rank. That framing is incomplete. I’ve read AI-generated blog posts that are well-structured, factually accurate, and substantively useful. By any reasonable quality measure, they’re fine. They still don’t rank.
The issue is signal, not quality. Google’s algorithm and LLM citation systems are looking for indicators that content was created from genuine expertise. Original data that exists nowhere else on the web. Specific examples drawn from real experience. Technical details that only a practitioner would know to include, or perspectives that contradict conventional wisdom with evidence.
AI-generated content lacks these signals by definition. An LLM can write a coherent article about B2B pricing strategy, but it draws from the same training data as every other LLM-generated article about B2B pricing strategy. There is nothing in the output that Google can use to differentiate it from the thousands of other AI-generated pricing articles that were published the same month. The content is correct but indistinguishable.
Google has been quietly adjusting its ranking signals to weight originality and information gain. Information gain is the delta between what a page offers and what already exists in the index for that query. When you search “B2B pricing strategies,” the first ten pages of results say substantially the same thing. They list the same pricing models (freemium, usage-based, tiered), cite the same examples (Slack, HubSpot), and offer the same advice (test your pricing annually). An AI-generated article on this topic adds zero information gain because it’s synthesizing the same sources. A practitioner article that says “I tested four pricing models over eighteen months and here’s what happened to our NRR” adds information that literally doesn’t exist elsewhere. That is what ranks.
The 80/20 content rule
The companies still winning with content have converged on a ratio, though most of them arrived at it by trial and error rather than by design. Roughly 80% of their content can be AI-assisted or fully AI-generated. The other 20% must contain something AI cannot fabricate.
The 80% is the foundation layer. Product documentation, feature comparison pages, integration guides, glossary entries, FAQ pages. This content needs to exist and needs to be accurate. AI produces it efficiently. It ranks adequately for branded and long-tail queries where the competition is other documentation pages.
The 20% is the differentiation layer. Original research. Case studies with named customers and specific numbers. Practitioner essays that draw on real experience and proprietary data. Content where the author’s identity and expertise are part of the reason to trust what’s written. This is the content that earns links and LLM citations, ranks for competitive queries, and builds the domain authority that lifts the other 80%.
The mistake I see repeatedly is companies using AI to produce 100% of their content and expecting the quantity to compensate for the missing 20%. It doesn’t. Two hundred AI posts without the differentiation layer perform worse than fifty AI posts plus ten practitioner pieces. The ten practitioner pieces generate the links and authority that make the fifty AI posts indexable. Without that authority anchor, the AI content floats in a void.
Why LLMs.txt is not the answer
A protocol called LLMs.txt has been circulating as a solution. The idea is similar to robots.txt: you put a file on your server that tells LLMs how to index and cite your content. The pitch is that by providing structured metadata, you can influence how AI systems reference your site.
I’ve seen companies invest engineering time implementing this. The evidence that it works is nonexistent. LLM citation behavior is driven by training data, retrieval-augmented generation (RAG) systems, and whatever ranking heuristics the LLM provider uses internally. A text file on your server does not influence any of these. ChatGPT does not check for LLMs.txt before deciding whether to cite your content. Neither does Perplexity, Claude, or Gemini.
The protocol might someday become relevant if major LLM providers agree to support it. That hasn’t happened, and there is no indication it will. In the meantime, implementing LLMs.txt is the equivalent of putting a “please rank me” note in your HTML and hoping Google reads it. The effort would be better spent creating content that LLMs cite for the same reason Google ranks it: because it contains information that doesn’t exist elsewhere.
The company blog is in trouble

company blog is in trouble as a maturity path.
Let’s address the elephant. The traditional content marketing playbook, where a company publishes three to five blog posts per week targeting informational keywords and builds organic traffic over time, is functionally broken for most B2B companies.
Three forces killed it. First, Google’s algorithm changes in 2024 and 2025 explicitly devalued content that exists primarily for search ranking rather than user value. The “helpful content” updates wiped out sites that had been publishing keyword-optimized articles without genuine expertise behind them. Many of these were the exact kind of content that B2B companies had been producing: technically accurate but strategically motivated by keyword volume rather than reader need.
Second, AI Overviews (Google’s AI-generated summaries at the top of search results) are absorbing clicks that used to go to organic listings. For informational queries, the AI Overview often answers the question directly, and the user never clicks through. The queries where blog content used to capture traffic are increasingly zero-click queries. You can rank #1 and still get no traffic because Google answered the question itself.
Third, the content saturation is overwhelming. Ahrefs estimates that over 90% of web pages receive zero organic traffic from Google. That statistic predates the AI content explosion. The percentage is certainly higher now. When every company in a category publishes weekly blog posts on the same topics, the mathematical reality is that most of that content will never be seen.
This doesn’t mean content marketing is dead. It means the easy version of content marketing is dead. The version where you keyword-research your way to a content calendar, assign articles to writers (or AI), and watch the traffic chart go up. That stopped working.
What still works
The content strategies that still produce results in 2026 share common characteristics, and all of them require something AI can’t provide on its own.
Proprietary data content outperforms everything else. Companies that survey their customers, analyze their product usage data, or run original experiments have content assets that are genuinely unique. A report that says “we analyzed 10,000 deals in our CRM and found that proposals sent on Tuesdays close 23% more often” is uncopyable. No LLM can generate that insight because it doesn’t exist in training data. No competitor can replicate it without access to the same data. Google ranks it because it adds information to the index. LLMs cite it because it’s a credible primary source.
Practitioner perspective content is the next tier. This is expert content where the author’s experience is the value. Not “5 best practices for outbound sales” but “I sent 50,000 cold emails last quarter and here’s what I learned about subject lines.” The specificity of the experience creates the signal. AI can write generic best-practice articles. It cannot write “I did this specific thing and here’s exactly what happened.”
The argument has been gaining traction for years: the future of content marketing is fewer, better pieces backed by original data and genuine expertise. The AI content explosion accelerated the timeline. What industry observers predicted would happen over a decade is happening in two years.
Comparison and evaluation content retains value when it’s genuinely objective. Buyers searching “product A vs product B” want an honest assessment, not a vendor’s self-serving comparison page. Companies that publish fair comparisons, including acknowledging where competitors are better, build trust signals that both Google and buyers reward. The counterintuitive finding is that pages acknowledging competitor strengths convert better than pages that don’t, because they pass the credibility test that buyers are applying.
Community-driven content, pieces that aggregate and curate insights from practitioners in a specific space, works because it provides perspectives AI can’t synthesize from training data. A roundup of how ten different SaaS companies handle churn, with specific details from each one, is a primary source that AI cannot replicate.
The content investment shift

content investment shift translated into operating choices.
The practical conclusion is that content budgets need to move from volume to signal. The math that justified hiring three writers to produce fifteen posts a month was based on a world where each post had a reasonable probability of ranking and generating traffic. That probability has collapsed. Fifteen generic posts now produce less traffic than three posts built around original data or practitioner expertise.
The budget that used to fund a content team producing volume should be partially redirected. A smaller portion funds AI-assisted production of the foundation layer: documentation, feature pages, long-tail content. The larger portion funds the activities that create signal: customer research, data analysis, expert interviews, original studies, case study development. The writing itself is the cheapest part. The inputs to the writing, the data, the experience, the access, are where the value lives.
Companies that make this shift early gain compound advantages. Original data content earns links, which builds domain authority, which makes AI-assisted content more likely to index and rank. The 20% lifts the 80%. Companies that keep producing volume without signal will watch their content investment flatline regardless of how much they scale production.
What Google and LLMs actually want
The alignment between what Google rewards and what LLMs cite is not a coincidence. Both systems are trying to surface content that adds something to human understanding. Google has been refining this objective for two decades. LLMs inherited it from training on web data where Google’s preferences shaped what was most visible and most linked.
The shared criteria come down to information gain, source credibility, and specificity. Does this content say something that doesn’t already exist in the corpus? Is there evidence that the author or organization has genuine expertise? Does the content provide concrete details and evidence rather than general advice?
AI-generated content fails on all three by default. It has zero information gain because it’s derived from existing content. It has no source credibility because there’s no practitioner behind it. And it trends toward generality because LLMs produce the statistically most likely output, which is by definition the most generic.
The path forward is clear even if it’s not easy. Use AI for the 80% where generic, accurate content serves a functional purpose. Invest human effort in the 20% that creates signal. Measure content performance by information gain, not by word count. The companies producing the most content are not winning. The companies producing the most differentiated content are.
Half the web is AI-generated. The half that matters isn’t.
Enjoying this essay?
Written by

Elom
GTM, growth, and revenue systems operator with 12 years across Fortune 500s, fintech, and B2B startups. Building at the intersection of AI, data, demand, and revenue.
Get the next deep-dive in your inbox
Essays on demand creation, GTM, growth engineering, and revenue systems. Free.


