For 20 years, the economics of publishing on the open Web rested on a measurable exchange. A search engine crawled your pages and sent you visitors, and you had instruments for both sides of the trade: server logs on one end, referrer headers and Search Console on the other. You could see what you gave and what you got back, down to the query. The agency I founded in 2015 places editorial coverage across a network of several thousand independent publishers, in more languages than I can read, and every pricing decision in that business, ad rates, content budgets, what a page is worth, was built on top of that visibility.
AI search removes half of the instrument panel. When an assistant retrieves a page and composes an answer, and the user never leaves the chat window, everything after the crawl happens somewhere the publisher cannot observe. The crawl side is still perfectly visible: open a log file and watch GPTBot, ClaudeBot, and their relatives arrive, and in the logs of the sites we manage, their share of requests grows every time I look. The usage side is dark. Was the page retrieved into a context window? Cited? Paraphrased without credit? Seen by 10 people or 10 million? No publisher-side log records any of it, because the event happens on infrastructure the publisher does not run.
The best public data we have is inference from network traffic. Cloudflare began publishing a crawl-to-refer ratio in July 2025: HTML requests from a platform’s crawlers divided by HTML requests arriving with that platform’s referrer. The first figures were memorable. In June 2025, OpenAI’s ratio was about 1,700 pages for every visit it referred back, and Anthropic’s was about 73,000. Cloudflare stated the caveat itself: traffic from native apps often carries no Referer header at all, so the ratios overstate the imbalance by an unknown amount. A year on, independent measurements of the same crawler span a factor of five or more depending on the window and the network doing the counting. When careful people with good data cannot agree on the first digit, that is not a measurement. That is the absence of an instrument.
The demand side is no better. Pew Research Center tracked the browsing of 900 American adults in March 2025 and found that when Google showed an AI summary, users clicked a traditional result on 8% of searches, against 15% without one, and clicked a link inside the summary itself on 1% of visits. Google holds the real numbers and, for two years, folded AI Overview and AI Mode activity into ordinary search totals. The Generative AI report added to Search Console in June 2026 counts impressions only; no clicks, no queries, no positions. It began rolling out in the U.K. first, under a conduct requirement from the Competition and Markets Authority. Partial disclosure, delivered on a regulator’s deadline, is the current state of the art.
Why should this community care about what sounds like a marketing problem? Because the missing layer is now load-bearing for questions that are not marketing at all. Content licensing deals and copyright suits are being negotiated and argued right now, and the central empirical question in all of them, how much a given corpus actually contributes to a model’s responses, can currently be answered only by the defendant. Publishers price their content blind. Courts are asked to weigh harms nobody can quantify.
The vacuum also fills with something worse than ignorance. An industry of AI visibility tools has grown up whose core method is sampling: fire thousands of prompts at the models, count the brand mentions, sell the count as share of voice. My own agency sells work in this category, so I have a commercial stake here, and I will still say plainly that the method sits closer to divination than measurement. Outputs are stochastic and personalized, and they shift with every model update, so the numbers cannot be audited or reproduced. Because nobody can see what actually gets retrieved and cited, generative engine optimization proceeds by superstition, which in practice means more content produced for machines, faster, with less care. Readers here can guess how that loop ends. Search already ran it once, in the spam wars of the 2000s.
None of the missing data is exotic. It sits in provider logs today. A minimal analytics layer would look much like what search settled on twenty years ago: aggregate, per-domain counts of retrievals and citations, reported by platforms to the publishers concerned, the way Search Console reports impressions and clicks without exposing any individual user. Add standardized attribution on outbound clicks from assistant apps, which today arrive with no referrer and get logged as direct traffic. Add third-party audit of the aggregates, roughly what podcast metrics eventually got from the IAB after years of vendor-invented numbers. Google’s impressions-only report proves the reporting pipes can be built. What is missing is a reason to build them properly.
The counterarguments are familiar. Privacy is the weakest, since domain-level aggregates require no user data at all. Competitive secrecy is real, retrieval statistics do reveal something about how a system sources its answers, but impression reports have already crossed that line without incident. The honest objection is incentive: a platform that discloses usage data creates the evidence base for its own licensing bill, so voluntary disclosure will stay minimal, and the first meaningful report arrived only when a regulator attached a deadline. I do not expect that to change on its own.
I would like to see this treated as an infrastructure problem rather than a feature request. The old exchange between crawlers and publishers was never written down as a protocol, but it behaved like one, and a workable content economy grew on top of it. The exchange has changed and the instruments have not. The people who build retrieval systems and measurement standards can define what honest aggregate reporting looks like before courts and regulators define it for them, badly. In my corner of the Web, the sooner the better.
Boris Dzhingarov is the founder of ESBO Ltd, a digital PR agency that has operated a multilingual network of independent publishers since 2015. He writes about search, AI visibility, and the economics of online publishing for Forbes, Entrepreneur, and Fast Company.
Join the Discussion (0)
Become a Member or Sign In to Post a Comment
Facts Only
* The economics of publishing on the open Web historically rested on a measurable exchange between search engines and publishers using server logs, referrer headers, and Search Console data.
* AI search removes visibility because user interactions with assistants occur outside the publisher's observable logs.
* Crawling activity by bots like GPTBot and ClaudeBot is visible in site logs for managing publishers.
* Cloudflare began publishing a crawl-to-refer ratio in July 2025, measuring HTML requests from crawlers versus those arriving via referrer.
* Initial measurements showed OpenAI’s ratio around 1,700 pages per visit and Anthropic’s around 73,000.
* Independent measurements of crawler ratios vary significantly depending on the window and network used, showing factors of five or more difference.
* Pew Research Center found that when Google showed an AI summary, users clicked a traditional result on 8% of searches, compared to 15% without one, and 1% with a link inside the summary.
* The Generative AI report added to Search Console in June 2026 counted only impressions, not clicks or queries.
* A core question in content licensing is how much a corpus contributes to a model’s response, which is currently only answerable by the defendant.
* A proposed minimal analytics layer would include aggregate counts of retrievals and citations reported by platforms to publishers, along with attribution for assistant app clicks.
Executive Summary
The economics of publishing on the open Web historically relied on a measurable exchange between search engines and publishers, facilitated by server logs and referrer headers. This visibility allowed for pricing decisions based on traffic visibility. The advent of AI search removes much of this instrument panel because generative AI processes occur within a context window that is not recorded by the publisher's side. While crawler activity remains visible in system logs, user interactions with AI assistants are opaque to publishers.
Data regarding web crawling is being measured through metrics like the crawl-to-refer ratio, where initial measurements showed significant differences between platforms. Demand-side data shows users exhibit different clicking behaviors when presented with AI summaries compared to traditional search results. A key tension exists because content licensing and copyright disputes rely on determining how much corpus contributes to model responses, a question currently answerable only by the defendant.
The narrative points to an infrastructure problem: the lack of standardized, auditable reporting mechanisms for data exchange between crawlers and publishers. While privacy concerns are noted regarding user data, the argument centers on establishing measurement standards for system interactions rather than just hiding personal information. The proposed solution involves treating this as an infrastructure issue where builders of retrieval systems can define honest aggregate reporting before external bodies mandate it.
Full Take
The core tension in this discussion lies between the historical, measurable exchange that underpinned the open Web economy and the current opacity introduced by generative AI. The argument shifts from a transactional reality—where data flowed between agents—to an epistemological crisis where the actual contribution of content to artificial intelligence is obscured, creating a vacuum for legal and economic disputes. This vacuum is filled by speculative, non-auditable methods of AI visibility, such as sampling prompts, which introduce stochasticity that resists measurement.
The pattern observed is a systemic resistance to accountability in novel technological interactions; systems evolve faster than the protocols governing their interaction. The reliance on aggregate data provided by platforms (like impression reports) shifts the burden of proof onto external regulators rather than providing intrinsic transparency, suggesting an incentive structure where voluntary disclosure remains minimal until regulatory pressure is applied.
This situation suggests a fundamental misalignment: the infrastructure that powers information exchange has evolved into a black box for value assessment. The move toward treating this as an infrastructure problem, rather than a feature request, attempts to restore agency by focusing on measurement standards among those who build the systems. The implication is that if measurement standards are established within the technical community, they may precede and guide legal and regulatory outcomes, addressing the difficulty of quantifying content contribution in copyright and licensing disputes. What remains unaddressed is how to institutionalize this infrastructural view against existing commercial incentives driving opacity.
Sentinel — Human
The text reads like an expert synthesizing complex, real-world economic and technical shifts concerning AI search visibility, using personal authority to argue for a specific infrastructural solution.