
AI answers are leaving the crawled web
Google's own cards became the #2 cited domain inside its AI Mode, and OpenAI licensed Yelp rather than crawling harder. The crawled page is becoming the fallback, not the source.
Two moves last week, from two companies that compete with each other, pointing in exactly the same direction.
Profound measured google.com's citation share inside Google's AI Mode rising 8.4x between April 15 and June 30, making Google the second most-cited domain in its own answers. That share flows almost entirely through Business Profile cards and Product Knowledge Panels — Google's own hosted structured data, not crawled pages. The local-intent analysis behind it covers more than 32 million google.com/searchviewer/ citation instances.
Then on July 23, OpenAI licensed Yelp's reviews, ratings, photos, and Request-a-Quote flow into ChatGPT. Non-exclusive, terms undisclosed. Yelp's CEO Jeremy Stoppelman conceded openly that OpenAI controls how users experience that content.
One company cited its own database more. The other bought somebody else's. Neither solved its weakest queries by crawling the web harder.
What that does to the page
The working assumption behind most AI-visibility work is that an answer engine reads pages, so better pages earn more citations. That assumption is getting narrower, and it is narrowing fastest exactly where commercial intent is highest.
For a local query — a restaurant, a dentist, a plumber — the answer is increasingly assembled from Business Profile fields and a licensed review corpus. For a product query, from product feeds and knowledge panels. Those are database records. They have owners, schemas, update cadences, and access terms. They are not documents you can out-write.
So the page has not stopped mattering. It has been demoted from source of record to fallback: the thing consulted when there is no structured record, or when the structured record is thin. That is a materially different job, and it justifies a materially different budget.
The rule I would hand a client
Not a prediction this time, a working rule, because this one is actionable today:
For any query class where a structured record exists — local, product, review, availability, price — treat the record as the primary surface and the page as support. Audit the record with the same seriousness you audit a landing page: who owns the fields, how fast corrections propagate, what the licence permits, and whether a competitor's record is more complete than yours.
Most organisations have never assigned an owner to their Business Profile fields. Marketing assumes operations maintains them; operations assumes marketing does. Meanwhile that record is now a cited source in the answer your buyer sees, and it is being read more often than the article you commissioned about the same topic.
The uncomfortable part of this rule is that it moves work away from the function that is used to owning it. Content teams write pages. Feeds, profiles, and catalogue data usually sit with operations or a data team who have never been told their fields are a marketing surface.
What I keep seeing in retainers
Where this shows up in my own work is in the scope of AEO engagements. A client buys a retainer, the retainer produces article pages optimised for answer engines, and the reporting shows citations climbing — for informational queries, which is where article pages genuinely still win.
Then you look at the queries that actually precede a purchase in that category, and the cited sources are cards, feeds, and review corpora the client does not control and, in several cases, has never audited. The retainer is working. It is working on the input whose share is shrinking on exactly the queries that matter most commercially.
The conversation that helps is not "stop writing." It is asking a client to name the person who owns each structured record their category depends on. That question usually lands in silence, and the silence is the finding — you cannot optimise a surface nobody owns.
One more signal, held loosely
There is a third data point circulating that I want to pass on carefully because I have not confirmed it independently. SE Ranking, reported via Search Engine Journal, found AI Mode showing text ads on 29% of tested commercial keywords while the advertiser's domain appeared among cited sources only 11% of the time, and the exact paid URL just 1.95%.
If that holds, paid placement does not buy citation — being in the ad slot and being in the answer are separate acquisitions. I would treat those as the study's own numbers rather than settled fact until someone reproduces them, but the direction is consistent with everything else here: what gets cited is decided by what the platform can source structurally, not by what has been paid for or published.
Two other things I could not pin down: Profound's total citation volume across all domains, since the 32-million figure covers only the local-intent subset, and the financial terms of the Yelp deal, which both parties declined to disclose.
The shape of the shift
The open web is not disappearing from answers. It is being repositioned — from the corpus that answers are built out of, to the corpus consulted when the owned and licensed ones fall short. Platforms prefer structured records because they are cleaner, cheaper to serve, and legally settled. That preference is not going to reverse.
Which means the durable question for anyone spending on visibility is no longer how well your page is written. It is whether you appear in the databases the answer is actually assembled from, and whether anyone in your organisation is responsible for keeping those entries true.
Optimising pages while answers are assembled from feeds is tuning the fallback — go find out who owns your records.


