— AEO Readiness Audit
An AEO readiness audit measures citation share below the surface, where schema markup alone does not reach.
Most AEO checklists measure the surface: schema tags, meta fields, a sitemap. Ahrefs tested schema markup on 1,885 pages against 4,000 controls and found no citation lift on AI Mode or ChatGPT, and a decline on AI Overviews. This guide introduces the Sounding, Dotfusion's own audit method, and the five-layer Waterline Framework it scores against, from the Answer above the surface to the Operations that keep it current below.
What the data shows
- Ahrefs tested schema markup on 1,885 pages against 4,000 matched controls (May 11, 2026) and found no measurable citation lift on Google AI Mode or ChatGPT, and a statistically significant 4.6% decline on Google AI Overviews.
- A February 18, 2026 Search Engine Land analysis of 18,012 ChatGPT citations found that 44.2% come from the first 30% of a page's content, and 78.4% of question-tied citations come from headings.
- Ahrefs found essentially no correlation (Spearman r equals 0.04) between word count and Google AI Overview citation across 174,048 pages, in a study dated December 3, 2025.
- Ahrefs found a 34.5% lower average clickthrough rate for the position-1 organic result when a Google AI Overview is present, across 300,000 keywords, dated April 17, 2025.
- 68.01% of US Google searches resulted in zero clicks from January to April 2026, per SparkToro citing Similarweb data, while the share of US desktop devices classed as heavy Google users rose from 84% in Q1 2023 to 87% in Q1 2025, desktop only.
- Across 167.9 million Google AI Overview citations measured June 15 to August 15, 2026, pages outside the organic top 10 supplied 57% of all citations, more than every ranked position in the top 10 combined, per Conductor.
An AEO readiness audit scores five layers of AI search readiness beneath the surface, because schema markup alone produces no measurable citation lift. A website can be beautifully built, fast, and clean, and still invisible to ChatGPT and Gemini. That is a scoring problem, not a build problem. Most AEO checklists measure what sits on the surface, schema tags, meta fields, a sitemap, and call it done. The data says otherwise: schema markup on its own produced no citation lift and a small decline on Google AI Overviews when Ahrefs tested it against a live sample. For the fundamentals of what answer engine optimization actually is, we wrote a guide to that here. This post is about what a real audit measures, and why the fix usually lives well below the surface.
We call the AEO readiness audit a Sounding. Like a ship's crew measuring the water beneath the hull before choosing a course, a Sounding measures how deep your AI search readiness actually runs, not just what shows on the surface. The output is a one-page depth chart: one score, one owner, and one set of metrics for each layer. Behind it sits a depth model, Dotfusion's own Waterline Framework, cited here as Dotfusion, The Waterline Framework, 2026, that separates what you measure from what you build.
Why does surface scoring fail?
Adding schema markup to a website produced no measurable citation lift on Google AI Mode or ChatGPT, and caused a statistically significant 4.6% decline on Google AI Overviews, when Ahrefs tracked 1,885 pages against 4,000 matched controls between August 2025 and March 2026, in a study published May 11, 2026.
Ahrefs ran four separate statistical tests on that sample: a t-test, a difference-in-differences model, an event study, and a symmetric-window difference-in-differences check. All four converged. The change on AI Mode was +2.4%, and the change on ChatGPT was +2.2%. Both are statistically indistinguishable from zero. "We tracked 1,885 web pages that added JSON-LD schema between August 2025 and March 2026, matched them against 4,000 control pages, and measured citation changes across Google AI Overviews, AI Mode, and ChatGPT," Ahrefs wrote in its May 2026 report.
We want to be precise about what that 4.6% decline means, because it is easy to overstate. The absolute effect is small, roughly a dozen fewer daily citations per page, and both the treated pages and the control pages were declining before the test even started. That confounds a clean causal read. Schema markup, added on its own, earns you nothing you did not already have, and the data does not show that it hurts you.
That single finding is why most surface-level audits miss the point. A checklist that stops at whether you added schema cannot tell a CMO whether the brand shows up in an AI answer next quarter. This post is about what to measure once the fundamentals from our AEO guide are already in place.
What does the evidence actually reward?
Search Engine Land's analysis of 18,012 verified ChatGPT citations, published February 18, 2026 and reporting Kevin Indig's own research, found that 44.2% of citations come from the first 30% of a page, and 78.4% of citations tied to questions come from headings, a ChatGPT-only measurement.
A second Search Engine Land article, published March 24, 2026 and reporting Kevin Indig's research on citation data from Gauge, covers roughly 98,000 citation rows across 1.2 million ChatGPT responses in seven verticals, and adds two more ChatGPT-only data points. Within a topic, roughly 30 domains capture 67% of citations; within product comparison topics specifically, the top 10 domains alone capture 46%, a figure scoped to that one vertical, not a general within-a-topic number. Pages ranking first in Google were cited by ChatGPT 43.2% of the time, 3.5 times more often than pages ranked beyond position 20. That figure is a ChatGPT-specific correlation, not a lever to pull directly: across 167.9 million Google AI Overview citations measured between June 15 and August 15, 2026, 57% went to pages outside the organic top 10, more than every ranked position in the top 10 combined (Conductor, Sept 25, 2026). Ranking helps your odds; it is not the path to citation on its own. Pages over 20,000 characters averaged 10.18 citations, against 2.39 for pages under 500 characters.
Put those two studies side by side and a pattern holds: heading structure, front-loaded direct answers, and page depth all move the needle, and ranking position correlates with citation without being the lever that earns it. Word count and schema tags, on their own, do not. The practical reframe: topic-cluster coverage, how deeply and credibly a site covers the query space, together with a self-contained, front-loaded answer, is what both ranking and citation reward at once. Chase that, not the rank itself. Ahrefs found essentially no correlation, Spearman r equals 0.04, between word count and Google AI Overview citation across 174,048 pages, in a study dated December 3, 2025 specific to that one surface. And as established above, adding schema markup produced no lift on AI Mode or ChatGPT and a 4.6% decline on AI Overviews, with the causal caution that both the small effect size and the pre-existing decline in both groups make the negative result modest, not damning.
| Moves citation | Does not move citation, or is not the lever |
|---|---|
| Answer placed in the first 30% of the page (44.2% of ChatGPT citations, ChatGPT-only, Search Engine Land, Feb 18, 2026) | Adding schema markup alone (no lift on AI Mode or ChatGPT, a 4.6% decline on AI Overviews, Ahrefs, May 11, 2026) |
| Question-led headings (78.4% of question-tied citations, ChatGPT-only, same Search Engine Land study) | Word count on its own (Spearman r equals 0.04 against Google AI Overview citation, 174,048 pages, Ahrefs, Dec 3, 2025) |
| Page depth (10.18 average citations for pages over 20,000 characters vs. 2.39 for pages under 500, same study) | Organic ranking position on its own: 57% of Google AI Overview citations go to pages outside the organic top 10, more than every ranked position in the top 10 combined (Conductor, Sept 25, 2026). ChatGPT's 43.2% citation rate for number-one ranked pages (3.5 times pages beyond position 20, Search Engine Land, Mar 24, 2026) is a correlation on one engine, not a lever you can pull on its own |
Source: Search Engine Land (Danny Goodwin, reporting Kevin Indig's research), Feb 18, 2026; Search Engine Land (Danny Goodwin, reporting Kevin Indig's research, citation data from Gauge), Mar 24, 2026; Ahrefs, Dec 3, 2025 and May 11, 2026; Conductor, Sept 25, 2026.
What is the Sounding, and what are the five layers below the waterline?
A Sounding measures how deep your AI search readiness actually runs, across five layers, and produces one depth chart with one owner and one set of metrics per layer, so no team is left guessing what to fix or who owns it.
We named it after the oldest way sailors measured what they could not see: dropping a weighted line to find the depth of water beneath the hull before choosing a course. The depth model underneath the Sounding is our own, cited here as Dotfusion, The Waterline Framework, 2026. It sorts everything that determines whether AI engines cite you into five layers, from what sits above the water to what runs underneath it.
Exhibit 1 is the depth chart itself.
| Layer | Owner | Metrics |
|---|---|---|
| The Answer (above the waterline) | CMO | Citation share on the priority query set; AI referral traffic; accuracy of what AI says about you |
| The Content | Editorial lead | Priority query coverage; authorship coverage; source density; median age of top pages |
| The Structure | Content architecture and development | Schema coverage; structured field population; parse cleanliness on a sampled crawl |
| The Platform | CIO, CDO, or engineering lead | Default versus manual ratio for schema, delivery, and entity data |
| The Operations (the seabed) | COO or marketing operations | Cycle time; items published outside process; measurement cadence held |
Source: Dotfusion, The Waterline Framework, 2026.
The Answer
The Answer is the only layer above the waterline: the extracted answer, the citation, and the brand mention that show up when someone asks an AI engine a question, and it is where you measure the Sounding, not where you build anything.
Everything below the waterline exists to earn what shows up here. The CMO owns this layer because the metrics are business metrics, not technical ones: citation share on a priority query set, AI referral traffic, and whether AI engines describe the brand accurately. A depth chart with a weak Answer score and strong scores everywhere else usually means the lower layers are healthy but under-measured. That is a different problem than a broken content or structure layer, and it calls for a different fix.
The Content
The Content layer covers semantic coverage of the query space, direct answers placed up front, question-led headings, named authors, sourced claims, and entity consistency, and it is where the ChatGPT-only heading and placement data from the previous section applies directly.
The editorial lead owns this layer, with four metrics: priority query coverage, authorship coverage, source density, and the median age of top pages. That last metric needs a caveat. AI assistants across ChatGPT, Perplexity, Gemini, and Copilot combined cite content that is on average 25.7% fresher than organic search results, 1,064 days old versus 1,432, across roughly 17 million citations tracked by Ahrefs, published July 28, 2025. But Google AI Overviews is the one surface that skews slightly older than organic results, plus 16 days, so freshness is not a blanket rule. It is an engine-specific preference, strongest for ChatGPT, and a content refresh plan treats it that way rather than assuming every AI surface rewards recency equally.
The Structure
The Structure layer covers the content model, template-level schema, heading hierarchy, FAQ blocks built as discrete units, clean HTML, and a llms.txt file, and it is the layer most likely to be mistaken for the whole audit.
Content architecture and development owns this layer. Its metrics are schema coverage, structured field population, and parse cleanliness measured on a sampled crawl. Here is the distinction that matters most in this entire post: structural cleanliness is about whether a machine can parse your page correctly. Citation lift is about whether an AI engine chooses to cite it. These are different claims, and the evidence does not let us conflate them. Ahrefs' May 2026 test of 1,885 pages against 4,000 matched controls found no measurable citation lift on Google AI Mode or ChatGPT from adding schema markup, and a 4.6% decline on Google AI Overviews, with the same causal caution noted earlier: the absolute effect is small and both groups were already declining before the test began.
So we score this layer, and we remediate it, because a machine that cannot parse your page cleanly will not represent it accurately even when it does cite you. We do not sell it as a citation lever, because the data does not support that claim.
The Platform
The Platform layer asks one question: is machine-readable output the default in your content management system, or a manual act someone has to remember to do, and the answer determines whether every other layer's gains hold up over time.
A CIO, CDO, or engineering lead owns this layer, measured by the ratio of default versus manual handling for schema, delivery, and entity data. Composable, API-first platforms tend to make structured output a property of the content model itself, so every new page inherits it. Legacy and template-bolted platforms tend to make it a task, and tasks get skipped under deadline. We cover the platform side of this question in more depth in our guide to CMS selection for answer engine optimization, and our headless CMS practice is built around making that default, not manual.
The Operations (the seabed)
The Operations layer is the supply chain that keeps content current, governed, and measured, including any agent workflows that touch it, and it sits at the seabed because nothing above it stays fixed once you stop tending it.
A COO or marketing operations lead owns this layer. Its metrics are cycle time, the number of items published outside the defined process, and whether the measurement cadence set for the Answer layer is actually held. A brand can score well on a single Sounding and still lose citation share within two quarters if nobody owns the cadence of republishing, re-checking, and re-measuring. We go deeper on what a governed content supply chain looks like in our guide to content operations for enterprise teams.
How do you run the surface measurement?
A surface measurement starts with a priority query set of 20 to 40 questions your buyers actually ask, sampled three times per engine, and a query counts as won when your brand is named in three of three samples on both engines measured.
Our panel for this measurement covers ChatGPT and Gemini. It does not cover Perplexity, and we exclude it in part because it behaves differently from the other engines: Perplexity's cited URLs overlap with Google's top 10 at 28.6%, compared with an average cross-engine overlap of only about 11.9% across ChatGPT, Gemini, Copilot, and Perplexity combined, per Ahrefs, August 11, 2025. Folding Perplexity into a blended number would hide that difference rather than reveal it. The three-samples-per-engine approach exists because AI answers are not static: the same question asked twice can return a different citation, and a single sample tells you almost nothing about your actual citation share. We built a fuller version of this measurement, including how to weight and re-run it over time, in our AEO ROI measurement framework.
Why does this matter now?
68.01% of US Google searches resulted in zero clicks between January and April 2026, up from 60.45% for all of 2024, while the share of US desktop devices classed as heavy Google users, ten or more searches a month, rose from 84% in Q1 2023 to 87% in Q1 2025, desktop only.
Read those two numbers together, not apart. Zero-click search is rising and, on desktop specifically, so is the intensity of Google use. That is not a contradiction, and it is not evidence that search is dying. It is evidence that search behavior is splitting: more of the answer gets delivered on the results page or inside an AI engine, and the people who search heavily are searching more, not less. The traffic cost of that shift is already measurable at the individual keyword level: Ahrefs found a 34.5% lower average clickthrough rate for the position-1 organic result when a Google AI Overview is present, across 300,000 keywords, in a study published April 17, 2025. Exhibit 3 puts the broader search-behavior figures side by side.
| Metric | Period | Figure | Scope |
|---|---|---|---|
| Zero-click share of US Google searches | January to April 2026 | 68.01% | All devices, per Similarweb clickstream data |
| Heavy Google users (10 or more searches a month) | Q1 2023 to Q1 2025 | 84% to 87% | Desktop only. Mobile devices are explicitly excluded from this figure |
Source: SparkToro (Rand Fishkin), citing Similarweb data, June 9, 2026, and Datos panel data, August 26, 2025.
For an enterprise brand, the practical read is this: you are not choosing between AEO and traditional search. You are competing for citation share inside a search behavior that already includes both, and a Sounding tells you where you actually stand in that split, not where you assume you stand.
What does a CMO, CDO, CIO, and COO each do first?
Each of the five Waterline layers has one named owner, from CMO to COO, and each owner has exactly one first action to take now, an action that does not depend on the other three layers being finished first or on the audit being complete.
The CMO owns The Answer. Run the priority query set, 20 to 40 questions, across ChatGPT and Gemini this month, and record a citation-share baseline before any remediation starts. Without a baseline, no one can prove the Sounding worked.
The CDO supports The Content layer's entity consistency and source density metrics, which the editorial lead owns per the Depth Chart. First action: work with the editorial lead to audit the top 20 pages by traffic for named authorship and sourced claims. That single audit usually reveals which pages are citation-ready and which are not, fast.
The CIO owns The Platform. Determine whether the current CMS makes schema and structured delivery the default, or a manual, inconsistent act that depends on someone remembering. That answer predicts whether gains from the other layers hold or decay.
The COO owns The Operations. Set a measurement cadence, how often the priority query set gets re-run, before anyone treats a single Sounding as a finished project. A Sounding is a snapshot. Citation share moves.
What do three remediation scenarios, sequenced by depth, look like?
Remediation works in the order the water gets deeper: fixing the Answer and Content layers first, Structure and Platform second, and Operations last, because operations is what keeps the first two layers from decaying once the initial fix ships and attention moves elsewhere.
Exhibit 4 sequences three scenarios by how deep the remediation goes, using the citation data already covered above rather than a projected outcome we have not measured ourselves.
| Scenario | What the data shows | What gets fixed first |
|---|---|---|
| Surface only: schema added, nothing else touched | No measurable lift on AI Mode or ChatGPT, a 4.6% decline on AI Overviews (Ahrefs, May 11, 2026) | Nothing. This is the scenario the evidence rules out as a strategy on its own |
| Answer and Content addressed: coverage, headings, and front-loaded answers fixed | Citation concentrates around topic depth and structure, not rank: roughly 30 domains capture 67% of ChatGPT citations in a topic (Search Engine Land, Mar 24, 2026), and across Google AI Overviews specifically, 57% of citations go to pages outside the organic top 10 (Conductor, Sept 25, 2026). Rank correlates with citation; it is not what earns it | Query coverage, heading structure, and answer placement in the first 30% of the page |
| Full depth: Structure, Platform, and Operations added on top | Holding citation share requires a held measurement cadence, since a single Sounding is a snapshot and citation share moves between measurements | Schema and delivery defaults at the platform level, and a governed republishing and re-measurement cadence |
Source: Ahrefs, May 11, 2026; Search Engine Land, Mar 24, 2026; Conductor, Sept 25, 2026; Dotfusion, The Waterline Framework, 2026.
Method note
Every figure and date in this post is current as of October 5, 2026, the date every source below was last fetched directly. The surface-measurement panel described in this post covers ChatGPT and Gemini only, not Perplexity, and this post treats AEO as a citation play rather than a volume play, meaning the goal is earning citation share inside existing search behavior, not replacing search demand.
Two figures in this post, the 44.2% first-30%-of-page figure and the 78.4% heading figure, are ChatGPT-only measurements from the February 18, 2026 Search Engine Land article reporting Kevin Indig's own research; Gauge is not named in that article. The 43.2% and 3.5 times ranking figures, and the 46% domain-concentration figure (scoped to product comparison topics), come from the separate March 24, 2026 Search Engine Land article, which names Gauge as the source of the citation data. None of these figures are blended across AI engines, and we have kept both the engine scope and the correct attribution explicit everywhere we cited them. Similarly, the word-count correlation figure, Spearman r equals 0.04, measured Google AI Overview citation specifically, not a cross-engine average.
Finally, the 4.6% decline in AI Overview citations after adding schema markup is not proof that schema markup actively hurts citations. The absolute effect is small, and both the treated and control groups in Ahrefs' study were already declining before the test began, which confounds a clean causal read. We have carried that caution everywhere the figure appears in this post.
This post uses Conductor's direct measurement of 167,867,680 US AI Overview citations collected June 15 to August 15, 2026, 43% from pages ranking in the organic top 10 and 57% from pages that do not, in place of earlier sampled overlap figures that moved sharply between versions of the same study. Conductor's own writeup names that drift as the reason the question needed a larger, unsampled measurement, and we are following their finding as the current, best-available number. Nothing in this post treats organic ranking as the path to AI citation; rank correlates with citation without being the lever that earns it, and the practical fix is topic-cluster coverage and a self-contained, front-loaded answer, not a ranking campaign.
Frequently asked questions
These five questions are the ones we hear most often from enterprise marketing and technology leaders trying to understand why their site is invisible to ChatGPT and Gemini, and each answer below reflects the same evidence and the same Waterline Framework covered in the sections above.
How do I audit my website for AI search readiness?
Run a Sounding: measure citation share on a priority query set of your real buyer questions across the AI engines that matter to you, then score the five layers beneath that result, The Content, The Structure, The Platform, and The Operations, to find out which one is actually limiting you. The output is a one-page depth chart with a metric and an owner for each layer.
What should an AEO audit checklist for enterprise websites include?
A good checklist measures citation share and AI referral traffic first, ahead of schema coverage. It checks content coverage and authorship next, then structural parse cleanliness, then whether structured output is a platform default or a manual task, and finally whether a measurement cadence exists to catch drift after the audit is done.
Why is my company not cited by ChatGPT or Gemini?
Usually because the content does not answer the question in the first 30% of the page, the heading is not written as the buyer's own question, or the topic is not covered deeply enough across the site for an AI engine to treat it as a credible source, not because of where it ranks. Across 167.9 million Google AI Overview citations, 57% go to pages that never reach the organic top 10 (Conductor, 2026), so ranking is not the lever to pull. ChatGPT-only data shows 44.2% of citations come from the first 30% of a page and 78.4% of question-tied citations come from headings.
Does adding schema markup improve AI citations?
No, not on its own. Ahrefs tested schema markup on 1,885 pages against 4,000 matched controls and found no measurable lift on Google AI Mode or ChatGPT, and a small, statistically significant decline on Google AI Overviews. Schema still matters for how cleanly a machine can parse your page, which is a separate claim from citation lift.
What is the Waterline Framework?
It is Dotfusion's own depth model for AI search readiness, cited as Dotfusion, The Waterline Framework, 2026. It sorts what determines AI citation into five layers: The Answer above the waterline, then The Content, The Structure, The Platform, and The Operations below it, each with one owner and one set of metrics.
Where does this data come from?
Every statistic, quote, and figure used in this post is drawn from one of the primary sources listed below, each fetched directly and verified before publication, and they appear here in the order they are first cited above, so any claim in this piece can be checked directly.
- Ahrefs (Louise Linehan), "Does Adding Schema Markup Increase AI Citations?" May 11, 2026. ahrefs.com/blog/schema-ai-citations
- Ahrefs (Despina Gavoyannis), on word count and AI Overview citation, December 3, 2025. ahrefs.com/blog/short-vs-long-content-in-ai-overviews
- Ahrefs (Ryan Law), on AI Overviews and clickthrough rate, April 17, 2025. ahrefs.com/blog/ai-overviews-reduce-clicks
- Search Engine Land (Danny Goodwin, reporting Kevin Indig's own research), on ChatGPT citation study, February 18, 2026. searchengineland.com/chatgpt-citations-content-study-469483
- Ahrefs (Louise Linehan), on AI engine overlap with Google rankings, August 11, 2025. ahrefs.com/blog/ai-search-overlap
- Ahrefs (Ryan Law), on content freshness and AI citation, July 28, 2025. ahrefs.com/blog/do-ai-assistants-prefer-to-cite-fresh-content
- Search Engine Land (Danny Goodwin, reporting Kevin Indig's research, citation data from Gauge), on domain concentration in ChatGPT citations, March 24, 2026. searchengineland.com/chatgpt-citations-domains-study-472349
- SparkToro (Rand Fishkin), citing Similarweb data, on zero-click search, June 9, 2026 (published date per site metadata; the on-page byline in Pacific time reads June 8, 2026, the same instant). sparktoro.com/blog/in-2026-less-than-one-third-of-google-searches-still-send-a-click
- SparkToro (Rand Fishkin), citing Datos panel data, on heavy Google use, August 26, 2025. sparktoro.com/blog/new-research-20-of-americans-use-ai-tools-10x-month-but-growth-is-slowing-and-traditional-search-hasnt-dipped
- Conductor (Jia-Rong Li), "AI Overviews and Organic Rankings: 57% of Citations Come From Outside the Organic Top 10," last updated September 25, 2026, underlying data measured June 15 to August 15, 2026. conductor.com/academy/aio-vs-organic-rankings
This post was conceived and written by Chris Bryce with our stack of AI research agents doing the heavy lifting on sources and data. If you want to talk about enterprise web strategy, AEO, or content operations, talk to us. We love this stuff.