22% of citation leverage in AI search sits outside your website.
Earned coverage and social citations, combined, pull harder than any single thing you can do to a page you control. That's from 1.13 million prompts we ran across ChatGPT, Claude, Perplexity, Grok, Gemini, and Google AI Mode between February and June 2026 on the Goodie plaform.
Meanwhile most of the AEO advice in circulation is about page rewrites.
I've spent the last two months getting the same question after workshops and in DMs, always some version of the one I set up in Marketing to Agents back in May. That piece argued agents are a buyer persona nobody in your org owns. Fine. So what do I change on the page?
I went looking for the answer in our own data. It isn't the one I expected.
Structure your content well enough to be read, then stop. Put the answer near the top of the page and in the first sentence of each section. Name your subject instead of pointing at it with "this" or "it." Keep the number in the same paragraph as the claim it supports. That's most of the available gain from restructuring, and it's a floor you clear once. The larger share of citation leverage lives off your site, in earned coverage and social, where almost nobody is spending.
What we actually measured
The fourth edition of our AEO Periodic Table ran 1.13 million prompts through six AI surfaces across 31 industries, scored every citation outcome against a candidate factor list defined in advance, and consolidated it into fourteen factors with explicit weights. It sits on a broader dataset of more than 58 million citations.
It's correlational. Not a controlled experiment, not causal, and the factors move together in ways observational data can't fully separate. We say so in the study and I'll say it here” directional, not decisive.
But 1.13 million prompts is enough to tell you where a category has its priorities backwards.
Where content structure genuinely helps
Answer placement works. Leading with the answer and keeping statements self-contained gives a real lift, because engines retrieve passages rather than pages, and a passage that depends on the three paragraphs above it can't be lifted cleanly.
That's the mechanism, and it's true. Four things follow from it, and they're worth doing:
Name the subject. A section that opens on "This," "It," or "The platform" resolves to nothing once it's pulled out of context. Say the entity.
Answer one question per section. Heading asks it, first sentence answers it, everything after is proof.
Keep the evidence with the claim. Number, date, and source in the same paragraph as the assertion they support.
Make it survive the lift. Paste a section into a blank document. Still true, still complete, still attributable? It passes.
Run that on your top ten revenue pages. It'll take an afternoon.
Then stop, because the next part is where the industry loses the plot.
Where the advice runs past the evidence
Answer-First Structure & Extractability sit in the lower middle of our table. It earns its place. It sits well below the substance, originality, and authority signals doing the heavier lifting.
The rewrite-everything posture — chunking every article, "AI syntax," restructuring long-form for extractability — correlates in our data more with the search rank and content quality underneath the page than with the rewrite itself. Teams restructure a page, the page was already good and already ranking, citations follow, and everyone credits the restructuring.
And the effect varies by content type in a way nobody mentions. Reference and how-to content responds well to sharper answer-first structure. Long-form editorial rarely does, and sometimes loses something in the rewrite.
That last part deserves more attention than it gets. If you take a piece of argued, textured writing and chop it into extraction-optimized blocks, you can end up with something a model can parse and nobody wants to read. You've optimized the passage and damaged the reason anyone would cite the page.
The honest version: structure is hygiene. Get it right, keep it right, don't confuse it with a growth strategy.
Originality beats schema, and it isn't close
Originality & Information Gain scores 86.7 across engines and carries 9% of total weight. Structured Data & Machine-Readability scores 75.8 and carries 4%.
A brand publishing schema-rich derivative content loses to a brand publishing schema-light original research. Schema is worth doing — it helps engines parse entities and qualifies you for rich results. It doesn't manufacture citations, and the most rigorous 2026 tests show engines reading visible HTML during retrieval while largely ignoring the JSON-LD.
Same story for llms.txt. There's no evidence today that publishing one changes whether you get cited, and Google has said it doesn't use it. I'm more optimistic about it than most people are — as a readme for machines on a large site, it's directionally sensible and costs nothing. Just don't expect it to move anything this quarter.
The file that actually matters is robots.txt, and it matters because its effect is binary. Block GPTBot, ClaudeBot, PerplexityBot, or Google-Extended and you're invisible to that engine no matter how good everything else is. Audit it first. It earns no citations on its own, which is exactly why it never makes the listicles.
The gap nobody is funding
Here's where the leverage actually sits.
Earned citations account for 72.2% of all citations in our social citations study, built on 6.1 million citations across 10 AI platforms. Social content drives 2.31x more AI citations than owned content across the full period, and 4.17x in the peak month.
Earned and social together hold 22% of total weight in the periodic table — more than any single on-page content factor.
Now look at how your program is staffed. If more than 70% of your AEO effort is on-page content work, you're mismatched against the data, and you're not unusual. SEO, PR, and social report to different leaders in most companies and share no analytics layer, so nobody can see the whole picture, let alone act on it.
That's the arbitrage. It gets fixed on a budget line rather than in a content brief.
The off-site picture also fractures by engine in ways an averaged program can't catch. X citations come from Grok 99.75% of the time. YouTube concentrates 82.47% into Google's surfaces. A brand with no presence on X is effectively absent from Grok regardless of how strong its site is.
And it's fragile. When Reddit sued Perplexity in October 2025, Perplexity's Reddit share of social citations fell from 19.51% to 2.67% — an 86% drop, with YouTube absorbing the difference almost overnight. Access rules move citation share faster than content quality does.
Rank is the best predictor and the worst lever
For engines grounding on live search, conventional organic rank is the strongest single correlation of citation we can observe. Ignoring that would be dishonest.
Treating it as a dial is the trap. Rank is downstream of relevance, authority, structure, and freshness. It's the scoreboard those inputs produce.
Fan-out is the part of this most teams haven't adjusted for. Engines expand one question into related sub-questions before assembling an answer, and Google's AI Mode is built around it — one thorough page covering a question and its follow-ups beats several narrow keyword pages. Your keyword tool will miss a lot of the queries that win you citations. They aren't in the index because nobody types them. The model writes them.
Domain authority has the same proxy problem as rank. Some analyses find it among the strongest predictors, others find almost nothing, and the reconciliation is that it rises alongside the things actually causing citations — entity recognition, earned coverage, primary-source presence. Optimizing for a proxy is a category error.
Optimize the causes. The proxies follow.
What the fix actually looks like
Here's a section from the May piece, live right now:
Layer 1: Are you reachable? "This is the boring infrastructure layer where most brands lose before the game starts."
Lift that out and it names nothing. The heading asks a question with no subject. The first sentence opens on "This," pointing back at a paragraph the engine didn't take.
The repair:
Layer 1: Can AI crawlers reach your site? "Reachability is the first thing an agent checks. It reads your robots.txt, your sitemap, and your IP-level bot protections, then decides whether to spend tokens on you at all."
Same claim, same voice, one version quotable.
That's the whole intervention. Ten minutes a page, real, small. What it isn't is the reason that article gets read or cited. That article travels because it made an argument nobody else had made yet, and because people linked to it.
Structure is seam work. Keep it in proportion.
What to do this quarter
Five moves, in order.
Audit robots.txt. Binary, consequential, ten minutes. Do it before anything else. Goodie's agent site audit checks crawler access automatically if you'd rather not read the file yourself.
Clear the structure floor on your revenue pages. Answer near the top, subjects named, evidence beside claims. One afternoon, then stop.
Build a source map. Take your top 50 category prompts, record which sources every engine cites, group by domain. The output tells you which surfaces you must be on and which ones competitors are using that you aren't. Higher leverage than any page optimization on this list.
Move budget off-site. Earned coverage, video, community. Reddit and YouTube are the two most cited social surfaces in our data and the two most neglected in most plans. If your split is 70/30 toward owned pages, invert it toward the middle and watch what happens.
Publish something only you could publish. Originality carries more than twice the weight of schema. Original data, a real framework, a genuine POV. This is the one that compounds.
How to measure it
Not with referral traffic. Click-based analytics fold AI surface traffic into general web traffic and won't tell you how often a given engine cites you.
Build a fixed set of 30 to 50 prompts that matter in your category. Record citation share per engine before you change anything. Make changes. Re-measure on the same list after 30 days. Fixed prompt set, fixed window, one variable at a time.
Measure per engine, never averaged. Claude and Gemini weight author authority far higher than AI Mode does. Perplexity and Grok reward freshness roughly ten points more than ChatGPT or AI Mode. A program built on the average underperforms one built per surface.
I wrote the longer version of this in How to Measure AI Search Visibility. If you're choosing tooling to automate it, the AI search visibility monitoring tools buyer's guide covers what's worth paying for.
