Guides
How Publishers Get Cited in AI Answers (2026 Guide)
Ranking in Google and getting quoted by ChatGPT are different games. Here is how AI answer engines actually choose sources, and what publishers can change this month.
Ten years of SEO taught publishers to compete for a position on a page. That skill is now only half the job. When someone asks ChatGPT or Perplexity a question, there is no page of ten blue links to win. There is one answer, assembled from a handful of sources, and either you are one of them or you are invisible.
The frustrating part is that the two things come apart. You can rank third for a query and never get quoted. You can rank eleventh and get quoted constantly. I have watched both happen on the same site in the same month.
So what actually decides it?
Answer engines quote passages, not pages
This is the single most useful thing to understand, and it changes almost everything downstream.
When a model answers a question, it does not read your article the way a person does. A retrieval system pulls a small number of chunks from a much larger index, and the model writes an answer from those chunks. A chunk is typically a few hundred words: a section under a heading, a table, a list.
Your page is not the unit of competition. Your paragraphs are.
That means a 3,000-word feature with a beautiful narrative arc can lose to a 600-word explainer, because the explainer has a section that answers the question cleanly and the feature buries the same information in paragraph nine of a story about a person in a café.
It also means your job is not to be comprehensive. It is to be extractable.
The five things that make a passage extractable
Answer first, then context. Under every heading, the first two or three sentences should answer the question that heading implies. Everything else can follow. If your section titled “How much does article narration cost?” opens with a paragraph about the history of synthetic speech, you have handed the citation to whoever answered the question in their first line.
Real numbers with real provenance. “Audio engagement has grown considerably” is unquotable. “Audio completion rates on our 4,200 narrated articles averaged 61% in Q2 2026” is quotable, and it is quotable specifically because it is checkable. Attach a date and a source to every figure. Models are trained to prefer attributable claims, and so are the humans who fact-check them.
Tables and lists as HTML. If your comparison lives in a screenshot, it does not exist. A real <table> element is one of the most reliably extracted structures on the web. Same for <ul>. This is not an aesthetic preference; it is the difference between machine-readable and not.
Say plainly who you are. Entity clarity matters more than most publishers think. Every page should make it unambiguous who published it, who wrote it, and what they know. An author byline with a real bio, an organisation schema block, a consistent name across your site and every external profile. Models resolve entities across sources, and ambiguity is expensive.
Structured data as a summary layer. Article, FAQPage and Organization schema will not get you cited on their own. What they do is give a parser an unambiguous version of what your page claims, which reduces the chance it gets your claim wrong or attributes it to someone else. Google mostly stopped showing FAQ rich results for commercial sites, so the SEO value has thinned. The machine-readability value has not.
Publishers already own the hard part
Here is what I find genuinely encouraging about this shift.
The things answer engines reward are things newsrooms have always produced and SaaS blogs mostly cannot fake. Original reporting. Named humans with verifiable expertise. Numbers you gathered yourself. An archive with years of consistent coverage on a beat.
A vendor blog can write “the best CMS for local news” from desk research. A local news publisher can write it from having migrated twice. When a retrieval system has both in its index, the second one tends to survive the cut, because it contains claims that appear nowhere else.
If you publish anything proprietary at all, subscriber data, audience surveys, internal benchmarks, put a slice of it on the open web in a table with a date on it. That single act does more for AI visibility than a quarter of keyword work. It creates a fact that can only be sourced to you.
The same logic applies to your beat archives. Depth on a narrow topic beats shallow coverage of a broad one, because retrieval is semantic. Fifteen articles on regional water policy make you the obvious source for regional water policy questions in a way that one article each on fifteen topics never will.
The citations you do not control
Some of your AI visibility is not on your site at all.
Ask ChatGPT to recommend tools or sources in almost any category and watch where the answers come from. A lot of it is forum discussion, community threads, review platforms and reference sites. These are heavily represented in training data and in live retrieval, and they are places where you are talked about rather than places you publish.
For a publisher this means a few unglamorous things are worth real effort. Being correctly described on reference sites. Being listed accurately in industry directories. Having staff who participate honestly in the communities where your subject matter is discussed, under their own names, without link-dropping. Having reviews somewhere that a model can read.
None of this feels like content strategy. It works anyway. If you want the full picture of how audio fits into a publisher’s workflow (/for/news-publishers), the same principle applies there: what other people say about your operation shapes what models say about it.
Measuring this without fooling yourself
This is where I have to be honest with you, because most advice on the topic is not.
Search Console will not tell you whether you are being cited. There is no impressions column for ChatGPT. Anyone selling you a precise AI visibility score is selling you a proxy dressed as a measurement.
What you can actually do:
- Build a list of 20 to 40 questions your audience genuinely asks, the kind that map to your beat. Run them through the major assistants once a month. Log which sources get named. It is manual and it feels crude, but it is real data and it takes an hour.
- Watch your referral traffic for assistant domains. The volumes are small and they undercount badly, since most people read the answer without clicking. Track the direction rather than the absolute number.
- Track branded search separately. If your organisation name starts appearing in search queries that you did not buy, something upstream is introducing people to you. That is often the first measurable trace of AI citation.
And accept a longer feedback loop than you are used to. Retrieval indexes update on their own schedule. A page you publish today may not surface in an answer for weeks.
Where audio actually fits, and where it does not
I want to be careful here, because there is a lazy version of this argument and I do not want to make it.
Audio files are not what gets cited. A model retrieving an answer is reading text. Publishing a narrated version of an article does not, by itself, make that article more quotable.
What does matter is what surrounds the audio. A narrated article with a full transcript on the page gives retrieval systems more text to work with, and transcripts of interviews or panel discussions often contain quotable specifics that exist nowhere else on the web. Summary text generated alongside narration tends to be exactly the answer-first structure described earlier, which is why it extracts well. We went into this in more detail in our piece on summaries and transcripts as article assets.
The honest framing is that audio is an audience play first and a visibility play second. It keeps readers on a page longer and reaches people who will not read. If it also produces transcript text that happens to be extractable, treat that as a useful side effect rather than the reason to do it.
What to do this month
Pick your ten highest-traffic evergreen articles. For each one, rewrite the opening two sentences under every subheading so they answer the implied question directly. Convert any comparison buried in prose into an HTML table. Add a date and a source to every number. Make sure the byline links to a real bio.
That is a day of work. It will do more than a rebrand.
Then publish one piece of original data. Something only you can know. Put it in a table, date it, and let people cite it.
FAQ
What is generative engine optimization?
Generative engine optimization is the practice of making content more likely to be retrieved and quoted by AI answer engines such as ChatGPT, Perplexity and Google’s AI Overviews. It overlaps heavily with SEO but optimises for passage-level extraction and citation rather than page-level ranking.
Is GEO different from SEO?
They share most of the same foundations: crawlable pages, clear structure, genuine expertise. The differences are in emphasis. SEO optimises a page to rank; GEO optimises passages within a page to be extracted and attributed. A page can do one well and the other badly.
Does schema markup help with AI citations?
Indirectly. Schema does not cause a citation, but it gives parsers an unambiguous machine-readable version of what your page claims, which reduces misattribution and factual drift. Article, Organization and FAQPage are the useful ones for publishers.
How long does it take to see results from GEO work?
Longer than SEO, and with less visibility into the process. Retrieval indexes refresh on their own schedule and there is no equivalent of Search Console to watch. Plan on measuring in months, using manual prompt testing rather than a dashboard.
Can I track whether ChatGPT is citing my site?
Not precisely. There is no official reporting. The practical approach is monthly manual testing against a fixed list of questions, plus watching referral traffic from assistant domains and changes in branded search volume. All three are proxies.
Does adding audio to articles improve AI search visibility?
Not on its own. Models retrieve text, not audio. The indirect benefit comes from transcripts and summaries published alongside the audio, which add extractable text to the page. Audio is primarily an audience and engagement decision.