A competitor made a run on ChatGPT for a query you've ranked for the last two years. Perplexity answered the query, citing a blog post by a competitor for a claim you made. Your CEO forwarded you a Semrush thread. Someone on the exec team asked whether next quarter should include "AI SEO."
I've had this conversation seven or eight times in the last six months. Every single one of them asked the same thing: not "should we do generative engine optimization?" but "is any of this a real lever, or just vendor pitch-deck with '96% LLM visibility' as a round number placeholder?".
Note: I have opted not to categorize the responses, like "LLM", "Perplexity", "Claude", "Copilot Search", "Google AI Overview" and "others", given that the average reader should be familiar with what these are. I'm putting my emphasis on the how. What is the mechanism by which the LLMs read your article and decide whether to cite you? What kinds of transformations to the structure of your page could tip the scales? And what sort of measurement could you perform that would let you know how well you did, across the various systems, this week, without having to pay for a service to do it for you? This is what I'm trying to convey in this article.
What AI Engines Actually Do With Your Article Before They Cite You §
All AI answer engines do the same three things: they fetch documents from an index, they sort or filter the documents based on the query, and they compose a response by drawing on one or more of the documents. The difference between the AIs is in how they do the steps, but they all do the same thing with your query and the documents in the index.
Your article's ability to be in the candidate pool for retrieval is what it means to be "indexed". For Google AI Overviews, this means the page is in Google's search index, and Google has no additional technical requirements for AIO or AI Mode beyond snippet-eligible indexing (Google Search Central). For ChatGPT Search, it's a combination of third-party search results (right now, only Bing) and content licensed directly from OpenAI's publisher partners (OpenAI). For Perplexity, the retrieved sources are the pages themselves, not a fixed snapshot of the top results (help.perplexity.com/hc/en-us). If your article is not in the relevant index (i.e., Bing for ChatGPT, Google for AIO, and Perplexity for its own indexes), nothing you write in the article will matter.
Ranking is when the query is broken down into subtopics, and the subtopics are matched against the retrieved candidates. Google, for example, does something similar called query fan-out where they perform multiple searches (on the subtopics and data sources) to construct a single response (Google Search Central). Ranking for a single head keyword is not enough here, as the engine is looking for entities and related questions that are distributed across multiple sub-searches. A thorough article covering all aspects of the entity will be picked more easily than a narrow one.
Composition is where citations are served. The engineering team at Microsoft AI recently formalized (in a blog post) what changes at this stage: "search indexing was built to help humans decide what to read. Grounding indexing is being built to help AI systems decide what to say." (Microsoft AI.) The unit of value shifts from documents to "groundable information — discrete, supportable facts with clear provenance." Read that twice. The model doesn't care about your page; it wants a discrete, sourceable fact that supports the sentence it is about to compose. The structural levers below — entities, lists, definitions, schema, fact patterns — are about producing those units in a form the model can extract without ambiguity.Microsoft AI, "the evolving role of the index" (May 2026): grounding indexing shifts the unit of value from documents to "groundable information — discrete, supportable facts with clear provenance."
AI answer engines don't retrieve pages — they retrieve groundable facts: discrete, supportable claims with clear provenance. Every structural lever in this article exists to package your claims as units a model can lift, name, and trust without ambiguity. If a sentence can't be extracted cleanly with its source attached, it won't be cited, no matter how good the page reads.
There are two other mechanisms that are worth briefly mentioning. First, the Microsoft paper points out that if the underlying data used to train the models is stale, then the model can respond with incorrect information. This is particularly problematic because, unlike a search result that is ranked lower, the model is making a strong claim about the world that is incorrect. If the model claims that a vendor's product supports a protocol it has never actually shipped, even if that isn't true, the cost of that incorrect information is very high. In some cases, if there is no information about a query, the model might simply refuse to answer it. In other words, if your article isn't in the index a given engine draws from, a competitor's article that is will get cited instead.
Engine by Engine, and What Each One Does With Your Article §
If you can only afford three engines this quarter, use Google AI Overviews, ChatGPT Search, and Perplexity. Add Claude and Copilot once the first three are up and running.
ChatGPT Search
ChatGPT Search is powered by a fine-tuned version of the GPT-4o model that was post-trained on a mixture of synthetic data generation and distillation from the o1-preview (OpenAI) model. The search component is powered by a third-party search provider, which in this case is Bing. That means if your site is not indexed by Bing, ChatGPT Search will not surface it through the search leg — separate from OpenAI's direct publisher-licensing partnerships. OpenAI has direct partnerships with some publishers, but they did not disclose how much weight they are giving to their content.
The citation surface is represented by the Sources button below each response, which, when clicked, reveals the underlying sources in an inline sidebar. You can baseline the citations made by ChatGPT by running your own tracked prompts and opening up the Sources button, and noting down which sources it lists. No tool needed.
Perplexity
Perplexity is designed to answer queries by source. That is, given an answer, you can look at each individual claim in the answer and see its source page. If you like, you can do the same for any other answer, and see exactly what changed. In this way, Perplexity is the easiest answer engine to inspect and replicate. If you can find a sentence in one of their answers that's present in another system, you can look at that sentence, figure out why it's easier to extract from one system than another, and make a note to yourself about how you can do better. Perplexity exposes their retrieval system via an open API call to an MCP server, which allows other AI systems to use it as a baseline retrieval system for their own answers (Perplexity Docs).
Claude web search
Anthropic added web search to Claude for U.S. paid customers on 2025-03-20, and made it available to all users globally on 2025-05-27. They don't disclose their retrieval source pool, so Claude's citations are only observable via the interface — you can run your own search with your own citations, and see if they seem to change over time. Claude is particularly relevant for buyers who use Claude in their own tools — in B2B, Claude is often part of the buying committee.
Bing Copilot Search
The first thing to note here is that the Microsoft Copilot Search used in Bing has an "intelligible" UX for citations, linking whole sentences to sources, and listing sources at the top and bottom of the response (Bing Search Blog). This makes being cited by Copilot feel more like being referenced in a normal referral than anything else. ChatGPT has Bing as its third party search API, so the same work done for Copilot citations would also be done for ChatGPT in the case when the user isn't a partner.
Google AI Overviews and AI Mode
On 14 May 2024, Google rolled out AI Overviews for all US users. AIO uses a Gemini model that is tailored to Google Search (Google, AIO FAQ). As noted above, the design of AIO assumes the use of query fan-out. Google states that they have no additional requirements compared to snippet eligibility, except for the need to match the data in the structured data with the data on the page (structured data can help with readability, but it doesn't guarantee a result will be picked as a source, per Google Search Central).
AIO share is volatile. AIO coverage went from 6.49% of Semrush's tracked queries in January 2025 to 24.61% in July and down to 15.69% in November 2025 (Semrush). Any given point in time is misleading for AIO coverage. Semrush also found that the informational share of AIO-triggering queries fell from 91.3% in January 2025 to 57.1% in October 2025 as commercial, transactional, and navigational AIOs took a larger share. In other words, AIO is no longer just at the start of your sales funnel. If your buyers are starting their journey in the middle, AIO is already there.Semrush: AIO coverage of tracked queries ran 6.49% (Jan 2025) → 24.61% (Jul) → 15.69% (Nov 2025). Any single snapshot misleads; measure the trend on your own prompts.
The Five Structural Levers That Change What Gets Cited §
The Princeton GEO paper (KDD 2024) did a test with structural interventions to generative-engine responses. They found that some lever combinations can boost content visibility by up to 40% in generative engine responses (Aggarwal et al., arXiv:2311.09735). This is why I suggest you look at per-topic tracked-prompt sets, not category-level "GEO scores".Aggarwal et al. (KDD 2024, arXiv:2311.09735) found structural GEO interventions lifted visibility in generative-engine responses by up to 40% — but very unevenly across domains, which is why category-level scores mislead.
The five levers below reflect the "discrete, supportable facts with clear provenance" framing Microsoft AI used. Each lever is something you can change within the article itself.
Entities
An entity is a named thing that an engine can resolve to a known concept, such as a company, a product, a person, a regulation, a paper, or a metric with an associated unit. The query fan-out is designed to expand a head query into a set of sub-queries. A page with dense, correctly named entities gets more coverage than a page with the same entities but in the form of pronouns and generic descriptions.
Before (blue-link version): "Several vendors are pitching special optimization strategies for AI search, but the search engine's official documentation doesn't confirm any of these as required."
After (AI-search version): "Vendors are pitching GEO markup, AI-search schema, and llms.txt as required inputs for AI Overviews. Google's own Search Central documentation says there are no additional technical requirements for AIO or AI Mode beyond standard indexing and snippet eligibility (Google, developers.google.com)."
Second version: Four entities (GEO markup, AI-search schema, llms.txt, Search Central) plus one canonical source, one specific claim. First version gives an engine nothing to work with.
Lists
Extractable numerals: lists (ordered or unordered), tables, and so on. If you write the five things you want to say, don't forget to write them as a list and define the terminology so the engine can know what it's talking about. If you want the engine to cite you on a comparison, it better be in a table with rows and columns and names. Buried prose won't work.
Before (blue-link version): "There are a few ways teams typically measure AI-search visibility, ranging from checking whether the engine cited you at all to comparing your citation position to competitors."
After (AI-search version): "Three measurement columns work with a shared tracked-prompt set: (1) cited-at-all (binary), (2) cited-by-name with your brand or article link (binary), (3) cited-above-competitors on the same prompt (ordinal, ranked in citation order)."
Definitions
A definition is a short, declarative sentence about a term, usually identifying what kind of thing the term refers to. Engines add definitions when users ask "what is X?" by returning the shortest, clearest, most self-contained sentence they can find. If you write an article about grounding indexing, starting out with "Grounding indexing is ..." and then go on to explain it in a single sentence, you've provided a definition. If you write an article about grounding indexing that starts with a five-paragraph story, there's nothing to extract, and nothing to work with.
The median AI summary was 67 words long, and in 88 percent of cases, the summaries invoked three or more sources, with only 1 percent citing a single source (Pew Research Center). Sixty-seven words is not a lot of space for a paragraph of setup; it's enough for three definitions, four short claims with three citations each, or two comparisons. Write like you had that much space.Pew Research Center (Jul 2025): the median Google AI summary ran 67 words; 88% cited three or more sources, and only 1% cited a single source.
Schema and structured data
As to the article/FAQ/HowTo/structured data, well, these are still good things to do, but not for the reasons vendors suggest. Google states that structured data should be a representation of the visible data on the page and not an invention to improve machine readability (Search Central). Similarly, Microsoft's AI explains that the units in the ground truth are not the same as the units of the structured data (Grounding Index explainer).
Implement Article and FAQ schema as baseline hygiene. Get it right, keep it aligned with the visible text, and move on. If someone points you to a tool that "adds GEO schema" and charges you $12k/yr, ask them to show you a specific first party documentation page that says that this is required. There is none.
Fact patterns
A fact pattern is a sentence like: "By [some date], [source] says, [link] that X." A fact pattern is a preferred way of writing for these systems, as it fits their "supportable facts with clear provenance" metaphor. One fact-pattern paragraph should include: the source name, the source URL, a date (or as-of) marker, a number (or term) and its unit — and no paraphrasing, rounding, or attribution to "reports show".
Before (blue-link version): "Recent Ahrefs research shows AI Overviews significantly reduce click-through rates on affected keywords."
After (AI-search version): "Ahrefs looked at 300,000 keywords (150,000 with AI overviews, 150,000 informational keywords without), finding that the presence of an AI Overview reduced the average CTR for the first page by ~34.5%. The data is aggregated GSC from March 2024 to March 2025. (Source: Ahrefs, 17 Apr 2025)"Ahrefs (17 Apr 2025): across 300,000 keywords, the presence of an AI Overview cut average first-page CTR by ~34.5% (aggregated GSC, Mar 2024–Mar 2025).
The second version can be copied and pasted verbatim into a grounded response. The first can't — the AI has to summarize a summary, and take a risk on the numbers.
A Measurement Frame You Can Run Without a Tool §
Most GEO pitches skip this part. You do not need to buy anything to baseline your citation position across the five engines. You need a tracked-prompt set, a spreadsheet, and about two hours a month.
A "tracked-prompt set" is 5–10 prompts written in the buyer's language (not your keywords). For example, if your keyword is "vector database benchmarking," the prompt set might look like: "which vector database should I pick for a mid-sized RAG stack with hybrid search?" According to Pew, 60% of question-word searches (who, what, when, why) resulted in an AI summary, and so did 53% of searches with 10+ words. Only 8% of one- or two-word searches got this treatment (Pew). The keyword rank tracker is not the AIs' battleground. The "prompt set" is.
Here is a possible 7-prompt worked set for a technical B2B company selling an infrastructure monitoring product:
- What is the difference between distributed tracing and application performance monitoring in 2026?
- How do I monitor a k8s cluster with more than 500 nodes without being tied to OpenTelemetry?
- Which observability platforms integrate cleanly with a self-hosted Prometheus and Grafana stack?
- When should a mid-sized team switch from open source to a commercial observability tool?
- What are the top failure modes when adding distributed tracing to an existing microservices architecture?
- How much does observability cost at scale (of nodes)?
- What is the best way to detect memory leaks in a long-running Rust service?
Each is 12–25 words. Each is a question. Each includes at least two entities (product category plus a technology or scale). Each is a real thing a buyer would type or say to a chat assistant.
The measurement frame itself is three columns applied against each prompt, per engine:
- Cited at all (binary). For this prompt, in this engine, does the response include any citation to your domain? Yes or no.
- Cited by name (binary). Does the response mention your brand or a specific article of yours in the body of the response — not just in a source list at the end? Yes or no.
- Cited above competitors (ordinal). In the source list or inline citations, rank the sources by position. Where does your source stand in relation to two or three competitors?
Run the set once. Then run it monthly, in the same order, with the same prompts, in the same clients (a signed-out browser or Perplexity's MCP tools for programmatic runs). In 90 days, you will know which prompts the AIOs argue about (not all prompts will elicit AIO action; Google says AIOs only act if they are additive to classic Search results, per Google Search Central), which engines cite you most consistently, and which changes in the design of the AIOs correlate to increases in citations across your prompts.
One final note of caution: currently, it is not possible to differentiate AI Overview clicks and impressions from other Search Console data (Ahrefs). What we're measuring here is a direct observation, something the GSC simply cannot provide.
When AI-Search Optimization Compounds Versus Replaces Organic §
Note: The framing of "AI search visibility = visibility from blue links (i.e., no search visibility)" is incorrect for the vast majority of technical B2B content. The picture is closer to "AI search visibility + organic visibility = visibility from blue links and citations across five AIs."
Does this replace blue-link SEO?
High-intent bottom-of-the-funnel (BOFU) pages are also unlikely to see much benefit from AI-search. BOFU pages tend to be more transactional, targeting pricing pages, specific product-comparison queries and transactional language. According to Ahrefs, 99.2% of pages that trigger an AIO will be informational (Ahrefs), suggesting that blue-link keywords still hold more traction in BOFU pages. Optimising pages targeting keywords that will trigger an AIO won't necessarily improve conversion rates, it will just add another layer of complexity.
For informational and mid-funnel content — the "how does X work," "what's the difference between X and Y," "when should we use X" queries — AIO exposure is now so high that ignoring it risks permanently ceding citation share from blue-link rank. The Ahrefs 34.5 percent CTR reduction on AIO-triggered informational keywords is a loss you need to mitigate if you care about citations.
The story is not symmetrical, and it's not all good news for the AIOs. Semrush did a before/after same-keyword analysis, and found that AIOs have decreased the zero-click rate for keywords getting an AIO from 33.75% to 31.53%. Semrush believe that AIOs focus on low zero-click keywords, so they don't have the same zero-click issues as the average query. Both Ahrefs and Semrush are right — the click rate just gets shifted and compressed. That's why you should measure your visibility based on citation share, not CTR.
Pew's data on "true" behavior supports the point: only one percent of their users clicked on a link inside the AI-generated summary, and the "influence" of AIO is more akin to a brand name than a referral.
What This Does Not Fix §
Structural optimization doesn't create credibility where it doesn't exist. If your product page says one thing, and your about page says another, and your authors are not real people with a reasonable claim on the subject, and your claims can't be found back in sources, AI search optimization just makes this inconsistency stronger.
Three specific things it will not fix:
- Bad authority signals. Thin author pages, unreviewed technical claims, and a domain history of contradictory positioning. AI engines are trained on the same public web that shows these problems; extractable prose does not paper over them.
- Thin product. If the buyer is asking whether your product does the thing or not, citation share on informational articles is downstream of a product-marketing problem, not a content problem.
- Contradictory canonical messaging. If your service page says one thing and your blog says another, the composition step will reveal the contradiction.
The corollary is the useful one: AI-search optimization rewards articles that already have a defensible thesis and verifiable claims. It does not manufacture them.
FAQ §
What does this not fix?
AI-search optimization does not fix bad authority signals, thin product, or contradictory canonical messaging. It rewards articles that already have a defensible thesis and verifiable claims; it does not manufacture them. If the claims aren't trustworthy, structural optimization just makes the weaknesses more extractable.
How long does it take to see citations?
In the Princeton GEO study, the impact of the structural changes on citation share was very uneven across domains (see Aggarwal et al., 2024), so it depends very much on the domain. Also, AIO coverage fluctuates a lot across months (from 6.49% in January 2025, up to 24.61% in July 2025, down to 15.69% in November 2025, see Semrush), so it would be naive to assume a "you'll see it in six weeks" timeline. If you take the AIO baseline, run it once a month, and interpret your own results, you can learn a lot.
Do schema and structured data still matter for AI engines?
For machine readability, yes; for citations, no. Google says they don't need anything else beyond indexing and snippet eligibility, and that their structured data should reflect what's on the page (Google Search Central). They don't mention anything about an "AI-search schema" being necessary; Article, FAQ, and HowTo types kept aligned with your visible content are the good-practice baseline.
If you want to convince a CFO, take a single hard number — the ~34.5% CTR reduction on AIO-triggered informational keywords — and build from there. For a CTO, it's the mechanism: the grounding-index framing from Microsoft AI. Both are better than "96 percent LLM visibility" as a starting point.