Technical guide

The Complete Guide to GEO: Generative Engine Optimization, From the Basics to Real Implementation

The comprehensive guide to GEO: how to get AI engines like ChatGPT, Gemini and Perplexity to cite you, recommend you and weave you into the answer. From the fundamentals to real implementation.

Co-founder and CEO of Fixel, acquired by Logiq in 2020. Head of specialization at Ono Academic College.

The key points

What it is
GEO is the set of actions that make an AI engine cite you, weave you into its answer and recommend you. It is a layer on top of SEO, not a replacement.
What moves the needle
Most 2026 research converges on one conclusion: brand strength and earned media beat technical tricks. 84% of AI citations come from third-party sources.
Where to start
Don't block the AI bots, render server-side, and get indexed in Bing. Then brand, original data, entity clarity and measurement.
Table of contents

By Etgar Shpivak, Head of specialization at Ono Academic College.

Editing: Dafna Harel Kfir, journalist, founder of Poenta (poenta.co.il).

Two or three years ago you searched for a product, got ten blue links, and clicked one. Today you ask ChatGPT and get a single answer with three brands inside it. If your brand isn't one of the three, it doesn't exist. That is the whole story of GEO, and everything else follows from it.

In short, what is GEO? GEO (Generative Engine Optimization) is the set of actions that get an AI engine like ChatGPT, Gemini or Perplexity to cite you, weave you into the answer and recommend you. It is a layer on top of SEO, not a replacement: instead of ranking high on a results page, the goal is to become a source the engine trusts and quotes inside the answer itself.

Who reads what:

PartWho it's forCan you stop?
OneManagers and executivesYes, after Part One
TwoMarketing managers and account leadsYes, after Part Two
ThreePractitioners doing GEOThe technical meat, to the end

What's in this guide, in short:

  • What GEO is, and why it is no longer futuristic but how people already discover things today.
  • The real market share of the AI engines, and why any single number someone quotes you is misleading.
  • How a manager should think about it, what to ask a vendor for, and how to spot bad work.
  • The meat: how to run an audit, how to build a page that gets pulled and cited, and how to become a name the engine recommends.
  • Brand search, social, content repurposing and FAQs, each with the mechanism and why it works.
  • Measurement, tested tools (must-have vs. nice-to-have), and expert level.
  • A 1-10 priority ladder, so you know what to do first.
  • Along the way: small tools that run in your browser, and infographics that explain each point.

The guide is built in three layers, and each one continues the last.

The first part is for executives and CEOs: what GEO is, how it affects them, and what to do at a high level.

The second part is for marketing managers and account leads who work with a GEO specialist: the vocabulary, what to ask for, and how to manage the process.

The third part is the meat, for the practitioner who does GEO themselves, from the audit to expert level. A senior manager can stop after the first part, a mid-level manager after the second, and anyone who wants the technical side reads to the end. I'd still recommend reading in order, because each part builds on the one before.

I've built small tools into the guide that you can use directly. They all run locally in the browser, in plain code with no AI behind them, and no data is pulled or sent anywhere. The information here is yours, and your privacy matters to me.

A note on how I've written this: GEO is a young field, full of hype. It changes fast, and plenty of numbers circulate without backing. I've tried to give only what's supported by sources, to link every study with a month and year, and to present two views where the field is divided. Where there's no solid number, I've said so.

Part One: For Managers. What GEO Is and Why It Already Matters to You

What is GEO?

The three letters GEO stand for Generative Engine Optimization. In plain terms: the actions that make an AI engine like ChatGPT, Gemini, Claude or Perplexity cite you, weave you into an answer, and recommend you when someone asks a question.

This isn't SEO under a new name, and it doesn't replace SEO. It is a layer on top of it. In SEO the question is "how do we rank high on the results page." In GEO the question is "how do we become a source the engine trusts and cites inside the answer."

The difference is subtle but critical: in a world of blue links, second place and seventh place both get clicks. In a world of AI answers, the engine cites a short list of brands, and whoever isn't on the list barely exists. It is closer to a pass-or-fail grade than to a gradual ranking.

Where the term came from

The term was coined in a 2023 paper (GEO: Generative Engine Optimization on arXiv, accepted to KDD 2024), which tested a simulated generative engine the researchers built and found up to a 40% rise in a visibility metric they defined themselves (a page's attributed share of the answer text), not in traffic and not in ranking. That is where the idea came from: citations, statistics and sources help, and keyword stuffing hurts. It is the foundational study, not the only one: controlled experiments have joined it since, chief among them a SIGIR 2026 study on 252,000 runs, and a critical 2026 review warning that a lab gain still doesn't prove real traffic.

SEO vs GEO: The shift from "ten links, pick one" to "one answer, three brands inside it."

What is AEO, and how is it different from GEO and SEO?

AEO, answer engine optimization, is the older name for getting your content picked as the direct answer, starting with Google's featured snippet, the answer box above the results. GEO is the name the research and Wikipedia use for being cited inside AI answers. SEO is the ground both stand on, and in practice people use AEO and GEO interchangeably.

Jason Barnard, who says he coined the term, dates it to 2017, when featured snippets began pushing aside the ten blue links. Wikipedia redirects its "answer engine optimization" entry to the GEO article and notes that, as of early 2026, no agreed definition in the research separated the terms. And Google's guide to its generative AI features (updated July 2026) says that, from where Google stands, both are still SEO.

SEO, AEO and GEO, side by side:

Compared onSEOAEOGEO
The goalRank high on the results pageBe picked as the direct answerBe cited, woven into the answer and recommended
Where the answer appearsThe ten blue linksThe featured snippet above the links, or a voice assistant's replyChatGPT, Gemini, Perplexity, Claude, Copilot, AI Overviews and AI Mode
The unit of workThe pageThe one passage that answers the questionThe paragraph, plus what others say about you (Earned Media)
The main leverRelevance, technical health and linksA short, direct answer right under a question headingBrand strength, mentions across the web, and pages an engine can quote
How success is measuredRankings, clicks and organic trafficSnippets and direct answers wonMentions, citations and Share of Voice on a fixed question set

The AI engines' market share: always ask "which metric"

There is no single number for the AI engines' market share, and that is the first point. ChatGPT's share ranges from 46% to 78%, and it all depends on exactly what you count. Anyone who quotes a single number without stating the metric is misleading you. There are three different metrics, and all of them are legitimate.

First metric: active audience (app and web)

Per Sensor Tower (its True Audience metric, which counts unique users across app and web, data through end of May 2026): ChatGPT holds 46.4%, Gemini 27.7%, and Claude 10.3%. This is the "ChatGPT drops below 50% for the first time" headline. It held above half until January 2026, dropped below 50% in March, and by late May stood at 46.4%.

Second metric: adoption across the population

A Pew study (a representative sample of 5,119 Americans, fielded February 2026 and published June 2026, the most reliable source here) found that about half of Americans use an AI chatbot (49% today versus a third in 2024).

ChatGPT is the leading chatbot (44% have ever used it, Gemini next at 24% and Copilot at 17%).

42% use a chatbot to look up information, and 60% of Americans read the AI summaries that appear at the top of search results.

Third metric: referral traffic to websites

This is the most relevant metric if you care about traffic to your site, and it tilts far more heavily toward ChatGPT. Per StatCounter (March 2026): ChatGPT sends 78% of the referral traffic coming from chatbots, Gemini sends 8.65%, Perplexity sends 7% and Claude sends almost 3%. A study by SE Ranking across 101,574 sites gives similar numbers: ChatGPT sends about 75% and Gemini about 11.5%.

The metric everyone forgets: the AI inside Google itself

When people talk about "traffic from AI" it is easy to forget that the biggest venue is not a separate chatbot, but Google's AI Overviews, the AI summaries that appear at the top of ordinary search results. Per BrightEdge they appeared on about 48% of their tracked queries in early 2026, and on broader panels the rate is far lower (Semrush: about 16% of US searches, and up to about 39% of informational queries).

Either way, 60% of Americans read them, and they reach more people than all the separate chatbots combined. The intent is different: someone reading a summary in Google usually isn't "conversing," they're searching. But in terms of visibility, this is where the most people encounter an AI answer and where the most traffic stops. So I wouldn't measure GEO by ChatGPT alone.

Same brand, three numbers: ChatGPT's market share is 46% or 78%, depending on the metric.

The bottom line: ChatGPT leads on every metric and keeps growing in absolute numbers (Pew: from 34% to 44% in a year), but its share is falling on every metric because competitors are growing faster. Gemini is the clear number two and the biggest gainer in percentage points. Claude grows fastest in relative terms but stays small. And most important: these are different metrics, so don't compare an "app usage" number to a "referral traffic" number. They measure different things.

Why does this matter to you?

Three things are happening at once, and each of them can affect you directly.

Where did the click go?

Per a SparkToro analysis (January to April 2026), about 68% of US Google searches end with no click to any site, up from 60% in 2024 (other sources measure lower, in the 58% to 65% range, so treat the number as a range). And when an AI summary appears, the effect on clicks is sharp: an Ahrefs study (February 2026, December 2025 data, 300,000 keywords) found that the click-through rate on the first result was tied to a drop of about 58% relative to expected when an AI Overview is present, nearly 1.7 times the 34.5% drop measured in April 2025. In other words, more and more people get the answer without reaching you at all.

Why is a mention worth more than a click?

If nobody clicks, what still matters is whether you were mentioned in the answer itself. Being cited has become more important than winning the click. This is a conceptual shift that takes managers time to absorb: the metric is no longer "how much traffic did we get," but "does the engine recommend us when someone asks."

Why a strong brand beats the tricks

And here is where whoever built a strong brand has the advantage. Most of the serious 2026 research converges on the same conclusion: what really moves the needle is brand strength and presence off your own site, not technical tricks. Ahrefs looked at 75,000 brands (May and December 2025) and found that brand mentions across the web, and especially on YouTube, correlate far more strongly with AI visibility than backlinks do (correlation 0.66-0.74 versus about 0.2). This is correlation, not proof of causation, but the direction is consistent.

Translated for managers: your investment in brand, in PR and in building awareness is exactly what makes the AI recommend you. GEO is, to a large degree, demand generation.

The funnel flipped: We used to start from traffic. Now we start from being known.

The four question types every manager should know

Not all questions are the same. A user asks different things depending on their situation, and the GEO strategy shifts accordingly. Four types, in my breakdown (it parallels the TOFU/MOFU/BOFU funnel that appears at GoVISIBLE and elsewhere).

General questions about the field

For example: "what is a CRM," "how does customer onboarding work."

Here the user isn't thinking about a brand yet. The engine looks for whoever explains the category with authority.

The win here is to be the source that explains the field, not to push your name.

For B2B companies this is usually the biggest long-term investment, but also the most expensive and the slowest, so it comes after you've closed out the brand and comparison questions. The customer asks these questions before they even know a vendor. In many cases there are no "vendors" involved at all, just a general explanation of the category.

Research and consideration questions

For example: "which CRM is best for a small real-estate business," "the right tool for a five-person team."

Here the user is already looking for a fit to their specific situation, but is still open. The engine builds a short list from criteria (size, field, budget), and usually draws it from third-party sources that map the options.

The win here is to appear in the comparisons and guides that map the market by need, and to hold your own content that answers exactly this dilemma. This is the stage where you enter the "consideration set," and the choice flows from there.

Brand questions

For example: "what does Etgar Shpivak do," "is this company trustworthy."

Here the engine builds the answer mainly from third-party sources, not from your site. Two critical questions: does the engine even know you, and does it describe you correctly. A wrong description of the brand in answers isn't a content problem, it's an entity and reputation management problem.

This is also where to start: the first part of GEO work should focus precisely on the brand questions, because these are people who already want to check you out, and you need to make sure they get the answer you want.

Comparison questions

For example: "X vs Y," "the best tool for a small SaaS," "alternatives to Z." These are the buying questions, the most competitive and the most valuable. Here the engine leans heavily on third-party lists and rankings.

Four question types, four strategies: Every stage of the journey has a different question type and strategy.

So what should managers do?

You don't need to become a GEO expert. You need to know where to point resources and what to expect. The order of priority in the process: start with the questions people already ask about you by name (brand questions), then the category comparison questions, and only at the end the general questions about the field.

Why invest in brand and entity, not tricks?

An "entity" is how the engine identifies you as one clear thing in the world: the name, the field, the people behind you, and the link between all of these and a single brand. If there is one thing that recurs in every 2026 study, it is that Earned Media (mentions and coverage you earned through your content and reputation, not advertising you paid for) and brand strength beat owned content and links. Muck Rack found (May 2026, 25 million links) that 84% of AI citations come from earned media. A budget that goes to real PR, to original research that gets cited, and to being present on sites the engine trusts works better than a budget chasing technical tricks.

Why the process is slow, and how to measure it right

GEO doesn't deliver results in a week. Citations usually start moving between week 6 and week 12, so a 30-day trial is meaningless. Siege Media estimates that AI's real impact is 100% to 200% higher than analytics shows. Other checks show it can be larger by an order of magnitude, more than 10x what you see in your site analytics.

The managerial takeaway: don't wait for the perfect number, start measuring now and watch three signals together: AI traffic, share of voice in answers, and a rise in searches for your brand name on Google.

What can you promise leadership, and what can't you?

If you report to a CEO or a board, don't promise a specific position, don't promise that a one-time optimization will hold over time, and don't promise precise attribution. No platform gives you that. What you can promise: systematic measurement and a trend that improves over time.

That's the foundation. Anyone who just wants the big picture can stop here. Anyone who needs to manage the process in practice, the next part is for you.

Part Two: For Account Leads and Marketing Managers. How to Work Effectively with a GEO Specialist

This part is for whoever doesn't do the GEO themselves but is responsible for it. You're the account managers, the account leads, the marketing managers who have to brief a vendor, approve a budget, and judge whether the work is good. To manage a process you need to understand the vocabulary, know what to ask for, and know how to spot bad work.

The vocabulary: terms you must know

Without understanding the words, you can't hold a professional conversation with a vendor. Here's the minimum, in plain terms:

  • Citation vs mention. A mention is when the engine names you in the body of the answer. A citation is when it links to your site as a source. Two different things you need to measure separately.
  • Share of Voice. Out of all the answers in your category, in how many are you mentioned versus your competitors. This is one of the most important metrics.
  • Entity. How the engine understands who you are: the name, the field, the founders, the competitors. Two examples to make it concrete: when you write "Apple," the engine needs to know whether you mean the technology company or the fruit, and that's decided by the surrounding context and links. And when it reads "Etgar Shpivak," the entity ties the name, the role at Ono, and the marketing work into a single identified person. A brand with a clear, consistent entity is easier to identify.
  • Crawlability. Whether the AI engines' bots can even reach your content and read it.
  • fan-out. Google, ChatGPT and other AI engines break one question into many sub-questions and look for an answer to each. We'll expand on this in the third part, but it's important that you know the word.
  • E-E-A-T. Experience, Expertise, Authoritativeness, Trust. Google's framework for assessing content trustworthiness. Trust is the most important of the four.
A GEO glossary: The six terms you must know to hold a conversation with a vendor, each in one sentence.

If a vendor throws these words at you without explaining them, that itself is a flag. A good GEO specialist knows how to translate the craft into language a manager understands.

The check you run before you start

You must not start a GEO process without capturing the starting point. Without a baseline, you can't know whether anything improved. A good initial audit checks four dimensions (Discovered Labs, July 2026):

  • Technical health (do the bots arrive, is the data structured)
  • Extractability (is the content built so a clean paragraph can be pulled from it)
  • Entity authority (does the engine identify you)
  • Off-site consistency (is what's written about you across the web accurate and consistent).

The baseline translates into three numbers you measure before you start, and all three are genuinely different from one another.

Mention rate is the percentage of answers in which the engine named you at all.

Citation rate is the percentage in which it also linked to your site as a source. It is usually lower than the mention rate, because you can mention a brand without linking to it (and occasionally the reverse, a link without naming the brand).

Share of voice is how many of the mentions in the category belong to you versus your competitors, that is, who leads the conversation. Example: out of 50 key questions you were mentioned in 15 and cited in only 5, so your mention rate is 30% and your citation rate 10%. If in those same questions the competitor was mentioned 30 times and you 15, your share of voice is a third.

Important: these must be questions genuinely relevant to your business, not generic questions that sound good in a report.

A point an experienced GEO specialist will raise: measurement is a sample, not something exact. AI answers aren't deterministic, so you measure a fixed set of questions over and over rather than trying to catch every possible answer (LLM Pulse, May 2026).

The four audit dimensions: A GEO audit checks four things, not just the site.

GEO Readiness Scorecard

Runs locally

What it does: a roughly 20-item self-check across six categories (crawler access, structured data, extractability, entity authority, measurement, and brand & off-site) that produces coverage by dimension (green/yellow/red) and the three most urgent gaps. Input: yes/partial/no answers. Output: coverage by category and an action list, exportable to Excel or CSV. Data is stored locally in the browser so you can see progress next time.

What to ask the vendor for, and what a good report looks like

Three deliverables you should get from a GEO vendor: an audit report (the starting point), a recommendations report (what to do, in priority order), and an ongoing monthly report. Here is what each looks like when it's done right.

What does the audit report look like?

The audit report, as I build it, has six to nine sections: an executive summary (if the vendor gives an overall score, demand to see its components), visibility and sentiment for each engine separately, citation and source analysis (which domains the engines cite in your category), competitive share of voice, content structure and citation readiness, structured-data gaps, entity and trust, technical readiness, and finally findings and a prioritized fix plan with an effort estimate.

What a GEO audit report looks like: The minimum structure to demand from a vendor in the audit report.

What does the recommendations report look like?

The recommendations report is organized around four pillars (the split I work with, parallel to Discovered Labs' four audit dimensions): technical (can the bots find and access the content), content (does the AI understand and extract it), entity (does the AI identify who you are), and brand authority (does the AI trust you and prefer you).

The first two are on-site work, the last two mostly off-site. Every recommendation gets an impact-vs-effort score, and is split into "quick wins" versus a "roadmap."

What does the monthly report look like?

The monthly report is where you see whether you're paying for real work. A clean structure, in my wording, answers three questions (the metric list per LLM Pulse, May 2026): how are we doing (share of voice per engine, average position, accuracy rate, versus last month and versus the baseline), how do we compare to competitors (share-of-voice gap against two or three competitors) and what's happening (new citing sources, position gains, hallucinations detected). And most important, citation tracking: which specific pages of yours the AI links to as a source. This is the deliverable that proves the work happened.

A useful discipline: ask the vendor to separate the measurement into three tracks: cited (with a link), mentioned (by name, without a link) and message accuracy (what the AI actually claims about you). This distinction prevents reports that blur "they mentioned us" with "they linked to us."

What to check on an ongoing basis and how to set targets

Here two vendors recommend different cadences. GenOptima recommends tracking the leading metrics (citation, mention, position) weekly. Analyticahouse builds a monthly reporting model. The statistical logic is simple: AI answers are volatile, so you sample often but decide slowly, to avoid reacting to noise.

The compromise I recommend: look at the leading metrics weekly or biweekly to know what's happening, but make decisions and report once a month, with a quarterly strategic reset. One thing is certain: if a vendor reports to you less than once a month, that's a red flag.

On targets, you need to separate leading metrics from outcome metrics. Leading metrics move fast and show whether you're on track, even before there's any money: mention rate, citation rate, and your position inside the answer. Outcome (lagging) metrics are the business result itself, and they move slowly. Three concrete examples: how much traffic reached the site from AI engines, how many of those visitors became paying customers, and by how much the number of people searching your brand name on Google rose.

The relationship between the two: after two months you already see you were mentioned in more answers (a leading metric rises), but only after six months does it start translating into more sales inquiries (an outcome metric). Realistic expectations (a common vendor estimate, not a measured figure): first citations in the early months, around three months a meaningful rise in citation rate, and only after six months and up a stable correlation between AI citations and a rise in brand searches.

A complementary lens for tracking, not by funnel stage but by the type of prompt the customer types: a recommendation prompt (asks for a list of options without naming a brand, "which tools are worth checking out"), a comparison prompt (between known options, or a search for an alternative) and a verification prompt (checking a detail before acting: price, hours, a technical limit, and this is where the AI is most exposed to error because its knowledge is stale). A rule for measurement: each prompt is measured at least ten times from clean accounts with no memory layer, because a single answer is an anecdote, not data.

Leading vs lagging metrics: Measure traffic fast, decide slowly, expect the business result slower still.

AI Share of Voice Calculator

Runs locally

What it does: enter how many times you and your competitors were mentioned across a set of questions, and the tool computes share of voice as a percentage, separates mentions from citations, and draws a chart. Input: mention/citation counts for you and competitors. Output: share of voice, mention vs. citation rate, chart. Export to Excel or CSV. Everything is stored locally.

GEO KPI Dashboard Template

Runs locally

What it does: a simple board to track, month over month, AI traffic, share of voice, the number of questions that cite you, and the branded-vs-unbranded ratio. Input: manual monthly data (copied from GA4 and the earlier tools). Output: KPI cards with the change from the previous month, export to Excel or CSV. Local storage.

Red flags: how to spot a bad vendor

This is the part that will save you money. Here are the signs that a vendor doesn't really know what they're doing (inspired by a post from Influencers-Time, September 2026, a marketing content site rather than research, and by my own experience):

  • Vague engine coverage. If they say "all the leading platforms" without naming names, that's suspect. A good vendor names the platforms relevant to you: ChatGPT, Perplexity, Gemini, AI Overviews, Copilot, Claude.
  • Reporting less than once a month.
  • Ownership of the deliverables. What happens to the content the vendor created off-site, and does the citation history stay yours if you leave?
  • Percentage case studies with no baseline. "A 50% rise" that is really from 2 to 3 citations is not the same story as 50% of 2,000.
  • A composite "visibility" score you can't break apart.

The most reliable sign, in my view: a weak vendor can't show you the exact questions they track, and can't name a single specific new citation they won in the last 30 days. Real work leaves traces you can see.

The five red flags: Five signs you're dealing with a weak vendor.

A full example: GEO on one brand, from audit to result

That's the components. Here is how they connect on a single brand. Take a fictional example: "Novopay," a payments tool for small businesses.

The starting point (audit). You run 50 key questions across five engines. Novopay is mentioned in 20% of the answers, cited (with a link) in only 6%, and its share of voice against two competitors is about a third across the whole set, and lower still on the comparison questions, where the real fight is. OpenAI's bot gets a 200 on the site, but the pricing page is rendered client-side only, so it is invisible to retrieval.

The three moves chosen (in build order). First, fix the rendering so the core content arrives in the initial HTML, and make sure the site is indexed in Bing. Second, write three comparison pages that answer the real buying questions ("Novopay vs X"), each with a direct answer up front and original data. Third, land three earned-media placements and publish a webinar transcript on your own domain.

What moved (weeks 6 to 12, illustrative numbers). The citation rate rose from 6% to 14%, share of voice on the comparison questions crossed half against competitors, and brand searches for "Novopay" on Google started to climb. None of this happened in the first week, and all of it was measured against the baseline captured in the audit.

Part Three: For the Practitioner. The Meat of GEO

From here it gets more technical, but I've written it assuming you're not a GEO expert, just a marketer who wants to become one. Every term is explained. Some passages are basic and some are genuinely technical, and that's fine. The goal is for you to finish this part able to do GEO yourself.

How to run a GEO audit

A real audit isn't a single screenshot of ChatGPT. It's a repeatable, reproducible assessment of how the brand appears in answers, across several engines. The reason: AI answers aren't deterministic, and the citation mix varies wildly. An audit of a single run simply isn't reliable.

The first step: a fixed question set

You build a fixed set of 30 to 100 questions, segmented by the types we covered: general questions, research questions, brand questions, comparison questions. You run them across ChatGPT, Perplexity, Google's AI Overviews, Gemini and Copilot, each one several times. For every answer you record: were we mentioned, were we cited, which competitors appeared, and which sources the engine drew from.

Where do the questions come from? You don't invent them out of thin air, you gather them from the places where real customers already ask.

Six sources:

  • Questions that recur in sales calls and support tickets
  • The "People Also Ask" box at the bottom of Google results for your keywords
  • Tools like AlsoAsked and AnswerThePublic that map the questions asked around a topic
  • Real discussions on Reddit and Quora in your field
  • The comparison questions against competitors ("X vs Y", "alternative to Z")
  • And finally, simply asking ChatGPT and Perplexity "what do people ask when they're considering a purchase in this category," and seeing which questions they suggest themselves.

From all of these you pick the questions genuinely relevant to the business, and lock them as a list that runs over and over.

A clean five-step methodology you can adopt: an AI-presence audit, entity and authority alignment, building citable "answer units" with original data, an off-site layer, and weekly citation tracking on a fixed question set.

The crawlers: who even reaches your content

A technical point that separates a good audit: the crawlers. There are three bot families, not two: a training bot that collects content to train the model, a search-index bot that builds the engine's search index, and a user-triggered fetcher that brings a page in real time when a user asks.

A critical point: robots.txt is a request, not a wall. It controls training and indexing, but user-triggered fetchers (ChatGPT-User, Perplexity-User) usually don't obey it. You can also accidentally block the index bot, and then you simply don't exist in answers. Here's who's who:

CompanyTraining botSearch-index botFetcher (user-triggered)
OpenAIGPTBotOAI-SearchBotChatGPT-User
AnthropicClaudeBotClaude-SearchBotClaude-User
Perplexity(none separate)PerplexityBotPerplexity-User
GoogleGoogle-Extended (training + grounding in Gemini)Googlebot (search + AI Overviews)(none)
Microsoft(none separate)Bingbot(none)
Meta / Apple / Common CrawlMeta-ExternalAgent, Applebot-Extended, CCBot(none)(none)

The most important nuance with Google: AI Overviews and AI Mode are served from the regular search index and use Googlebot, not a dedicated AI bot. Blocking Google-Extended doesn't remove you from Overviews, but it does remove you from both training and the grounding of the Gemini app, the second-largest engine. Anyone who wants to appear in Gemini doesn't block it. Blocking Googlebot, by contrast, removes you from everything.

A point from the OpenAI, Anthropic and Perplexity bot documentation that saves an expensive mistake: each bot obeys only the most specific group of rules that matches it in robots.txt. If you blocked everything with User-agent: * and explicitly opened only for Googlebot, the AI bots (which have no dedicated rule of their own) fall under the general blocking rule, and you've blocked them without noticing.

AI-Crawler Access Checker and robots.txt Generator

Runs locally

What it does: paste your robots.txt file (or start from scratch), and the tool shows a table of which AI bots are allowed and which are blocked, and generates a valid file for the stance you chose (allow answers, block training, block everything). Input: robots.txt content + a stance. Output: a bot table, a ready-to-download file, and plain-language warnings (including the reminder that user-triggered fetchers usually don't obey robots.txt, so "block everything" doesn't really block everything). It all runs in the browser in plain code, with no AI behind it, because the tool doesn't access your site, you paste the file.

GEO Audit Checklist

Runs locally

What it does: an interactive checklist covering all the audit steps (defining a question set, baseline runs across every engine, scoring mention/citation/share of voice/sentiment, source analysis, crawlability checks, on-page and structured data, off-site entity, competitive gap, and prioritized findings), marks what's done, and produces a report. Input: checked items. Output: percent complete, a gap list, and export to Excel or CSV. Local storage.

The differences between the main engines

Every engine retrieves and cites differently, so you can't run one strategy for all of them, and you can't average data across engines. Three points that change your planning:

  • ChatGPT started on Bing's index, but that's changing. A small Seer Interactive study (February 2025, 100 SearchGPT queries) found that 87% of citations matched Bing's top results, versus about 56% for Google. But in 2026 ChatGPT built its own index (Peec AI, September 2026), via the OAI-SearchBot crawler and alongside external results providers, and Bing remains mainly in Deep Research. The practical conclusion: don't block OAI-SearchBot, and make sure you're indexed in Bing as a safety net. For Copilot, Bing is still the engine. It's something most marketers don't touch at all.
  • Google works with fan-out. The query is broken into parallel sub-questions, each retrieves and ranks sources separately, and Gemini synthesizes. So a page ranked sixth that answers a sub-question excellently can get cited. Here topical coverage and paragraphs beat ranking alone.
  • Perplexity is the densest in citations and always retrieves live. Every claim maps to a numbered footnote, and source freshness and authority win there.

The upshot: run a per-engine playbook (ChatGPT: domain authority and freshness; Perplexity: dense data and original research; Gemini: lean on Google's ecosystem and your existing SEO; Copilot: the Bing index; and AI Mode: cover the entire fan-out tree). Don't assume that what works on one works on all.

How do you show up in ChatGPT's answers?

ChatGPT's search answers draw on pages its OAI-SearchBot crawler has read. OpenAI's crawler documentation says a site that opts out of that bot will not be shown in ChatGPT search answers, only as a navigational link, and that a robots.txt change takes about 24 hours to register.

What to do: open yoursite.com/robots.txt in the browser (or paste it into the checker above) and confirm that nothing blocks OAI-SearchBot. Then ask ChatGPT, with search on, three questions about your brand and note which of your pages it cites.

How do you appear in Google AI Overviews and AI Mode?

They are two different places on one index. AI Overviews is the summary Google shows above the ordinary results, and only when it judges one adds something. AI Mode is the separate mode for longer questions and comparisons, and the two may run on different models, so the links they show differ. Google's documentation sets one condition for both: a page must be indexed and eligible to be shown in Google Search with a snippet, and there is no extra file or markup to add.

What to do: put your most important page into URL Inspection in Search Console and check that it is indexed. If it is indexed and still never appears, ask your developer whether a nosnippet or max-snippet tag is limiting it.

How do you get cited by Perplexity?

Every Perplexity question triggers a live web search, and every claim in the answer gets a footnote, so it can only cite a page its PerplexityBot crawler is allowed to reach. Perplexity's crawler documentation warns that a site behind a web application firewall (a WAF, the security layer many sites sit behind, such as Cloudflare's) may need an explicit rule to let its bots in.

What to do: ask whoever runs your Cloudflare or hosting to confirm that PerplexityBot is allowed. Then ask Perplexity one question your page answers and check whether you are in the footnotes.

How do you appear in Claude's answers?

Anthropic runs three bots, and the one that indexes your pages for Claude's search is Claude-SearchBot. Anthropic's help article says that disabling it may reduce your site's visibility and accuracy in user search results, and that Anthropic's bots will not try to get past a CAPTCHA, so a bot challenge on your site keeps Claude out as surely as a robots.txt block.

What to do: allow Claude-SearchBot and Claude-User in robots.txt, and check that your bot protection does not challenge them.

How do you get cited in Microsoft Copilot?

Copilot answers from Bing's index, so the work starts in Bing Webmaster Tools, which is free. Its AI Performance report shows how your pages appear in AI answers across Copilot and Bing, and in June 2026 Microsoft added intent, topic and citation-share views to it, in preview.

What to do: verify the site in Bing Webmaster Tools, submit your sitemap, and check the report once a month for pages that stopped being cited.

On-site: how to build a page that gets pulled and cited

A lot of what you did for SEO is still very relevant to GEO. Clear content, a clean structure, a fast site accessible to crawlers, all of these stay. GEO doesn't replace them, it arrives as a layer on top. So don't throw out what worked, just add to it. Here's what's genuinely different.

Why the retrieved unit is the paragraph, not the page

An engine retrieving in real time doesn't "read your page" start to finish. It searches (in Bing, in Google, or in its own index), fetches pages, and uses only the paragraphs that answer the sub-question. Many systems work on content chunks, not on a whole document. The conclusion: the unit that gets retrieved and cited is the paragraph, not the page. Judge every decision on the page by the question, "how does this paragraph read on its own, detached from everything around it" (a handy rule of thumb: a short paragraph that stands alone). Proper optimization is for clarity at the level of the standalone paragraph, not for ranking at the document level.

How do you write a paragraph that gets retrieved?

Under every heading, open with a sentence that fully answers the question, in one short paragraph (the field's rule of thumb, about 40-60 words, not a formal threshold), before any preamble. Standalone sentences, with no pronouns pointing elsewhere ("this," "as we saw above"). Headings in question form. Facts and numbers with source attribution. The winning format by question type: a comparison in a table, a process in a numbered list, terms in a definition list.

Format is a gateway to retrieval that precedes keywords. And indeed, most AI Overviews contain at least one list, and statistic lines and tables get pulled the most, because you can't safely rephrase a number.

"Answer-First" and Citability Analyzer

Runs locally

What it does: paste a paragraph, and the tool scores it: is the answer in the first sentence? Is it 40-60 words long (a rule of thumb, not a measured threshold)? Standalone? Does it contain a number or a data point? And it suggests improvements. Input: a paragraph or passage. Output: a score, a checklist, and phrasing suggestions. It all runs in the browser in plain code, with no AI behind it.

Rendering: how sites disappear quietly

Vercel and MERJ analyzed (the foundational study in this area, December 2024, and still valid) 569 million GPTBot requests and found zero JavaScript execution. The bots of OpenAI, Anthropic and Perplexity fetch raw HTML only. Worth knowing: Gemini and Google's AI Overviews do render JavaScript through Googlebot's infrastructure. The implication: if your content is rendered only client-side (CSR, Client-Side Rendering, meaning the JavaScript builds the page in the visitor's browser), it is invisible to ChatGPT, Claude and Perplexity. Serving the core HTML from the server (SSR or static HTML) fixes this.

A controlled test built identical pages with fictional data on AI-based website builders (Lovable, Base44, Bolt, v0) and checked whether ChatGPT and Claude could read them. The result: body content surfaced only when the platform served real HTML. Platforms that returned an empty div simply failed.

How do you check this yourself, without writing a line of code? Simplest: right-click the page and choose "View Source" (or Ctrl+U). That shows exactly the raw HTML the bot receives, before the JavaScript runs. Don't use "Inspect," because it shows the page after the JavaScript built it, which is not what the bot sees. Note that a bot sometimes gets a different response than your browser (for example a Cloudflare 403 block), so if you have access to server logs, that's where you see what the bot really got. Google Search Console's URL inspection tool shows the HTML after Google's rendering, so it fits Google but doesn't necessarily reflect the other bots. The rule: if the core content doesn't appear in "View Source," the bots probably don't see it.

How do you build a machine-readable structure?

Beyond rendering, the structure itself needs to be machine-readable. A clean heading hierarchy (one H1, H2 for topics, H3 for subtopics), real semantic tags (article, section, nav) and not just div, and each idea in its own paragraph. The clearer the structure, the easier it is for the engine to chunk it correctly and pull the right paragraph.

How do you build a GEO page, step by step?

  1. Start from the question the page answers, not from a keyword.
  2. Write a heading in question form.
  3. Open with a direct 40-60 word answer.
  4. Add attributed data or a statistic on a standalone line.
  5. Break it into subtopics with subheadings.
  6. Choose a format by question type: a table for a comparison, a list for a process.
  7. Add an FAQ block from the real questions that came up.
  8. Mark up schema that matches what's visible on the page.

Site speed and user experience: plenty of myths

Speed is an entry ticket, not a lever. This needs precision, because it has two halves that sound contradictory and are both true. On one hand, a Discovered Labs study (August 2026, about 2 million citations) found that once you control for domain strength, Core Web Vitals (Google's speed and stability metrics) has no independent effect on citations, so better speed doesn't buy you citations. On the other hand, a site that's too slow does hurt: AI crawlers don't wait for a slow page and may abandon it before reading it. There's no official number for this, so the practical rule is simple: aim for a server response time (TTFB) under 0.8 seconds (Google's threshold), preferably less, and make sure the main content already arrives in the initial HTML.

After that, stop and invest the rest in authority and content, not in more speed improvements. And user experience? The engine doesn't see your analytics, so dwell time and bounce rate don't feed a citation. UX matters for converting the traffic you already won, not for winning the citation.

Schema and entity: where the lever really is

Schema (structured data) doesn't create authority, it removes ambiguity so the machine knows exactly who you are. The value is in consistency: the same name and the same details everywhere, and a sameAs property that links your entity on the site to the identifying profiles across the web: Wikipedia, Wikidata, LinkedIn, Crunchbase. The types that matter: Organization, Person (with the author and their credentials), Article, FAQPage.

An iron rule: the schema must match what's visibly on the page, otherwise it's a guidelines violation.

Where it actually sits: a JSON-LD block inside the page's <head> (or right before the end of the <body>), not something the visitor sees on screen.

And how to check it yourself, without touching code: paste the page URL into Google's Rich Results Test, and the tool shows which schema types Google detected on the page and whether they have errors. And be honest with yourself: valid schema is a baseline and hygiene, not a GEO win. The real lever is entity clarity, and most of it is free.

Organization and Person Schema Generator

Runs locally

What it does: fill in a form (name, site, logo, role, a sameAs list), and the tool produces a valid JSON-LD block (the format in which structured data is written) to copy, with an alert for missing fields. Input: the organization/person details + a profile list. Output: ready-made schema code. It all runs in the browser in plain code, with no AI behind it.

"Key Facts" and FAQ Block Generator

Runs locally

What it does: enter question-answer pairs, and the tool produces both clean HTML for the page and FAQPage schema, with an alert if an answer is too long or has no number. Input: questions and answers. Output: HTML + JSON-LD. It all runs in the browser in plain code, with no AI behind it.

Knowledge Panel and Wikidata, without confusing the terms

Knowledge Graph is Google's internal database of entities in the world; the user doesn't see it. Knowledge Panel is what you do see: the info box on the side of the results page. Wikidata is a structured, open database that feeds the knowledge graph and is absorbed by the LLMs (the large language models behind engines like ChatGPT).

How it connects: consistent sources like Wikidata feed Google's knowledge graph, and when Google is confident enough about who you are, it shows a Knowledge Panel for you.

A surprising point: a Wikipedia entry is no longer required to get a panel, a valid Wikidata item can support one. But it doesn't produce a panel on its own, and it requires notability: an item a brand creates about itself with no serious external sources (press, a company registry) gets deleted.

You may edit your own entity as long as every claim rests on an external source, and paid editing must be disclosed. A local business gets a panel first through a Google Business Profile.

From Wikidata to Knowledge Panel: How consistent information becomes the info card you see in Google.

Entity Consistency Checker (NAP: Name, Address, Phone)

Runs locally

What it does: enter your brand details as they appear in each source (name in every language and variant, description, category, founding year, founders, city, profiles), and the tool flags the inconsistencies that confuse the engine. Name, address and phone (NAP) matter mainly for local businesses. Input: the same fields from each source. Output: highlighted mismatches + a suggested canonical version. It all runs in the browser in plain code, with no AI behind it.

Off-site: becoming a name the engine recommends

A large part of GEO work happens off your site. The engine asks "who is trustworthy and what's the consensus about them," and it solves that through what the web says about you.

What do the 2026 numbers say?

Muck Rack (May 2026, 25 million links): 84% of AI citations come from Earned Media (third-party sources you earned through your content and reputation, mainly press but also academia, government and Wikipedia, not advertising you paid for), and only 0.3% from paid advertising.

Ahrefs (about 75,000 brands, May and December 2025): brand mentions across the web correlate far more strongly with AI visibility than backlinks (correlation, not causation), and Stacker with Scrunch (March 2026, a study by the distribution vendor itself on 30 brands): distribution through third-party sites produced a median 239% rise in citations, versus the same content on your own site alone.

The rule: the domain that says it about you matters more than the words.

There's also an honest piece of evidence that calls for caution. Experiments by Zeeshan Yaseen in Search Engine Land (September 2026) found that "best of" lists a brand published about itself contributed only 14% of the citations, while third-party names delivered nearly 86% of the citations. And also: about half the sources stopped being cited within 30 days. A one-time publication doesn't hold. The conclusion: being included in others' "best of" lists beats publishing such a list about yourself.

The red line: why you must not fake presence

The temptation to fake presence, plant reviews or assemble "best of" lists on sites you hired is strong, because lists are a large share of citations (Evertune research in Search Engine Land, May 2026: 63% of citations in the sample pointed to lists). But fakery gets caught. GPTZero published an investigation that found most of the sources in EY Canada's report were fabricated or broken, and the report was pulled. How do you tell real from fake? Real presence comes from an external party with no stake in the outcome: a site with a name, a real author, a visible ranking method, and links that lead to sources that exist. Fake is identified by the same phrasing recurring, a wave of publications in a narrow time window, accounts with no history, and links that return 404. And beyond that, faking reviews is a legal risk: the FTC warned ten companies (December 2025) and noted possible penalties of up to $53,088 per violation. The simple rule: if the presence wouldn't survive a check by a skeptical person with an open browser, you're better off investing in a real result someone will mention of their own accord.

Which sources do the engines cite, and how did it change in 2026?

Until early 2026, Reddit was the most-cited domain (Peec AI, March 2026). Already by late 2025 YouTube overtook it in the share of citations from social networks (from 18.9% to 39.2% between August and December 2025, per Goodie AI), and in August 2026 Reddit's share of ChatGPT citations collapsed from 3.8% to 0.5% after a change in how ChatGPT searches (the reason hasn't been confirmed). One live tracker (LLM Pulse, September 2026) puts YouTube first and Reddit around third. The important point: every engine has a completely different source mix, and only about 11% of domains are cited in both ChatGPT and Perplexity. Per Profound (2024-2025 data): ChatGPT leaned then on Wikipedia and press (Wikipedia's share has since dropped below 20%); Perplexity leans heavily on Reddit (about 47% of its top ten sources, not of all citations), and it's the only engine where Reddit remains king; Google AI Overviews balances between Reddit and YouTube; and Copilot is a complete outlier (Ahrefs, September 2026): it leans on shopping sites, with Amazon and Walmart leading, and no Reddit. The takeaway: you can't run one source strategy for all engines, and this ranking also shifts fast, so verify it against current data before you build on it.

Where citations come from: 84% of citations come from earned media, not from your own site.
Top cited sources, by engine: Every engine has a completely different source mix.

Reddit, reviews and YouTube: what's really worth it?

Reddit is still among the most-cited domains, especially in Perplexity and AI Overviews, so a real presence there (not spam) is worth it, even though it has almost vanished from ChatGPT since August 2026. For "best of" questions, review sites like G2 for B2B and Trustpilot for B2C are sources the engine lifts from. And YouTube: most engines don't watch the video, they read the transcript. So a long video with a published transcript is a real GEO asset, which is why YouTube is climbing fast among cited sources.

Why brand strength is the backbone of everything

Everything we've written here converges on one point: the AI recommends whoever the web already knows. So the most impactful investment in GEO is usually not technical but marketing, building a brand people talk about. And that leads straight to the next part.

Brand search and social media

Brand strength is one of the biggest levers for GEO, and brand search is the metric that reveals it. Brand search and AI visibility feed each other in both directions: an AI recommendation increases searches by name, and search and brand mentions increase AI visibility. A Similarweb study (June 2026, analyzed by Rand Fishkin of SparkToro) found that a brand the AI recommends is 2.5x more likely to get a site visit the following week, most of it through brand search. This is a loop, not a funnel. Caveat: the figure is aggregate across large brands, not proven for a small brand.

Broad-reach brand advertising that builds mental availability, so people search for you by name later. PR and earned media, which pay twice: a jump in brand searches and indexed mentions on trusted domains (what the AI cites). Thought leadership and visibility for the founder or executive, because people search for the brand because of the people (content published by a person clearly gets more engagement than a post from a brand page). Recurring original data, which is at once PR bait, a driver of brand search, and the snippet the LLM cites. And consistent brand assets (name, logo, colors), because that same consistency lets the AI identify you as one entity.

When is brand search a healthy signal, and when an illusion?

A healthy signal: a rising, sustained search share versus competitors, unbranded discovery that turns into branded search over time, and branded search that converts (a PPC managers' rule of thumb, not research: about 2-3x an unbranded search). An illusion signal: a one-time jump from a PR crisis, bot traffic that inflates vanity metrics, bought branded traffic (branded PPC) counted as organic demand, or branded search that rises while conversions stay flat (a sign of confusion or a clashing name).

Which social platforms feed citation, and which only demand?

The load-bearing insight: AI cites indexed, readable text that reads like a standalone answer. It doesn't cite a "platform," and it doesn't cite engagement. From here the platforms split into two jobs.

TrackPlatformsWhat it does
Direct citation (indexed text)Reddit, YouTube (long-form video + transcript), LinkedIn, QuoraRetrieved and cited directly in answers
Indirect demand (mostly visual)Instagram, TikTokDrive awareness and brand search, cited far less than YouTube

A clarification on terms: Stories, Reels and Shorts are formats within the platforms, not platforms in their own right. What determines whether something is cited is the text the format produces, not its name. Stories disappear after 24 hours and so have no permanent address to index, and are almost never cited. Reels and Shorts are video with little text: on YouTube, Shorts were only 5.7% of citations versus 94% for long-form video (OtterlyAI, March 2026).

By contrast, a feed post or carousel on Instagram does enter Google's index (since July 2025) and its caption can feed an answer.

A point worth correcting: engagement and comments don't cause citation. A post with tons of comments doesn't get retrieved more because it's popular, only if its text is indexed and readable as an answer. In OtterlyAI's study of 100 million citations (March 2026), likes, views and subscribers showed no correlation with citation frequency (41% of the cited videos had fewer than 1,000 views).

What does move the needle: description length and content freshness. Reddit is the exception: there the comments themselves are the text that gets retrieved, because the discussion is the content. On Facebook, Instagram and TikTok a high comment count doesn't make the post citable.

And what about X (Twitter)? It is almost never cited by ChatGPT, Perplexity or Gemini, because it blocks crawling in its terms of service. Its data flows directly to Grok through xAI's internal integration, not through the open web. So X is a channel to Grok only, not a general visibility channel.

How to work with social, by business type

First choose the citation-feeding platforms by audience.

  • B2B: first and foremost LinkedIn (the most-cited domain for professional queries, Profound March 2026) together with YouTube and Reddit.
  • B2C: Reddit (opinions on products), YouTube (reviews and explainers with a transcript) and review sites like Trustpilot and G2.

In both cases, Instagram and TikTok are an awareness layer that builds the brand the AI will find elsewhere, not the citation itself.

Make sure a consistent entity exists on every profile, and make sure every substantial post also produces an indexed text version on your domain: a full transcript for YouTube (published on a page on your site or in the video description), a text post, or a written summary.

Measure what actually matters: the brand's Share of Search and citation share versus competitors, not followers and likes.

Content repurposing: one item, many assets

Repurposing is how you turn one item into many assets, which both grow interest on social and grow presence in AI.

The trick: knowing which assets feed citation and which only feed awareness. Social clips buy attention and brand search (indirect GEO value). Indexed text assets (a blog, a transcript, an FAQ) are what the AI actually retrieves and cites (direct GEO value). A repurposing plan that produces only ephemeral video feeds the first and starves the second.

The model: anchor, break down, native format, hub

One rich anchor item (a podcast episode, a webinar, a long video) containing at least three distinct ideas breaks down into many single-idea assets, distributed in a native format for each platform, and reunited on a pillar page you own. This is Ross Simmonds's "Create Once, Distribute Forever" model. Simmonds estimates that 85% of marketers spend 90% of their time on creation and neglect distribution.

Why is repurposing core to GEO?

Most AI engines don't see audio or video, they read the text around them: the transcript, captions, title, description (Gemini is the exception that processes video directly).

A Reel or a raw podcast episode, on its own, is invisible to retrieval. The move that creates GEO value is to convert the ephemeral anchor into indexed, structured text. A tracking experiment found that clean transcripts were cited about twice as often as raw automatic transcripts (a small, directional sample).

How do you repurpose correctly?

One idea per asset. Adapt the format to each platform, don't publish the exact same post everywhere (what's called cross-posting, copying instead of redesigning). Always publish a clean transcript (punctuation, paragraphs, an entity-rich summary paragraph) on a domain of your own, that's the highest-leverage move and where most marketers fail. Build a pillar page that unifies everything, extract the best statistic and the expert quote as standalone attributed lines, and build an FAQ from the questions that came up.

Repurposing: what feeds citation: The same anchor splits into two tracks, and only one of them gets cited.

FAQ: how to build it right, and when it's overdone

A question-answer structure is one of the highest-value formats for AI, because it maps directly to how people ask and to how the engine breaks a prompt into sub-questions (fan-out). A standalone question-answer block is close to the ideal unit for retrieval. But it's easy to turn into fake promotional content, and that's the distinction.

What happened to FAQ schema?

A point worth getting right: Google removed the rich FAQ display from search results (in 2023 it limited it to government and health sites, and in May 2026 it dropped it for everyone). Meaning there's no FAQ visual in results for almost any site. FAQPage schema is still valid and does no harm: Google reads JSON-LD as structure for entity identification, and the other retrieval engines turn the page into text and mostly ignore the block. In both cases, what gets retrieved is the visible text on the page. Don't add schema to "get a rich result," there is none.

How do you build an FAQ right?

Real questions, not invented ones (from support tickets, sales calls, "People Also Ask," Reddit, and directly querying ChatGPT). A primary answer of 40-60 words, direct in the opening sentence, standalone. Specific, with numbers and dates. Visible and accessible on the page, not hidden. And if you add FAQPage JSON-LD, every question and answer must appear word for word on the page too.

When does an FAQ become overdone and read like promotion?

This is the important question, because these patterns read as spam: questions stuffed with keywords, invented leading questions that are really an ad ("why is [the brand] the best"), an identical FAQ block pasted on every page, and marketing copy dressed as a question and answer. The test is simple: would a real person ask this in these words? Does the answer inform first and sell (if at all) second? And can it stand alone as a cited sentence, correct and useful out of context? If you answered "no" to one of these, fix it.

Fan-Out analysis

Fan-out is a Google technique in AI Mode. Instead of answering your question directly, the engine breaks it into many sub-questions, runs them in parallel, and assembles an answer. Google confirmed this (May 2025). The implication: ranking for the main term no longer guarantees an appearance, because the answer is composed of many hidden sub-questions, and the engine cites the best paragraph for each sub-question.

It's not only Google. ChatGPT also decomposes queries, just differently, and more selectively. Every engine decomposes differently, so the same question produces different sub-questions in each engine.

Two data points on Google's AI Mode: per Ahrefs (September 2025 data) the overlap between the pages cited in AI Overviews and in AI Mode is only about 13.7%, and per Semrush (June 2026) AI Mode mentions 2.5x more names of brands and people. That is, even within Google itself, appearing in one place doesn't guarantee appearing in the other, and AI Mode leans even further toward the known brand.

How this connects to the ranking collapse: an Ahrefs study (March 2026) found that only 38% of the pages cited in AI Overviews rank in the top ten, down from 76% in July 2025 (Ahrefs note that their citation detection improved between the two checks, so some of the drop comes from the method and not only from the engine). There's a disagreement worth knowing here: BrightEdge measured an overlap of only 17%, and 5W measured a collapse from 70% to under 20%. Everyone agrees the overlap collapsed, the exact number is disputed. The accepted explanation is fan-out, even if part of the gap comes from changes in measurement methods.

What you do with this: pick a commercial target query, generate its sub-questions (by simulation or by watching AI Mode's source panel), map which of them you cover and which you don't, and brief content for each gap. It's better to add a standalone paragraph to an existing page than to spin up thin new pages.

Sub-Question Generator (manual fan-out)

Runs locally

What it does: enter a topic, and the tool suggests a list of likely sub-questions by aspect (definition, how, cost, comparison, risks, alternatives), as a coverage checklist. Input: a topic. Output: a sub-question tree + export to Excel or CSV. An honest note: real fan-out happens on the engine side; this is a manual aid for ideas, not a reveal of the actual queries. Runs in the browser in plain code, with no AI behind it.

E-E-A-T: building trust the machine reads

E-E-A-T is Experience, Expertise, Authoritativeness, Trust. Two points worth getting right, which Google states explicitly: in the Creating Helpful Content document, E-E-A-T is not a ranking factor in itself but a conceptual framework; and in the quality rater guidelines (September 2025 version), Trust is the most important, and everything else builds it. So you check trust first.

Why this matters doubly for GEO: a ranking algorithm can weigh hundreds of signals and still rank a weak page seventh. An answer engine cites few sources, usually two to five. So a trust signal that only nudges a ranking in Google can serve as a gateway to an AI citation. And the engine reads trust from machine-readable signals: a byline (the credit line stating who wrote it, "By [name]"), an author bio page, Person schema, sameAs, and external corroboration.

How to build it in practice: a named author with real credentials on every substantial page, linked to a dedicated bio page; first-hand experience (original photos, "we tested this ourselves," original data you produced); citing primary sources; and on the trust side, a real About page, contact with an address, visible update dates, and facts that don't contradict across pages.

Measurement and ongoing tracking

How do you measure for free, on your own?

The DIY method is a fixed set of 15 to 25 buying-intent questions, run monthly across the engines, and logged in a sheet: date, engine, were you cited, position, which competitors. It's a manual version of what paid tools do automatically, and without any details leaving your hands.

Prompt Library Generator

Runs locally

What it does: enter a brand, category and competitors, and the tool assembles about 10 to 15 prompts in brand, category and comparison clusters, ready for tracking. Input: brand + category + competitors. Output: a prompt list + export to Excel or CSV. Feeds the next tool. Runs in the browser in plain code, with no AI behind it.

Prompt and Citation Tracking Log

Runs locally

What it does: a private log to record, for each prompt/engine/date, whether you were mentioned, cited, and the sentiment. Exports to Excel or CSV, so the "database" is portable with no server. Input: prompt, engine, date, result. Output: a table + summaries + export. This is the free alternative to $200-a-month tools, because running prompts automatically against the engines requires a paid API and sends data out. This tool runs in the browser in plain code, with no AI behind it, and that's exactly the point.

What do you get out of the official systems?

In Google's Search Console, clicks and impressions from AI Overviews and AI Mode have been folded, since 2025, into the regular report with no way to separate them; a dedicated report from June 2026 (Generative AI performance) separates impressions only, without clicks and without queries.

Bing Webmaster Tools shows Copilot citations only (not ChatGPT), but it's the way to confirm you're in the index Copilot draws from. GA4 has, since 2026, an "AI Assistant" channel, but it's partial: Perplexity still falls under Referral, and clicks from AI Overviews count as Organic, so it's still better to set up a custom regex (a regular expression, a text pattern that catches the AI engines' addresses).

One workaround for the GSC limitation: filter Search Console for long, prompt-style queries (10 words or more) with regex, and that still works.

The second technique, tracking the Text Fragment parameter in the URL (#:~:text=) that arrived from an AI answer via Google Tag Manager, stopped working in May 2026 when Google removed the fragment from AI Overviews links, so it's worth knowing mainly for historical reasons. Credit for these techniques goes to Brodie Clark (who identified the Text Fragments method) and Dana DiTomaso (the GTM and GA4 setup).

AI-Traffic Regex Generator for GA4

Runs locally

What it does: choose which engines to include, and the tool generates the exact regex and the GA4 setup path, with an explanation of each part. Input: engine selection. Output: regex + setup instructions (including catching session_source and utm_source=chatgpt.com, not just the referrer, and matching GA4's new AI Assistant channel). It all runs in the browser in plain code, with no AI behind it.

And remember the dark-traffic problem: a large share of AI visits arrive with no referral source and fall under "Direct." Every AI-traffic number is a floor, not the truth. That's why you look at three signals together: AI traffic, share of voice, and a rise in brand searches.

The expert's tools: must-have, nice-to-have, or skip

What separates an expert from someone who buys every tool that comes out is the ability to prioritize. Most of the real, cheap leverage is in entity clarity and the ground truth of server logs. Schema and most structured-data tools are hygiene, not a lever. Here's the ranked verdict, as of September 2026 (prices are illustrative, verify them before quoting a client).

Must-have, and almost all free

  • The free base as a unit: Google Search Console, Bing Webmaster Tools (feeds ChatGPT and Copilot), GA4, and manual querying of the engines. The fastest reality check before any paid tool.
  • Cloudflare Agent Readiness Score (free, isitagentready.com): a scanner that gives any URL a 0-100 score for how readable it is to AI agents. Run it as a first pass on every client. Caveat: it's a checklist scanner, it tells you whether the door is open, not whether anyone walked in.
  • Cloudflare AI Crawl Control (free, if the site is on Cloudflare): shows which AI crawlers actually arrived, how many requests, and whether they respected robots.txt. This is server-side ground truth, the hardest signal to fake. Important: Cloudflare began blocking AI crawlers by default on new domains in July 2025, so many sites are invisible to AI without knowing it. This is where you diagnose that.
  • Screaming Frog SEO Spider (about £199 a year): the critical feature is turning off JavaScript rendering to see whether the content exists in the raw HTML.
  • Wikidata and the Google Knowledge Graph Search API (free): the entity lever.
  • The free schema validators (Rich Results Test, Schema Validator): QA hygiene.

Nice-to-have, depending on scale

Log analysis at scale (JetOctopus, about $549 a month; or Screaming Frog Log Analyser at £99 a year for the cheaper option), one content tool (Surfer, Frase or Clearscope), and Kalicube Pro for entity-focused clients with a budget.

What to skip

Tools that sell "rich schema = more GEO" at an enterprise price (expensive hygiene), and Cloudflare's blocking and monetization features (Managed robots.txt, AI Labyrinth, Pay per crawl). Your job with the latter is only to make sure they aren't accidentally hiding the site from the crawlers you want.

Myths and what doesn't work

It's a shame to waste budget on what doesn't move the needle. Here's what's overhyped:

  • llms.txt. A file that's supposed to guide AI engines. Ahrefs checked and found that 97% of the files got no requests at all. The deeper evidence points the same way: Google's own documentation, Lighthouse classifying it under "agentic browsing" rather than SEO, and server logs from dozens of large sites and SaaS sites showing the file fetched in fewer than 1 in 10,000 bot requests, with near-zero use even at the scale of a website builder's blogs. Cheap to add, but don't sell it as a tactic.
  • Schema alone doesn't buy citations. It's hygiene that helps entity identification, not a direct lever.
  • Keyword stuffing hurts. The original GEO paper found this was the one tactic that reduced visibility. Search-engine tactics don't carry over one-to-one to a generative engine.

Expert level: what else you need to know

The layer that separates a professional from an amateur.

Two games, two clocks

There's a difference between being in the model's memory from training (what it "knows" about you, frozen at the knowledge cutoff, changing very slowly) and being retrieved in real time (fast, controllable, and where the citations live). Winning a citation is page- and paragraph-level work. Changing what the engine "believes" about you when search is off is long-term entity and reputation work.

How do you manage the engine's hallucinations?

Sometimes the engine describes you incorrectly. A Tow Center study for the Columbia Journalism Review (March 2025, 1,600 queries across eight engines) found that more than 60% of answers attributed journalistic citations to the wrong source. The study examined article attribution, not brand descriptions, but the mechanism is the same: the engine answers confidently even when it's wrong, and sometimes invents rather than admitting it doesn't know.

A hallucination about a brand is usually an old or contradictory source the engine repeats, but an information vacuum is no less dangerous, because then it fills the gap itself. Your site is only a small part of what the model reads about you, most of the picture comes from external sources (a common estimate in the field, not a figure from a controlled study), so the fix is mainly to correct the external sources and make the entity unambiguous.

Multimodal: why is YouTube the story?

Most engines don't watch video, they read the text layer: the transcript, title and description. Gemini is the exception, it processes video and audio directly, and from September 2026 Google is adding such a capability to Ask YouTube too.

Even then, what gives the other engines a paragraph to cite is the clean transcript on a page of yours. YouTube is cited a lot, mainly through AI Mode, AI Overviews and Perplexity, and most citations are long-form video, not Shorts. So a video with a clean, published transcript is one of the highest-leverage assets.

The agentic web: why does it matter already?

An emerging world where AI agents act on the user's behalf: searching, comparing and even buying. There are already new protocols in this space: Anthropic's MCP connects an agent to tools and data, and ACP (OpenAI and Stripe) and AP2 (Google) are commerce and payment protocols built on top of it.

WebMCP is an experimental interface that lets a site expose ready-made actions to an agent instead of the agent guessing where to click. Its status as of 2026: experimental only, running as an Origin Trial in Chrome (versions 149-156) that ends in November 2026, and only some agents support it. The recommendation: don't deploy it to production yet, but do experiment with a single flow. What you tell a client: the purchase may move from a human browsing to an agent acting, so machine accessibility and feeds may outweigh a pretty page. But in fairness: OpenAI shut down Instant Checkout in chat in March 2026, after only dozens of merchants joined (the ACP protocol continues). Prepare the infrastructure, don't rely on it for the coming quarter.

Content and news sites: a different clock

If you publish news content, the game is different. The essentials: a news article has a short window to be included in the News Sitemap (about 48 hours), so you need fast indexing (a News Sitemap together with WebSub, which notifies Google of every article), NewsArticle schema with accurate publish and update dates, and server-side rendering so the headline and body are already in the initial HTML. Speed is a strong operational recommendation, but it's not a formal threshold for Top Stories: Google states explicitly that a page can be included regardless of its Core Web Vitals score. And a point that ties straight to GEO: in mid-2026 large publishers (Reddit, Gannett, Reuters) considered limiting Google's access because of traffic lost to AI Overviews, and UK regulation (the CMA) requires Google, from June 2026, to allow opt-out with no ranking penalty (with nine months to implement).

Priority order: what to do first

If there's one thing that separates an expert from someone who buys every tool that comes out, it's the ability to prioritize. Here's my priority order, from most important and urgent to least impactful. 10 is "do this first," 1 is "barely matters." The number is an estimated impact score, not a continuous ranking, which is why some numbers are skipped.

  • 10: Make sure the bots arrive and see content, and render server-side. A precondition. If the bot gets a 403 (usually from a WAF, not from robots.txt) or doesn't see content, you get zero regardless of quality. Blocking training bots is a separate business decision.
  • 9: Brand strength and earned media. The big lever, though slow and expensive. This is where 84% of citations sit.
  • 8: Original data and statistics in the content, an answer-first structure, getting into third-party lists, and a YouTube presence with a transcript.
  • 7: Entity clarity (Wikidata), a Reddit presence, freshness, content repurposing, and prompt research and measurement.
  • 6: Author E-E-A-T, review sites, format by question type, and a real FAQ.
  • 3: Schema. Hygiene, not a lever.
  • 1: llms.txt. Skip it.
The 1-10 priority ladder: Not everything is equal. Here's what to do first.

Experts worth following

Rand Fishkin (zero-click research, "audience over traffic"), Mike King of iPullRank (technical GEO, "relevance engineering"), Aleyda Solis (the #SEOFOMO newsletter), Lily Ray (E-E-A-T and AI Overviews analysis), Kevin Indig (in-depth GEO essays), Marie Haynes (trust signals), and Dan Petrovic of Dejan (LLM retrieval experiments).

Where to start: a 90-day plan

If the whole guide feels like a lot, here is the practical order, without reading all of it.

  • Week 1, the technical gate. Confirm in the logs that the search and retrieval bots get a 200 and that the core content is in the initial HTML. Without this the AI simply can't see you, and nothing else matters. In parallel, make sure the site is indexed in Bing, which feeds some of the AI engines.
  • Weeks 2 to 4, entity and a measurement baseline. Align the entity (name, profiles, sameAs) so it is consistent, and capture a baseline: 30 to 50 key questions, how many mention and cite you versus competitors.
  • Months 2 to 3, brand and extractable content. Earned media, original data you can quote, comparison pages with an answer-first opening, and a YouTube presence with a published transcript.
  • Ongoing. Track the leading metrics weekly, decide and report monthly, and verify every number against current data.

Ten mistakes that quietly kill your GEO

  • Accidentally blocking the retrieval bot with a User-agent: * rule, then vanishing from answers without knowing why.
  • Serving the core content client-side only (CSR), so the bots see an empty page.
  • Averaging data across engines instead of running a separate strategy for each.
  • Chasing schema as if it were a lever, instead of investing in entity clarity.
  • Spinning up thin new pages instead of adding a standalone paragraph to an existing one.
  • Judging an engine on a single run instead of a fixed question set run again and again.
  • Building a "best of" list about yourself instead of getting included in others' lists.
  • Selling a client a guaranteed position or exact attribution. No platform gives that.
  • Measuring only ChatGPT and forgetting Google's AI Overviews, where the most traffic stops.
  • Waiting for a perfect number before you start measuring, instead of starting now and fixing as you go.

Summary

GEO is neither a bug nor a buzzword, it's how a growing part of your audience discovers things. The most important point recurs in almost every 2026 study: what has always worked in marketing, a strong brand, a real reputation and content that genuinely helps, is exactly what makes the engine recommend you. The difference is that now you also have to make it machine-readable. Start from the fundamentals, measure right, and don't chase the tricks. This is a process of months, not a week, and whoever starts now will be there before most of the market even begins.

Went through the guide and spotted gaps you want to close? That is exactly what I do with clients. If you'd like a hand, get in touch.

About the author. Etgar Shpivak is Head of specialization at Ono Academic College. He co-founded Fixel in 2018 and ran it as CEO, and Logiq acquired it in 2020. He now works with seed and Series A founders on marketing and go-to-market, including how AI engines describe their company. More about him on his bio, and more guides on the writing shelf.

This guide was written in September 2026 and its data is current as of then. The field moves fast, so verify every number before you present it to a client.

The tools in this guide

All thirteen tools from the guide, in one place. Every one runs in your browser and sends nothing anywhere. Come back to any of them, and copy a direct link (to save or share) from inside the tool itself.

Questions and answers

What is GEO?

GEO (Generative Engine Optimization) is the set of actions that get an AI engine like ChatGPT, Gemini or Perplexity to cite you, weave you into the answer and recommend you. It is a layer on top of SEO, not a replacement: the goal is to become a source the engine trusts and quotes inside the answer itself, not just to rank high on a results page.

What's the difference between GEO and SEO?

In SEO the question is how to rank high on the results page and win the click. In GEO the question is how to become a source the engine quotes inside the answer. Much of what you did for SEO (clear content, clean structure, a fast site that bots can reach) still applies, but GEO adds a layer of brand strength, entity clarity and an easily extractable structure.

How big is ChatGPT's market share?

There is no single number, and that is the point. The share runs from 46% to 78% depending on the metric: about 46% of active audience (Sensor Tower, May 2026), 44% adoption among Americans (Pew, February 2026), and 78% of referral traffic to sites (StatCounter, March 2026). Anyone who quotes one number without naming the metric is misleading you.

How long until GEO shows results?

Technical fixes show up within days to weeks, but brand and off-site work takes months. There is no precise research on this, it is a vendor estimate, but a 30-day trial only tests the easy part. A realistic expectation: first citations in the early months, and a stable correlation to a rise in brand search only after six months or more.

What matters most for AI visibility?

Most 2026 research converges on the same conclusion: brand strength and earned media beat technical tricks. Per Muck Rack (May 2026, 25 million links), 84% of AI citations come from third-party sources, and only 0.3% from paid media. Investment in your brand, in PR and in content others cite is exactly what makes an engine recommend you.

Does schema (structured data) help GEO?

Schema does not create authority, it removes ambiguity so the machine knows exactly who you are. It is a baseline and hygiene, not a lever: it helps the engine understand you, but does not make it prefer you. The real lever is entity clarity and consistency, and most of it is free.

Where do most AI citations come from?

From earned media: third-party sources like press, academia, Wikipedia and review sites, not from your own site. Each engine has a different source mix: ChatGPT leans on Wikipedia and news, Perplexity on Reddit, and Copilot on shopping sites. You cannot run one source strategy for every engine, and the ranking shifts fast, so verify it against current data.

What do you do first in GEO?

By priority order: first the technical gate (don't block the AI bots, render server-side, and get indexed in Bing). Then brand strength and earned media, original data and an answer-first structure, and only then schema hygiene. On question types, start with brand questions (people already searching for you), then comparison questions, and finally the general ones.

Do you need an llms.txt file?

llms.txt is a proposed standard meant to give AI engines a clean content map of your site. As of September 2026 no major engine confirms it reads the file, so don't let it crowd out real fixes. It is cheap to add and harmless, but prioritize it only once the basics (server-side rendering, indexing, an answer-first structure) are already in place.

Can you measure AI visibility for free?

Yes, at first. Define a set of 30 to 50 key questions, ask them manually in each engine once a month, and count how many answers mention you and how many link to you. That gives you a real baseline for share of voice and citation rate without paying for a tool. Paid tools save time and widen coverage, but they are not a prerequisite to start. The tools in this guide do the math for you.

I may have blocked the AI bots by accident, how do I check?

The most common mistake: a robots.txt that blocks everything for User-agent: * and only opens up for Googlebot, so the AI bots (which have no dedicated rule) fall under the blocking rule. Check your robots.txt against each engine's user agents (GPTBot, OAI-SearchBot, PerplexityBot and others), and confirm in your server logs that they actually get a 200. The tool in this guide checks this for you.
Etgar Shpivak, marketing and go-to-market advisor

About Etgar Shpivak

Founder and marketer. Co-founded Fixel in 2018 and ran it as CEO; Logiq acquired it in 2020. Head of specialization at Ono Academic College and host of the podcast Founders' Marketing Compass. Works directly with seed and Series A founders, starting with the offer.

Cite as: Etgar Shpivak, "The Complete Guide to GEO: Generative Engine Optimization, From the Basics to Real Implementation", shpivak.co.il, 16 September 2026. https://shpivak.co.il/writing/geo-guide

Tell me what stopped working

I help seed and Series A founders run their marketing, week by week.