GEO Infrastructure: The Technical Checklist for Getting Recommended by AI
By Zaira Céspedes, SEO & GEO Account Manager
The initial panic about AI search has settled. “SEO is dead” isn’t the take anymore, and thank goodness, because it’s the wrong one. SEO isn’t dying, but at the same time, the strategies that worked five years ago are becoming obsolete, and what’s replacing them needs a much harder push on technical integration and cross-channel strategy.
SEO used to be a fairly blunt pursuit: get ranked and be visible on a page of blue links. Now it’s contextual. Subtle signals matter, and users actively avoid anything that seems too commercial.
When online search was new (which was around 20 years ago), we learned that we had to adapt to the technology to the point we didn’t even notice anymore. We learned to compress what we actually wanted into two or three keywords and hoped the results page understood us. Generative AI flips that. For the first time, we can just talk, the way we’d talk to a friend, a neighbour, or someone behind a counter, asking what they’d recommend. That’s the real shift, and it’s why adoption has moved so fast.
Which raises the question this checklist is built to answer: if users aren’t typing keywords anymore, and are instead having real conversations, how do brands position themselves inside those conversations?
The Reality of AI Discovery
Search behaviour is decoupling as users research within AI answers, cross-check across platforms, and convert without ever visiting a website. In early 2026, AI search engines handled 12–18% of English-language informational queries, up from just 4% two years ago. But while traditional results pages show 10 links (at best), most AI answers cite only 2 to 7 sources. That creates a dramatically smaller pool of winners per query.
This is the part that should pique the attention of SEO teams: zero-click behaviour isn’t a fringe case anymore; it’s becoming the default for a growing share of informational queries. A user asks a question, gets a synthesised answer with two or three citations and either acts on it immediately or moves on, no visit, no session, no pixel fired. For years, we measured success in traffic and rankings because that’s what we could see. Now, the most valuable outcome is being the brand that an AI model chooses to name, and produces no trackable clicks at all. That’s forcing us to face an uncomfortable situation this year: the KPI that used to prove SEO was working is becoming the wrong KPI. We’re having to build new baselines, brand-inclusion tracking, citation-share audits and sentiment-in-answer scoring, which measure a result that used to be invisible to dashboards.
The brands falling out of that pool aren’t necessarily the ones with weak content. Often it’s legacy digital architecture, sites built for a crawler that indexes pages, not for a system that has to retrieve and synthesise an answer in a few hundred milliseconds. That’s the gap this checklist addresses.
Re-architecting for Retrieval: Crawlers, Rendering and Machine Readability
Traditional search indexing and AI retrieval are not the same job, even though they both start with a crawler.
Can AI actually see your content?
Before an AI can recommend you, it has to be able to read your page. That sounds obvious, but it’s where many sites quietly fail. A browser runs code that fills a page with content. Most AI crawlers don’t. They read the raw page and move on. Picture a shop whose window display only appears once the lights switch on. Google switches the lights on. Most AI crawlers just see the dark shop. So if your pricing, product specs or key claims only load through JavaScript, then ChatGPT or Perplexity may never see them. The fix is server-side rendering (SSR), where the server sends the page with the content already built in. For enterprise sites on heavy JavaScript frameworks, it’s worth doing on the pages you most want AI to quote.
How AI picks what to quote
Once AI can read your page, the next question is what it uses. Tools like ChatGPT and Google’s AI Mode don’t just rely on memory. They look things up first, then write their answer from what they found. This is called retrieval-augmented generation, or RAG. The important bit is that AI doesn’t quote a whole page. It quotes the small section that answers the question. So it’s not enough for your page to be good overall. Your answer needs to be easy to find and easy to lift out. If it’s buried in a long paragraph that covers four other topics, it may never get picked.
Managing Bots and llms.txt
Bot management is now more granular than a single robots.txt entry per company. OpenAI, for example, runs separate crawlers for training (GPTBot), for its search index (OAI-SearchBot) and for on-demand fetches when a user pastes a link into ChatGPT (ChatGPT-User); Anthropic runs the equivalent split with ClaudeBot, Claude-SearchBot and Claude-User. That granularity matters strategically: a brand can choose to opt out of training crawlers while deliberately keeping search/retrieval crawlers open, which is a coherent and increasingly common position rather than an all-or-nothing block.
On llms.txt specifically, it’s worth setting clear expectations rather than over-relying on it. A large-scale analysis showed that roughly 300,000 domains found no statistically significant correlation between having an llms.txt file and AI citation frequency for general queries, and among the 50 most AI-cited domains in the study, only one had the file. Google’s own guidance is blunt: AI Overviews (AIOs) and AI Mode draw from the standard Search index, and no special file changes are required.
John Mueller was asked directly in January 2026 whether llms.txt helps, answering simply that it does not. What has changed is that Google’s Chrome Lighthouse added an llms.txt audit under a new agentic-browsing category, so it’s shifting from a “citation hack” into low-cost hygiene for AI agents navigating a site, not for AI answer visibility. Our recommendation: implement it, but position it as infrastructure housekeeping rather than a lever that moves citation numbers on its own. The actual leverage sits in the sections below.
How to Define Entity Authority by Moving Beyond Keywords
Once a page is retrievable, the next question an AI engine asks is: Who is this, really? That’s an entity question, not a keyword question, and it’s answered with structured data rather than copy.
Schema markup as an entity graph: Organisation, Product and FAQ schema don’t just decorate a page for rich snippets anymore; they give AI models an explicit, machine-readable definition of who you are, what you sell and what you claim. A page with strong prose but no schema asks the model to infer identity from context; a page with a complete schema states it outright.
Connecting authority nodes: This is where many enterprise sites fall short. Schema markup describes your brand on each page, but the model still has to work out whether all those descriptions are about the same company. You can help it in three ways:
- Use @id tags. Think of an @id as an ID badge for your organisation. Give it one unique badge, then reuse it everywhere, so the “Organisation” in your Product schema and the one on your About page are clearly the same company. Without it, the model has to guess.
- Add sameAs links to your official profiles, such as LinkedIn, your Wikipedia page and your Wikidata entry. These are sources AI models already know and trust, so they work like references that vouch for who you are.
- Keep your facts consistent. Pricing, product specs and key claims should match everywhere: your own site, Wikidata, review platforms and any partner listings. Models cross-check. If your pricing page says one thing and a directory says another, the model becomes less sure of you, and it’s less likely to recommend you.
Writing Content That LLMs Can Actually Lift
Getting crawled and getting recognised as an entity both feed into the same final test: can a model actually lift a clean, accurate, quotable answer out of your content?
Structuring “Answer Capsules”
The most extractable pages include a concise, direct summary of two or three sentences that fully answer the implied question immediately under each H2/H3, before any scene-setting or narrative lead-in. This isn’t just good practice for AI; it’s exactly what featured snippets have been rewarded for years. It matters more now: one analysis of citation patterns found that 44% of all LLM citations are pulled from the first 30% of a page’s text, so burying the answer under three paragraphs of preamble is a direct, measurable cost.
Structured Formats
HTML tables, bulleted lists and explicit Q&A blocks are easier for a retrieval system to chunk cleanly and quote accurately than dense prose paragraphs, where the “answer” might be split awkwardly across a sentence boundary.
Demonstrating Depth and E-E-A-T
Original data, named expert sourcing and genuinely comprehensive coverage of a topic are rewarded, not because “AI likes long content”, but because thin, generic pages simply don’t survive the relevance-scoring step of RAG when better-sourced competitors exist in the retrieval set.
How LLMs Validate Knowledge: The Consensus Signal
This is arguably the most under-discussed part of GEO, and is worth explaining making sure we are all on the same page, because it changes where the budget should go. Models don’t just trust your domain because you say something about yourself; they look for agreement across independent sources before they’ll cite or recommend a brand.
AI platforms gain citation confidence when a brand appears consistently across multiple independent sources, looking for agreement rather than picking the single most authoritative source. If a product is described consistently on the brand’s own site, in a Reddit thread, on a review platform, in an independent comparison article and in a YouTube review, the model has several independent data points converging on the same answer, and that agreement is what tips the model from “aware of” to “willing to recommend”.
Digital PR, trusted vertical review hubs, and community platforms (Reddit in particular) carry real weight with several engines specifically; Reddit threads are weighted heavily by Perplexity, Gemini and Google AIOs, and long-term presence in two or three relevant subreddits compounds over time. This is also a useful moment to note that GEO isn’t a single, uniform target.
Decoding AI Perception
Beyond simple visibility tracking, the more valuable audit now is qualitative: asking the models directly how they characterise a brand’s strengths, weaknesses and points of difference relative to named competitors, across several prompts and engines. It’s often the first time we can see, in plain language, how an AI model is currently summarising brands to their own prospective customers, and it’s frequently not what the brand’s own messaging says.
GEO isn’t a replacement for SEO; it’s the next layer on top of it, pulling organic visibility, paid search, digital PR and brand equity into a single, less siloed strategy than most organisations currently run.
The starting point for any brand serious about this isn’t a content sprint. It’s a baseline: a technical audit of crawlability and rendering, a schema and entity-consistency check, and an initial brand-inclusion score across the engines that matter most to that audience, so that every piece of work afterwards is measured against a number, not a feeling.
Book a discovery call with our team to understand how you can adopt GEO best practices.
Sources:
-
- neuron. What is llms.txt? The New Technical SEO Standard for AI Crawlers. 2026. Available online at: https://neuronwriter.com/?p=28194
- Vercel. The Rise of the AI Crawler. 2026. Available online at: https://vercel.com/blog/the-rise-of-the-ai-crawler
- ahrefs. Retrieval-Augmented Generation (RAG) Explained: How AI Decides Which Pages to Search & Cite. 2026. Available online at: https://ahrefs.com/blog/retrieval-augmented-generation/
- Godberry Studios. llms.txt vs robots.txt: The New AI Web Standards Every Site Owner Needs to Know. 2026. Available online at: https://godberrystudios.com/posts/llms-txt-vs-robots-txt-ai-web-standards-2026/
- GEO Toolbox. llms.txt: What It Is and Whether It’s Actually Worth It. 2026. Available online at: https://geotoolbox.ai/blog/llms-txt
- Averi. The GEO Playbook 2026: Getting Cited by LLMs (Not Just Ranked by Google). 2026. Available online at: https://www.averi.ai/blog/the-geo-playbook-2026-getting-cited-by-llms-%28not-just-ranked-by-google%29
- Derivative. What Makes a URL More Likely to Appear in LLM Citations: 6 Page-Level Factors, Ranked by Impact. 2020. Available online at: https://derivatex.agency/blog/what-makes-url-likely-llm-citations/
Suggested articles:
