TL;DR: Traditional SEO gets people to click your link. GEO — Generative Engine Optimization — gets AI models to cite your brand inside their answers. The optimization targets are completely different: passages instead of pages, entities instead of keywords, crawler access instead of crawl budget. This is the complete playbook I use, including what I found when I audited my own site with it — the embarrassing parts included.
The moment the game changed
When someone asks ChatGPT "what's a good tool for custom merch," the answer contains three or four brand names. Either you're one of them, or you don't exist for that user. There is no page two. There isn't even a page one — just an answer.
That's the entire case for GEO in three sentences. The traffic entry point is moving from the search box to the chat box, and the currency is moving from ranking to citation.
I run SEO and GEO for a living — I'm the founder and CEO of Custyle, and I built FasterGEO, an open-source GEO toolkit. This post is the field guide I wish existed when I started: what actually changed, the three levers that work, and the results of turning the audit cannon on my own site.
SEO vs GEO: same goal, different physics
Both disciplines want the same thing — your brand in front of a person with intent. But the mechanics barely overlap:
- SEO optimizes ranking signals: backlinks, keywords, Core Web Vitals, crawl efficiency. The unit of competition is the page, and the prize is a click.
- GEO optimizes citability: structured facts, self-contained passages, entity clarity, AI-crawler access. The unit of competition is the paragraph, and the prize is being quoted — often without a click at all.
The overlap is real but thin: clean HTML, fast responses, and honest content help both. Everything else diverges. A site can rank #1 on Google and be invisible to every AI engine, because the things that make a page rank are not the things that make a passage quotable.
Here's the sentence I repeat to every client: models cite passages, not pages. A cleanly structured FAQ answer can outperform a 10,000-word pillar page in AI answers, because the model needs a self-contained unit of meaning it can lift and attribute — not a scroll-depth journey.
Lever 1: Let the crawlers in — all of them
The first lever is embarrassingly mechanical: most sites block or throttle the crawlers that feed AI engines, often by accident.
The detail almost everyone misses is that training crawlers and retrieval crawlers are different user agents with different jobs:
| Job | User agents |
|---|---|
| Model training | GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Meta-ExternalAgent |
| Live retrieval (answers cite you today) | OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot |
Blocking training bots while allowing retrieval bots is a legitimate policy choice. Blocking retrieval bots means AI answers can never cite you, period. My robots.txt enumerates both layers explicitly — not because the wildcard rule wouldn't cover them, but because a policy you can read is a policy you can audit.
Two files extend this lever:
- llms.txt — a curated map of your site for AI agents, per the llmstxt.org convention: who you are, what each page contains, one line per article with a real description.
- llms-full.txt — the extension I'd call underused: every post, full text, as one plain-markdown file. My entire blog is ~70KB of markdown. One fetch, zero HTML noise, no JavaScript required. For a retrieval pipeline paying per token, that's the difference between parsing 30KB of markup per page and reading the actual words.
One more mechanical point: my article pages are ~30KB of HTML of which about 2.5KB is visible prose — the rest is markup and inlined styles. Crawlers strip tags, but token-metered pipelines still pay for the noise. Static HTML with the full text server-rendered is table stakes; a plain-text full-content endpoint is the courtesy that gets you read more often.
Lever 2: Write passages a model can lift
A passage earns citations when it survives being ripped out of context. My working rules:
- Self-containment. Each paragraph answers one question completely. If a paragraph needs the previous three to make sense, the model can't use it alone — so it won't.
- Definition sentences. "GEO is the practice of optimizing content to be cited by generative AI engines" is liftable. A vibes-based intro is not. Every key concept on your site should have exactly one canonical "X is Y" sentence somewhere.
- Numbers with provenance. "3,034 founder profiles, 99.4% LinkedIn coverage, scraped in one evening with plain curl" — from my own dataset — is the kind of specific, attributable fact models love to quote. A number without a source is a liability; models increasingly skip unattributed claims.
- Questions as headings. A heading phrased as the question users actually ask, followed by a direct answer in the first sentence, is the passage shape that wins both featured snippets and AI answers.
- First-party data beats commentary. Models have infinite commentary in their training sets. What they lack is fresh, specific, verifiable fact. Publish the numbers only you have, and you become the citation instead of the echo.
Lever 3: Become an entity, not a string
This is the lever almost nobody works, and it's where I have the most uncomfortable first-hand data.
While auditing this site, I searched my own name. Every result for "Arron
Young" was someone else — a court case in Alaska, a real-estate agent in
Texas, a high-school athlete. As far as the knowledge graph was concerned, I
didn't exist. And the reason wasn't missing schema on my site — it was that
my entity graph had only outbound edges: my site declared sameAs links to
GitHub, X, and LinkedIn, but nothing pointed back. My GitHub profile
said just "Arron" with an empty website field. My own product's site,
fastergeo.co, mentioned my username eight times and my actual name zero
times.
Search engines and AI models do entity reconciliation the same way: they
trust bidirectional identity claims. A sameAs graph with no return
edges is an unverified assertion, and unverified assertions lose to a Texas
realtor with a complete profile.
The practical checklist that follows from this:
- Give your Person schema an
@id(mine ishttps://arronyoung.com/#person) so every article'sauthorfield, on every site you write for, can reference one canonical node. - Point the edges back. Your GitHub name and website field, your X bio link, the author box on every site that publishes you — each one is a vote that the string "Arron Young" resolves to this entity.
- Disambiguate proactively. My schema carries
alternateName: ["Arron", "arronyounging"]because models silently "correct" Arron to Aaron. If your name, brand, or product has a common misspelling, declare it — don't let the model average you away. - Keep sameAs semantically honest. Products you built are not you;
I moved FasterGEO and SnapCue out of
sameAsand intoowns. Sloppy entity claims make the whole graph less trustworthy.
What I found auditing my own site
Build in public means auditing in public. This site — built by a GEO practitioner, weeks old — failed its own audit in ways worth sharing:
- Three indexable copies of everything. The apex domain, the www subdomain, and the hosting provider's preview domain all served the full site with no canonical tags and no redirects. Textbook duplicate content, invisible until you look.
- A trailing-slash tax on every crawl. The sitemap and every internal link pointed at URLs that 308-redirected to their trailing-slash form. Every crawler paid double for every page.
- An internal QA scorecard published by accident. One syndicated post shipped with the content pipeline's self-review table still inside — including a claim that FAQ schema existed when it didn't. Machines read everything; they would have read that too.
All fixed the same day. The meta-lesson: GEO failures are mostly silent. Nothing looks broken in a browser. You find them by fetching your site the way a machine does — with curl, with a crawler UA, with no JavaScript — or by running tooling that does it for you.
Why I built FasterGEO
The GEO tools I could buy in 2025 fell into two buckets: consulting deliverables wearing a SaaS interface, or single-engine dashboards priced for enterprise contracts. Neither survived contact with my actual workflow, and none of them handled the half of the internet I also care about.
So I open-sourced the workflow instead. FasterGEO is 11 npm packages plus an MCP server, built around three ideas:
- A dual-engine matrix: CN + global. Most GEO tooling watches ChatGPT, Perplexity, and Gemini and stops. But Chinese AI engines — Doubao, DeepSeek, Kimi, Yuanbao — are a parallel answer ecosystem with different crawlers, different sources, and different citation behavior. A brand selling in both worlds needs visibility in both matrices; almost nobody measures the second one.
- An entity funnel, not a keyword list. The pipeline goes from "does the model know the entity exists" → "does it describe the entity correctly" → "does it cite the entity for the queries that matter." Each stage has different fixes; collapsing them into one "AI visibility score" hides the actual work.
- Machine-verified tickets. Every recommendation ships as a concrete change with a check the machine can re-run. "Improve your content" is consulting. "llms.txt missing per-article entries — here's the diff, and here's the validator that turns green when it's live" is engineering.
The MCP server matters more than it looks: it means an AI agent — the same kind of fleet that runs my publishing pipeline — can run detection, propose fixes, apply them, and verify, without a human in the loop. GEO work itself is becoming agent work.
The 2026 playbook, in order
If you run a brand site and want the shortest path:
- Fetch your key pages with
curlusing GPTBot's and PerplexityBot's user agents. If you don't get a 200 with full server-rendered text, fix that before anything else. - Enumerate training vs retrieval crawlers in robots.txt as deliberate policy.
- Ship llms.txt with one described line per important page; ship llms-full.txt if your content fits in a few hundred KB.
- Collapse your duplicate hosts and unify your URL forms. Every silent redirect is a tax on being read.
- Give every page one canonical "X is Y" definition and at least one first-party number with a source.
- Anchor your entity: Person/Organization schema with
@id, then spend an afternoon making the edges point back — profiles, author boxes, bios. - Re-check monthly. Answers move faster than rankings ever did.
FAQ
What is GEO in one sentence?
GEO (Generative Engine Optimization) is the practice of making your content and brand citable by AI engines — ChatGPT, Perplexity, Gemini, Doubao, DeepSeek — so their answers mention you, the way SEO makes your pages rankable by search engines.
Does GEO replace SEO?
No. Search boxes still drive enormous traffic, and the disciplines share a foundation of crawlable, honest, well-structured content. But they optimize different units (passages vs pages) for different judges (models vs ranking algorithms), and budgets are already splitting between them.
What's the fastest GEO win for a small site?
Crawler access plus llms.txt — usually under an hour of work. Most small sites have never checked whether AI retrieval bots can read them at all, and a described map of your content is the cheapest way to make one fetch count.
How do I measure whether GEO is working?
Ask the engines your buyers' questions, on a schedule, in every market you sell to — and track three stages separately: does the model know your entity, does it describe you correctly, does it cite you for queries with intent. That funnel is exactly what FasterGEO automates, in both the global and Chinese engine matrices.
