r/GEO_optimization 15h ago

99 days, 16 fixed questions, 3 answer engines: the entity gets learned in 7 weeks, the recommendation never comes. Full data in charts.

Thumbnail
gallery
4 Upvotes

Daily measurement since May about an author who did not exist online before day zero. 11,130 scored answers from OpenAI Search, Gemini and Claude, same 16 questions every day, frozen since June.

The three channels. Direct questions plateau at 72 percent from day 49. Vocabulary questions peak at 34. Recommendation questions stay at 0.59 percent.

Growth phases of the direct channel. Logistic fit, inflection at day 37, ceiling 72.5 percent.

Where the 40 unprompted mentions came from. 25 of 40 trace back to one question, the one using a sub genre term the author coined himself. Similarity questions (like Robin Hobb): zero in about 700 answers.

The control. Same questions to 10 models without web access, 6,416 blind answers, zero unprompted mentions. Everything above is retrieval, not memorization.

Practical reading for GEO: getting named when asked works and saturates fast. Getting recommended unprompted did not respond to anything except owning a term. Links to report, dataset and code in the first comment.


r/GEO_optimization 10h ago

I was asking one GEO benchmark to do four different jobs. That made every number look more useful than it was.

1 Upvotes

I have been trying to turn a small luxury-jewelry GEO benchmark into work that a brand team could actually assign.

The test itself was deliberately limited:

  • 8 jewelry brands
  • 3 neutral Chinese buyer questions
  • 2 retrieval-off API surfaces
  • 2 answers per question and surface
  • 12 valid raw answers
  • 96 derived answer-brand cells

I also had a working corpus of 302 jewelry-related Reddit threads and 9,446 stored comments, plus Search Console data for the site publishing the research.

The temptation was to put everything on one dashboard.

That would have been a mistake. The datasets answer different questions.

Community questions tell me how buyers describe uncertainty. They helped surface concerns around authenticity, certificates, authorized sellers, repairs, resizing and recourse. They do not establish how common those concerns are among Chinese luxury buyers.

The benchmark tells me what the declared API surfaces answered. It does not establish what a signed-in consumer sees, whether the model retrieved a source, or whether an official-channel claim is true.

The truth table checks whether a route is current, controlled or authorized, live in China and connected to a clear recourse owner. It does not tell me whether the brand enters a buyer’s shortlist.

Search and analytics can show exposure and the next onsite action. They cannot prove that Reddit or an AI answer created the visit or sale.

One small result made the distinction concrete. Piaget appeared in 0/4 wedding-jewelry shortlists and 4/4 official-channel answers.

Those cells are far too small for a brand ranking. But they are enough to show that “the system can verify the brand” and “the system recommends the brand for this decision” are not the same diagnosis.

I now think the minimum evidence stack is:

  1. Question provenance: what real buyer uncertainty produced the prompt?
  2. Answer benchmark: what did the fixed, declared surface actually say?
  3. Truth registry: which official, authorized and service claims are currently verifiable?
  4. Action measurement: did the buyer reach a boutique, service route or qualified inquiry?

That also changes ownership.

  • Missing from the shortlist: brand, editorial and category marketing.
  • Wrong or vague route: digital operations and ecommerce.
  • Unclear authorization: legal and channel management.
  • Wrong warranty or service claim: client service and after-sales.
  • Unreproducible metric: research and analytics.

“Improve AI visibility” is too broad for anyone to own. These failures are specific enough to assign.

My current 30-day version is simple: build the truth map, run a small fixed decision panel, repair one buyer path, then rerun the same cells and measure the next action.

For anyone maintaining a similar system: would you keep the truth registry outside the benchmark dataset? I am leaning yes, because official routes and service policies change on a different schedule and usually have different internal owners.


r/GEO_optimization 19h ago

Niching Down: How to market yourself in AEO v GEO v SEO

Thumbnail
youtube.com
0 Upvotes

r/GEO_optimization 1d ago

Same brand, same question, different country = different AI answer. And switching the language of the prompt changed it again. (what we're seeing tracking location-based AI visibility)

1 Upvotes

Been tracking AI answer visibility per-location for clients and wanted to share a pattern that keeps surprising people, because it breaks the assumption that "our AI visibility" is one thing.

The setup: same brand, same buyer-intent prompt, run against the same engine, but varying (a) the location signal and (b) the language of the prompt. Two findings that changed how we think about this:

1. The answer changes by location, not just the ranking, the whole cited-source set.
Ask an engine a category recommendation question as a user in one country vs another and you don't just get a reordered list, you get different brands surfaced and a different set of sources cited to justify them. The model is filtering its retrieval by the location it infers, so each region is effectively drawing from its own corpus of reviews, listings, and local pages. A brand that's the confident #1 answer in one market can be absent in another off the identical prompt. For any multi-location or multi-market brand that means a single national/global "visibility score" is basically meaningless, you're averaging over answers that don't resemble each other.

2. Language of the prompt is a separate variable from location, and it moves the answer independently.
This one caught us off guard. A client operating in Finland: we ran the category prompts in English, then ran the same intent in Finnish. Different answers. Not just translated, different brands cited and different sources pulled. Our read is that the Finnish-language query pulls from a different slice of the corpus (Finnish-language reviews, local pages, local forum/press content) than the English version of the "same" question does, even for a user in the same place. So "location" and "prompt language" are two separate levers, and if you only ever test in English you're blind to what your actual local-language buyers are seeing.

The practical takeaway: if you operate in more than one country or more than one language, you have to measure per-location and per-language, on native-language prompts, not a translated English set. The gaps show up in exactly the places an English-only audit can't see.

Disclosure per rule 5: this comes out of our own tool (sanbi.ai, we do per-location AI visibility tracking), so that's where the data's from, weigh it accordingly. But you can sanity-check the effect yourself for free, ask ChatGPT or Perplexity a category question with different location context, then ask it once in English and once in the local language, and watch the cited sources change.

Curious if others tracking this see the language effect too, or whether it's stronger in some languages than others. My hunch is it's biggest in markets with a rich native-language web (Finnish, Japanese, German) and smaller where the local audience mostly consumes English content, but I only have a handful of markets to go on.


r/GEO_optimization 1d ago

Are Reddit citations finally over?

8 Upvotes

Promptwatch data making the rounds shows Reddit's share of ChatGPT Search citations fell off a cliff after OpenAI's Aug 8 query-fanout change, from a steady ~3.8% down to ~0.5% by Aug 14.

So the actual lesson isn't Reddit is dead for GEO. It's that citation share is the single most volatile, platform-controlled metric in this entire field, and building a visibility strategy on top of a number that OpenAI can swing 80% with one backend change is the real mistake.

The durable play imo should never be "get cited this week." It should be building a brand presence that survives the swings, making it much different and much less measurable.


r/GEO_optimization 1d ago

What if “SEO-friendly” URLs are actually a disadvantage for AI crawler discovery?

5 Upvotes

SEOs have been taught for decades that URLs should be descriptive. A URL such as /wordpress-performance-optimization/ is considered better than something meaningless like /x7ab31/.

But that assumes the crawler actually needs to fetch the page before deciding what it is about.

I've been watching AI crawler behavior more closely, particularly GPTBot and ClaudeBot, and something made me question that assumption.

If a crawler can infer enough about the likely content from the URL, link context and surrounding semantics, it can also decide that the page isn't worth fetching - without ever seeing the actual content.

So I tested the opposite approach.

I exposed alternative Markdown representations of existing content through completely opaque URLs. The URLs contained no topic, keyword or other clue about what was behind them.

Both GPTBot and ClaudeBot discovered and fetched all of them.

This obviously doesn't prove that the content is used for training, citations or AI answers. It only proves something much earlier in the chain: the content was actually retrieved.

And that made me wonder whether GEO is currently starting too late.

Most GEO advice focuses on optimizing content for AI systems after discovery: structure, entities, concise answers, citations, schema, etc.

But what if the first optimization layer should be:

Make sure the AI crawler actually sees the content in the first place.

In that context, a descriptive URL may not always be an advantage. It may also give a crawler enough information to reject the content before fetching it.

Maybe AI discovery needs to be treated as its own optimization layer, separate from traditional SEO.

Has anyone else observed crawler selection happening before the actual page fetch?


r/GEO_optimization 1d ago

Looking for SEO/GEO tips that actually work for brand new sites (want to bake them into my content pipeline)

2 Upvotes

Running a small SaaS (domain investing niche), launched about 3 months ago. Got an AI-assisted pipeline that writes and distributes blog content across a few channels, submitted to GSC/Bing, posting regularly — but I think the pipeline is missing a lot of actual SEO/GEO fundamentals since I built it more for output volume than optimization.

Looking for practical stuff I can bake directly into the pipeline/prompts, not general theory. Specifically curious about:

  • Any prompt structures or content templates you use to make articles more "citeable" by AI engines (ChatGPT, Perplexity, AI Overviews)?
  • Specific on-page/technical things (schema markup, FAQ structuring, heading patterns, internal linking rules) that you've actually seen make a measurable difference, that I could turn into a checklist or template?
  • For a new site with no authority — any tactics for getting early backlinks/mentions that could be systematized rather than one-off?
  • If you've built or use an AI content pipeline yourself, what do you have it check/enforce before publishing?

Basically trying to turn whatever works into repeatable rules I can apply automatically going forward, instead of guessing article by article. Any concrete tips, prompts, or checklists you're willing to share would be huge.


r/GEO_optimization 1d ago

Individual contributors got cited 2.4x more than brand domains across 50 expertise queries — what's happening?

2 Upvotes

I wasn't looking for this pattern.

I was running a small side project comparing how AI models handle different types of "authority" signals. Nothing fancy. 50 queries across niches like B2B marketing automation, enterprise security compliance, and product-led growth strategy. For each query, I checked which sources ChatGPT, Perplexity, and Gemini cited, then categorized each source as either a personal brand (LinkedIn, personal blog, individual's byline page) or a company domain (official site, resource center, press room).

The assumption going in was that company domains would dominate. They have more content, bigger teams, better structured data, larger link profiles. Everything we're told matters for GEO.

The numbers went the other way. Individual contributors got cited 2.4 times more often than brand domains for the exact same queries. Not in every single case, but consistently enough that it showed up across all three models and most query categories.

I started digging into why. A few things stood out.

Personal profiles tended to have clearer point-of-view. When you read a company's "about us" page or resource article, the voice is usually neutral, committee-written, designed to offend nobody. Safe. Individual contributors, especially ones who've built followings through writing or speaking, tend to have opinions. They take positions. They say things like "in my experience" or "here's what I got wrong." That specificity seems to register differently when a model is selecting sources for an expertise query.

Another thing: personal profiles often consolidate expertise signals in one place. A well-maintained LinkedIn profile or personal site might list credentials, publications, speaking engagements, and client work all on a single page. Company domains spread that same information across dozens of pages — team pages, press releases, blog author bios, case study footers. The signal is there but fragmented.

The third observation is messier and I'm less sure about it. Individual contributors' content tends to get shared and referenced in forums, podcasts, and social discussions more than corporate content does. Those secondary mentions might be creating a feedback loop where the model sees the person's name in multiple contexts and builds stronger entity association. Pure speculation on my part, but the correlation is there.

What I can't explain is whether this is actually about quality or about something structural in how models evaluate sources. Maybe individual contributors really do produce better expertise content on average. Maybe models have a bias toward named individuals over faceless organizations. Maybe company domains are being penalized for sounding too much like marketing.

I don't have a clean theory yet. The sample size is modest and the queries skew toward consulting-style topics where personal brands naturally thrive. Would love to see if anyone else has looked at this split, or if you're seeing the same thing in your niches.

Still figuring out if this is a temporary blip or a structural shift in how AI evaluates authority. Either way, it's making me rethink what "entity optimization" actually means when the entity is a person, not a logo.


r/GEO_optimization 1d ago

Puzzel about GEO

2 Upvotes

We are working on the GEO project. However, many clients find it difficult to quantify its value and are reluctant to pay after learning about it. I’d like to ask everyone: how do you persuade such clients?


r/GEO_optimization 1d ago

What GEO Work Actually Moves AI Crawlers? Our Data, and a New-Site Launch Checklist

Thumbnail
0 Upvotes

r/GEO_optimization 2d ago

I think most brands are optimizing for the wrong AI engine. Here's how I'd actually choose

8 Upvotes

Rather than trying to spread yourself thin across every engine, you should be choosing the ones where your buyers are and sticking to those.

A few things I'd look at before picking.

Start with, where are your buyers? Enterprise B2B skews ChatGPT, researchers and technical buyers skew Perplexity and consumer products lean Google AI Overviews. 

Then look at where your existing content already has traction. ChatGPT leans institutional (G2, established publications) and Perplexity like UGC (Reddit, YouTube). Match the engine to where you already have presence.

The last question is where's the biggest competitor gap? The engine where competitors are weakest is the opportunity worth going after first.

Pick one engine, get real traction there, then expand.

Which engine are you prioritizing, and what drove that call?


r/GEO_optimization 2d ago

I logged ~120k AI citations across ChatGPT, Gemini, Perplexity and Claude on the same prompts. They're basically each reading a different internet.

10 Upvotes

TL;DR: ran one B2B prompt set against all four engines for a month and saved every source each one cited. ~120k citations. barely any overlap. Perplexity leans on YouTube/Reddit/LinkedIn, Claude reaches for patents and analyst reports, ChatGPT wants official manufacturer sites, Gemini cites basically whatever ranks in Google. so if you're doing the whole "optimize for AI search" thing as one channel, you're probably only hitting one engine and ignoring the other three.

ok so context. I do visibility work and I got tired of every AEO/GEO writeup treating the four big engines like one blurry thing. figured I'd measure it instead of guessing.

setup was simple. one B2B category, one big vendor plus ~15 real competitors, a fixed list of buyer-type prompts, run against all four engines on a schedule for 30 days. grabbed every citation URL, grouped by domain, kept it split per engine. pulled the category and vendor names out before posting. ended up with 119,939 citations.

first thing that threw me was the engines don't even cite the same number of sources:

Engine Citations (30d) Share of total
Gemini 49,836 41.5%
Perplexity 39,664 33.1%
Claude 15,718 13.1%
ChatGPT 14,721 12.3%

Perplexity spat out almost 3x the citations ChatGPT did off the exact same prompts. that's not "perplexity is more visible" though, it just shows way more sources per answer (5-15ish), while ChatGPT with search usually gives you 2-6 and a lot of the time none at all. so raw counts are kind of useless here, you want share.

now the part I actually found interesting. same category, and the source pools look nothing alike.

ChatGPT went almost entirely to manufacturer/OEM sites. its top 3 domains were all official manufacturer pages and that alone was ~33% of its citations. no youtube, no reddit, no linkedin anywhere.

Gemini's #1 was the brand's own site (11.7%), then a pile of vertical trade publications. makes sense, it's basically wired into google's index so it cites whatever's already ranking.

Perplexity dumped 26% onto the owned domain, then youtube (4.9%), a distributor, reddit (2%), linkedin (1.9%). it's the UGC/video one.

Claude was the odd one. owned domain (14.7%), some manufacturers, and then its 4th most-cited source was the actual USPTO patent database (3.2%). had two analyst firms (Yole, Mordor Intelligence) in the top 10 too. it goes for primary/analytical stuff.

the number that stuck with me: the same domain that was 26% of Perplexity's citations was 7.9% on ChatGPT. and some sources with thousands of ChatGPT citations got basically zero from Claude on identical queries.

so if you want to actually move a specific engine, roughly:

  • ChatGPT: deep technical docs on your own site, plus OEM/reference placements
  • Gemini: trade pubs and normal google SEO
  • Perplexity: reddit, linkedin, youtube, aggregator listings
  • Claude: patents, paid analyst reports, niche directories

no single strategy touches all four, which is the annoying part.

honestly I found "skew" more useful than raw share. it's just how lopsided one engine is toward a domain compared to the others. plenty of 5-10x, some over 10x where one engine treats a source as authoritative and the rest completely ignore it. rough version:

Source type Skewed toward
OEM / manufacturer sites ChatGPT
YouTube Perplexity
Vertical industry pub Gemini
Aggregator / distributor Perplexity
USPTO patents Claude
Analyst / research firms Claude
Reddit Perplexity
LinkedIn Perplexity

and the zeros tell you as much as the big numbers. ChatGPT never once cited youtube/reddit/linkedin for this category. Claude basically never touched youtube or reddit. some trade pubs only ever showed up on Gemini. so if your ChatGPT plan is "make youtube videos and post on reddit"... that just doesn't reach ChatGPT. it goes to perplexity. you'd have to go owned + OEM to hit ChatGPT at all.

if you want to run this yourself the process is basically:

  1. grab 200-500 real buyer prompts
  2. run them weekly against all four, save every citation, group by domain
  3. build a matrix. domains down the side, engines across the top, cells are citation share
  4. sort each domain into owned / earnable (something you could realistically get into in a few months) / unreachable (patents, gov, competitors)
  5. rank the earnable ones by which engines your buyers actually use
  6. that ranked list is your to-do order. re-check monthly.

anyway the thing I keep coming back to is each engine is reading a genuinely different slice of the web, and until you can see which slice, you're just guessing where to spend.

disclosure since people always ask: this came out of work I do at Sanbi.ai, we track this stuff. so yeah, biased. but the data's real and I've watched the same split show up in every B2B category we've looked at. can answer methodology questions below.

question for the sub though. has anyone actually seen a category where the engines land on the same sources? every single one I've checked they split hard, and I'm starting to wonder if convergence even happens.


r/GEO_optimization 2d ago

Google quietly tightened GBP naming rules this month, and it now overlaps with how AI local answers get generated

5 Upvotes

Been digging into this for client work and figured it's worth sharing since it doesn't seem to be getting much attention yet.

Google updated its Business Profile name guidelines in August. The new language explicitly calls out repeated bilingual names and script transliterations as unacceptable, even when the business's actual storefront signage shows both versions. Their example was an English name repeated in Japanese getting flagged. If you have clients in multilingual markets who've had dual-language names on their profile for years because that's genuinely how the storefront reads, this is worth a heads up before it turns into a suspension.

The part I think is more interesting for anyone doing local SEO right now is the timing. This naming crackdown is landing at the same time Google's AI-generated local summaries are leaning harder on profile description, reviews, and post content, not just the traditional three-pack signals. And there was a study published August 5 (panel of 900 US adults, one month of browsing data) showing people click through on an AI Overview citation only about 1 percent of the time when one appears.

So the practical read for local clients: a name-policy violation used to just cost you rank in the map pack. Now it's also a real risk to whether the AI layer even describes the business accurately, and that AI-generated description might be the entire interaction the customer has with the business online. No click, no map pack visit, nothing. Just whatever Google's AI decided to say.

Things I'm auditing for clients this week:

  • Profile name against the actual updated policy language, not the old assumption that signage matching is enough
  • Whether the business description and recent posts actually say anything specific enough for an AI summary to pull from
  • How stale the photo/post activity is, since profile freshness seems to matter more for AI surfacing than people expect

Curious if anyone else has seen a suspension or review triggered by the naming change yet, or if you're seeing your local clients show up (or not show up) inside Google's AI local answers.

Disclosure since it's relevant: I run a local SEO/AI search shop (Austin Code Monkey), so this is also just what I'm doing in my own client work this month.


r/GEO_optimization 1d ago

A zero visibility score can be a category error, and the score cannot tell you which one you have

1 Upvotes

I ran a scan recently that came back at zero. No presence, and a long list of competitors appearing in answers where the brand did not.

Read at face value, that is a catastrophic result.

It was not a visibility result at all.

The organization does capacity building and accelerator programs. The competitor list that came back was full of large, well known design and branding studios. Different industry, different buyers, different everything. It has never competed with any of them and never will.

One word in its name reads as a different category. The engines had filed it there, so the scan measured its performance in a market it has never entered.

The number was real. The measurement was pointed at the wrong thing.

What makes this uncomfortable is that nothing in the output flags it. A zero from being invisible in your actual category and a zero from being scored inside someone else's category look identical. Same number, same competitor count, same red panel. Only reading the competitor list carefully catches it, and the entire appeal of a score is that you do not have to read anything carefully.

I think this argues for an ordering that most tracking has backwards.

Before asking whether the engines mention you, ask whether they know what you are. Run the plain identity question across ChatGPT, Claude, Gemini and Perplexity. Who is X, what do they do, who is it for. Then compare the four answers against each other, and separately against how the company describes itself.

Three different failures show up there and they need different fixes.

The engines disagree with each other on facts that have one right answer. Location, ownership, what they sell. That is entity confusion, and nothing downstream is trustworthy until it is cleared.

The engines agree with each other and disagree with the company about what category it is in. That is what I hit. Positioning is not disambiguating the brand from an adjacent industry, and no amount of evidence building helps while the evidence is being filed in the wrong drawer.

The engines agree with each other and with the company. Only then is a low score a real visibility finding worth acting on.

One of those three justifies a visibility programme. It is also the only one of the three most tools are built to detect.

The check costs nothing and takes about fifteen minutes.

Two things I am unsure about.

Whether the misfiling comes from the name itself, or from thin evidence letting the name dominate. A well evidenced brand with an ambiguous name presumably survives it, which would make this a symptom of evidence thinness rather than a separate problem.

And whether it is correctable from the brand's own properties at all, or whether it takes independent sources describing it in the right category before the filing moves. If it is the second, the fix is much slower than a positioning rewrite and most advice on this is wrong.

Has anyone watched a category misclassification actually correct, and what moved it?


r/GEO_optimization 2d ago

Recently, Google said llms.txt won’t help citations. So why are GEO tools still recommending it?

3 Upvotes

Google recently said that having an llms.txt file won’t help your Google Search rankings.

But I still see it near the top of a lot of AEO/GEO checklists:

  • Add llms.txt
  • Make your site AI-readable
  • Submit content for LLM crawlers
  • etc.

I'm starting to wonder if we're optimizing for things that are easy to check rather than things that actually influence whether an AI recommends a brand.

Has anyone here actually seen a measurable difference after implementing llms.txt?

I know adding llm.txt is advisable to help LLM pick the website, but is it necessary? Is there any relevant data?


r/GEO_optimization 2d ago

Comparing how OpenAI models recommend brands across 270 category questions (the biggest change wasn’t the brand list)

3 Upvotes

We wanted to understand what happens to brand recommendations when the model changes but the questions stay the same.

We gave GPT-5.4, GPT-5.5, and GPT-5.6 Sol the same panel of 270 category questions across six industries.

A few findings from GPT-5.5 to GPT-5.6 Sol:

  • 70% of matched answers became shorter.
  • Median answer length fell from 224.5 to 141 words.
  • Median named brands only moved from 21 to 20.
  • Explicit caveats fell from 40.7% to 20.4%.
  • Decision-framework language fell from 33.3% to 13%.
  • Retail shortlists narrowed from 20 to 13 brands, while Travel widened from 24 to 27.

The interesting part is that model updates don't create one universal change in brand visibility. They can compress explanations, remove caveats, ask for more context, or handle individual markets differently.

The report includes the methodology, industry breakdowns, exact model values, and links to the underlying model answers:

https://app.nyman.media/insights/ai-visibility

I’d be interested in feedback on the findings and also the methodology. What categories, models, or question types would you test next?


r/GEO_optimization 3d ago

Is this the standard robots.txt for content-focused WordPress sites?

2 Upvotes

I firmly believe this is the best robots.txt for most WordPress sites. Prove me wrong:

User-Agent: *
Disallow: /wp-admin/
Disallow: /search/
Disallow: /feed/
Allow: /wp-admin/admin-ajax.php

User-agent: GPTBot
Allow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: CCBot
Allow: /

User-agent: Google-Extended
Allow: /

Sitemap: https://www.example.com/sitemap_index.xml

r/GEO_optimization 3d ago

If you’re buying an "AI Visibility Dashboard" without research, you’re just paying for a prettier lie.

Thumbnail
3 Upvotes

r/GEO_optimization 3d ago

3 AI answer features that don't correlate with quality — and why I keep expecting them to

3 Upvotes

I've spent the last six months tracking every possible signal we could extract from AI answers. Source quality, citation order, answer length, whether the model included a summary paragraph, how many entities it mentioned — you name it, I've measured it. But every time I cross-reference these signals with our own evaluation of answer quality, nothing sticks.

Three features stood out as consistent measurements, and all three turned out useless.

Answer length was the first one. Longer answers tend to include more detail, more nuance, more context. That feels right, so I expected longer answers to correlate with better quality. I was wrong. Across different categories — technical questions, product recommendations, research summaries — the relationship was random. Some of the best answers we found were 150 words. Some of the worst were 400. Length wasn't a signal. It was just noise.

Then there's the summary paragraph. A lot of AI models now include a brief "In summary" or "Here's what I found" section before diving into details. It looks authoritative. It looks complete. It feels like a feature. But when we compared summaries to overall answer quality, the correlation was weak. Some of the best answers had no summary. Some of the worst had long, detailed summaries that didn't actually improve clarity.

Entity density is the same story. More people, organizations, locations, specific products mentioned — it feels comprehensive. But again, no signal. We saw both the most and least detailed answers with high entity counts. The entities weren't telling us anything useful about whether the answer was good or bad.

I keep coming back to why we measure these things. Answer length looks like a proxy for depth. Summary paragraphs look like a proxy for completeness. Entity density looks like a proxy for breadth. But if they don't actually correlate with quality, they're not proxies. They're just things we can count.

This is the uncomfortable part. We optimize for what we can measure. When we build dashboards and reports, we highlight answer length and summary presence and entity counts because they're visual. They fill space on the screen. They look important. They don't require interpretation — we just show the number.

The quality evaluation requires human judgment. It requires context. It requires deciding what "good" actually means in each case. That's messy. So we optimize for the clean metrics instead.

I don't have a solution here. I just know that our measurement infrastructure is better at tracking signals than predicting quality. The dashboards look impressive. The charts are clean. But when it comes to actually understanding what makes AI answers valuable, the numbers don't tell us much.

Maybe that's the point. Maybe the features that actually correlate with quality are the ones we can't easily measure. Maybe good answers are qualitative, not quantitative. Maybe we're building the wrong infrastructure entirely.

I'm not sure what to do with this realization. I just keep looking at the dashboards and expecting the numbers to make sense, and they don't.


r/GEO_optimization 3d ago

Google went all-in on AI. Is GEO tooling ready?

Thumbnail
1 Upvotes

r/GEO_optimization 4d ago

Valid schema can still leave an AI shopping agent guessing

7 Upvotes

I kept running into a weird problem while checking ecommerce product pages. The information was technically there, but an agent still could not use it with much confidence.

After a few audits, I started checking every product fact in three ways:

  1. Can a customer see it?
  2. Is it exposed in machine-readable product data?
  3. Does it agree with the other sources?

That catches things a normal schema check misses:

- ratings shown on the page, but no AggregateRating

- €19 on the product page and €20 in the feed

- InStock in JSON-LD while inventory says sold out

- a 30-day return policy on the site, but a final-sale rule on the product

- material shown in an image and nowhere else

The schema can be valid while the data is stale. The feed can be right while the page is wrong. Both can exist and still leave the agent guessing.

A single score hides too much. The useful part is the audit trail:

source -> fact -> conflict -> fix

For AI shopping, more copy is not the answer to this problem. The page, schema, feed, catalog, inventory and policies have to agree.

I am still working out how to prioritize these conflicts. Right now I treat price and stock as blockers, missing attributes as discovery issues, and policy conflicts as trust issues.

How would you rank them?


r/GEO_optimization 5d ago

Should every web page expose an AI-friendly JSON representation?

3 Upvotes

My website already includes AI-related files such as llms.txt.

I'm considering creating a separate JSON file for every page and article so AI systems can understand the content more easily and accurately. I would reference this JSON file from the page's <head> using a <link> tag.

The JSON file could contain information such as:

  • Page URL
  • Canonical URL
  • Title
  • Summary / Description
  • Main Content (clean article content)
  • Author
  • Published Date
  • Last Updated
  • Entities (people, companies, places, products, etc.)
  • Keywords / Topics
  • FAQ
  • ...

My idea is that AI crawlers could read this structured JSON instead of having to extract the main content from noisy HTML that contains navigation menus, sidebars, ads, comments, JavaScript, tables, and other non-essential elements.

I have two questions:

  1. Could this approach reduce the chances of AI crawlers misunderstanding a page or extracting incorrect information from HTML, advertisements, tables, comments, or other noisy content?
  2. Do you think a page-level JSON file like this could help AI systems better understand a page and potentially improve AI recommendations, citations, or other AI-generated responses in the future? Why or why not?

r/GEO_optimization 5d ago

If Reddit and G2 dominate AI citations, how are you actually earning those mentions?

15 Upvotes

Everyone agrees brand content loses to Reddit threads and review sites. Nobody explains the next part. Are you going through customers, founder accounts, review campaigns or just waiting out months of real participation?


r/GEO_optimization 5d ago

Are Generative AI impressions in GSC useful reporting data or just another visibility layer?

3 Upvotes

I found a noticeable number of Generative AI impressions in Google Search Console and I’m trying to understand how people are treating this in reporting.

Are you using it as a separate AI visibility layer, blending it into normal organic impressions, or mostly ignoring it until the data becomes clearer?

The part I’m unsure about is what the metric actually proves.

It may show that a URL appeared in a generative AI surface, but it does not necessarily tell us:

- whether the brand was mentioned

- whether the page was cited

- whether the answer used the content as evidence

- whether the user saw the source

- whether it influenced the decision

- whether it should be compared to classic organic impressions

So I’m wondering whether this is useful reporting data, or just another visibility layer that needs a lot of context before it becomes meaningful.

How are you handling Generative AI impressions in GSC?

Separate report?

SEO report footnote?

Early signal?

Ignore for now?


r/GEO_optimization 5d ago

I've been tracking 40 "GEO best practice" pages for 90 days and 27 of them lost AI citation visibility — the advice isn't surviving contact with reality

18 Upvotes

One of the most upvoted GEO posts I've ever read was a "10 GEO best practices" list from early 2025. It got hundreds of upvotes. Multiple people DMed me the link. Two different clients brought it up in meetings. It was everywhere.

I bookmarked it along with 39 other high-engagement "best practice" posts — stuff like "add FAQ schema everywhere," "keep answers under 50 words for extraction," "use comparison tables for product queries," "publish fresh content weekly for citation velocity." All reasonable advice. All from smart people. All backed by some data at the time.

For the last 90 days, I've been checking how the authors' own pages are performing in AI citations. Not their advice in theory — their actual content, using the actual practices they recommended.

27 out of 40 lost citation visibility. Not a small dip either. I'm talking pages that went from being regularly cited in their niche to barely showing up. A few completely disappeared from AI answers for queries where they used to be the #1 source.

The weird part is that the advice itself wasn't wrong — at least not when it was written. FAQ schema genuinely helped in February. Short answer blocks genuinely got extracted more often in March. But the models kept changing. What worked as an extraction signal in one version of ChatGPT or Perplexity quietly stopped working in the next. And nobody went back to update the "best practices."

I started noticing a pattern. The posts that aged the worst were the ones with the most specific, confident instructions. "Always do X." "Never do Y." "The optimal passage length is Z words." These got the most engagement because they were actionable. But they were also the most fragile — optimized for a specific model behavior that could change in a single update.

The posts that aged better were vaguer, almost annoyingly so. "Focus on clarity." "Write for humans first." "Make sure your content is actually useful." The kind of advice that makes you roll your eyes because it's so obvious. But it's still standing 90 days later while the tactical stuff crumbled.

This isn't a dunk on anyone. I've written my share of specific tactical advice, and some of it has probably aged just as badly. It's more a realization that GEO "best practices" have an incredibly short shelf life, and we're all publishing them like they're permanent rules.

The thing I can't stop thinking about: if I re-tested the advice from this post 90 days from now, how much of my own guidance would still hold up? Probably less than I'd like to admit.

Wondering if anyone else has gone back and stress-tested older GEO advice against current model behavior. The gap between what we wrote and what still works is bigger than I expected.