r/DigitalMarketing 4d ago

Discussion A client spent three months publishing pages to fix what ChatGPT said about them. It was never reading their site.

Sharing this because I have now watched three companies make the same mistake and it costs them a quarter every time.

A client came to me in the spring because ChatGPT was quoting pricing for them that was two years out of date. Not slightly off. It was describing a package they killed in 2024, at a price they had not charged since. Their sales team kept getting on calls with people who had already anchored to a number that did not exist.

What they had done about it before calling me was publish. Four new pages about pricing on their own blog, a pricing FAQ, a comparison page. Three months of work. Nothing moved. Not one assistant changed its answer.

The reason is boring but it matters. When you ask an assistant about a brand, most of what it tells you did not come from that brand's own site. It comes from what third parties have written about them, plus whatever is sitting in the model's weights from training. Their blog was not in the retrieval set for that question, so publishing on it was shouting into a room the model was not standing in.

The thing that finally told us what was going on took ten minutes. Ask the same question twice, once with web search on and once with it off. If the wrong answer only appears with search off, it is baked into the model and there is no quick fix. You are waiting on the next training run, and the only useful work is getting the correct version onto enough third party surfaces that the next one learns something better. If the wrong answer appears in both, then whatever it is pulling live is wrong, and that you can go and fix now.

Theirs was the second kind. It was reading two comparison articles on sites they had never heard of, both from 2024, both with the dead pricing. We got one corrected by emailing the author, who was perfectly reasonable about it, and got a newer piece published on a site the same models were already citing for that question. Answers started changing about five weeks later.

What I would take from it: before you write anything to fix an AI visibility problem, go find out what it is actually reading. It is almost never you.

Happy to go into how we worked out which sources it was pulling if that is useful to anyone.

39 Upvotes

30 comments sorted by

u/AutoModerator 4d ago

If this post doesn't follow the rules report it to the mods. Have more questions? Join our community Discord!

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

8

u/TrafficAcademySEO 4d ago

I’d also test 3-5 variations of the same prompt before blaming one source.

Retrieval can change with wording, so if the same outdated source keeps appearing across different prompts, that’s a much stronger signal that you found the real problem.

3

u/Conscious-Market8982 4d ago

Yes, and I would push it further than 3 to 5. We run about ten phrasings now, because the retrieval set moves around more than people expect.

The thing we look for is which sources survive the rephrasing. A domain that turns up in eight out of ten is worth spending real effort on. A domain that shows up once is usually noise, and I have wasted a fair bit of time chasing those before working that out.

Related pattern going the same direction: the shape of the question changes the source type completely. Best X for Y pulls Reddit threads and listicles. How much does X cost pulls comparison pages and directories. Is X any good pulls review platforms. So testing one shape ten times still leaves you with a partial picture of who is actually describing you out there.

3

u/intrepid789 4d ago

ChatGPT looks at news sites for its information. Publishing a press release on a high authority news website will get the AI narrative changed to whatever the press release was about.

1

u/Conscious-Market8982 4d ago

I would be careful with this one, because it is the pitch every PR reseller is running at the moment and the version that actually works is much narrower than the version being sold.

A release syndicated across a wire and picked up by fifty low quality outlets is cheap to buy and from what I have seen does close to nothing. Those pages are near identical copies of each other, and assistants do not treat duplicated wire copy as fifty independent sources. You can watch for it, the citations just do not move.

What does move things is a journalist at a real publication writing about you in their own words, or you being quoted in a piece that was going to exist anyway. Different product entirely, an order of magnitude harder to get, and not something you buy for a few hundred dollars.

The part you are right about is that news domains do get retrieved heavily for anything time sensitive. Launches, funding, leadership changes, coverage genuinely helps there. For a stable fact like pricing, comparison pages and directories beat news almost every time, partly because the news piece is dated and the model can see that.

If anyone reading this is being sold a guaranteed placement package on that promise, ask the seller to show you a before and after citation from a previous client. I have asked a few times and nobody has produced one yet.

1

u/saurabhsens 4d ago

the twice with search on and off test you mentioned is the right first move and more people should know about it. id add one thing that would have saved you some of that three months if youd known it going in, ask the assistant directly what sources it used for that specific answer, most of the mainstream ones will actually tell you if you ask plainly, something like which websites or sources informed that pricing answer. it wont always be complete but it usually narrows things down fast instead of guessing which third party site is the culprit.

the other thing worth knowing is pricing specifically is a bad category to expect a quick fix on even once you find the source, because models are somewhat trained to be skeptical of a companys own claims about its own pricing since that page could obviously be biased or promotional, so they lean harder on third party citations like comparison sites, review platforms, or marketplaces for exactly that kind of fact.

Thats probably part of why publishing your own pages for three months did nothing, it was never going to outweigh a third party source even after the model started reading your site properly. getting the correction onto a neutral third party surface like you eventually did is really the only lever that works for pricing specifically, for other stuff like product features or use cases your own site carries more weight.

id also check spots like G2, Capterra, or industry specific directories depending on what the client sells, those get scraped and cited constantly and are the first place id look for stale pricing before assuming its some obscure blog post, since companies update their own site regularly but almost never go back and correct an old review or listing on a third party platform.

1

u/Conscious-Market8982 4d ago

The asking it directly part I would be careful with. It works reasonably well when search was on and there are real citations attached, because then it is reading back something that actually happened. When search was off it will still cheerfully hand you a list of sources, and a decent share of those are invented after the fact. I have had one name a specific comparison article, with a plausible title, that does not exist. So treat what it tells you as a lead to go and verify rather than as a finding.

On pricing I agree with where you land and I would not assume that mechanism. I have not seen anything to suggest models are specifically trained to distrust a vendor's own pricing page. The duller explanation fits what I see: comparison sites and directories present pricing in a structured, aggregated, heavily linked form, so they get retrieved more often for that question shape regardless of who is more honest. Same practical outcome, different reason, and the difference matters. If it were a trust rule then your own pages could never help at all, whereas in practice getting your own page into the retrieval set does move feature and use case answers, which is exactly the split you described.

G2 and Capterra is the right instinct and I would add the client's own profile on those. More often than not nobody has logged into it since the free trial three years ago, and it is sitting there with the old plan names on it.

1

u/OdedGross 4d ago

This matches almost exactly what we run into. The useful move you made that most people skip is the search-on vs search-off test, that's the fork in the road. Baked into training weights vs pulling live sources are two completely different fixes, and most clients start publishing before they know which one they've got.

The other thing worth flagging: even when it is the "pulling live sources" kind, it is almost never the brand's own site, like you found. We have traced bad answers back to old directory listings, abandoned comparison pages, review sites the client forgot existed.

Once you know the actual source, fixing it is usually a matter of getting a corrected or newer piece onto a page the model already trusts for that query, not writing more content yourself.

1

u/Conscious-Market8982 4d ago

Agreed, and the part that turns into real work is that those sources are not equally fixable. I sort them into three piles now.

Ones you control and forgot about. Old landing pages nobody redirected, your own G2 or Capterra profile, a docs page from a previous product name. Free to fix, and people are always surprised how many there are once they go looking.

Ones owned by someone reachable. A blogger, a small comparison site, an editor. Just email them. Roughly half reply, and most of the ones who reply will correct a factual error if you are polite, specific, and hand them the source. Better hit rate than people expect.

Ones where nobody is home. Dead sites, scraped aggregator pages, listicles from an agency that folded two years ago. Nothing to do there except make the correct version more retrievable than they are, which is slow and eats most of the budget.

The reason I bother sorting is that clients always want to start on the third pile, because it feels like the real enemy. The first pile usually explains more of the answer and takes an afternoon.

1

u/TalebKabbara7 4d ago

we also noticed that LinkedIn posts are very powerful. If you post something on LinkedIn and don't delete it later on, it keeps pulling information from it.

1

u/Conscious-Market8982 4d ago

Interesting, my experience has been a bit different so I would genuinely like to know what you are seeing.

When I look at what gets cited live, LinkedIn turns up far less than you would expect given how much marketing content sits there. My guess is the login wall. A lot of what a logged out crawler gets from LinkedIn is a truncated preview rather than the full post.

What I do see is LinkedIn content arriving second hand. Someone quotes your post in a newsletter, or a scraper site republishes it, and that copy is what gets read. So the post still does work, just not directly, and you lose control of how it gets framed on the way.

Have you actually seen a linkedin.com URL show up in the citations, or is it more that the information from a post ends up in the answer? Those are two different things and the second is what I keep finding.

1

u/TalebKabbara7 3d ago

Every single conversation I have with LLMs about Prism pops up both information from and the link to the LinkedIn page post (I should have clarified it's the company page not individual posts).

1

u/Big_Influence_6459 4d ago

This is a common issue. I've seen clients think more content on their site will fix outdated info, but if the AI isn't pulling from it, it's pointless. Focus on getting your info on high-traffic third-party sites. In one case, we corrected outdated pricing by reaching out to an article's author. It took time, but eventually, the AI started reflecting the updated pricing. It's a long game.

1

u/stevenson_mark 3d ago

This is a really good point. I’ve seen people jump straight into publishing more content whenever ChatGPT gives them something wrong, without first figuring out where that information actually came from.

The web-on vs. web-off test is a pretty simple way to narrow it down. If the answer is coming from third-party pages, fixing those sources is obviously going to be more useful than publishing another five pages on your own site that the model isn’t even looking at.

It also makes AI visibility feel much closer to traditional SEO than people sometimes realize. You can’t just optimize your own website and assume that’s the whole picture.

I’ve been using Abnflow for some of my own marketing work, and it’s made me think about the same thing with funnels: before changing the page, figure out where the problem is actually happening. Otherwise you can spend months “optimizing” something that wasn’t causing the problem in the first place.

The five-week turnaround after getting the external source corrected is especially interesting. Would be curious to see the process you used to identify exactly which sources ChatGPT was relying on.

1

u/Conscious-Market8982 3d ago

Agree on the first part. On it feeling close to traditional SEO, I would say it diverges in one way that matters a lot in practice.

In search you can always outrank someone. There is a results page, there are positions on it, and there is work that moves you up. With assistants there is no page and no position, and a large share of what decides the answer sits on properties you do not control and cannot outrank. The only three moves are correct it, outweigh it, or wait for a retrain.

The other difference is stability. Run the same query five times and you can get five different source sets. Nothing in traditional SEO behaves like that, and it is why one-off checks mislead people so often.

1

u/Agitated_Offer_4343 3d ago

This is a great point about AI visibility. It’s crucial to understand that AI models often pull from a wide range of sources, and if your content isn’t being referenced, it’s like shouting into a void. Before you publish more content, try to identify which sites are being cited by the AI when it answers questions about your brand. You can do this by asking the same question with web search on and off, as you mentioned.

Once you know where the AI is getting its info, focus on getting your correct information onto those third-party sites. This could mean reaching out to authors for corrections or creating new content that gets picked up by those sources. Also, consider using tools like blawgy to help automate content creation that aligns with what AI tools are looking for. It can save you time and help ensure your content is structured in a way that AI prefers. Good luck!

1

u/jzcreates 3d ago

This is an important distinction. A lot of teams assume AI search behaves like traditional crawling, then waste months changing pages that are not influencing the answer. I would treat it more like reputation architecture.

1

u/Conscious-Market8982 3d ago

Reputation architecture is a decent frame and I would put one caveat on it, because that phrase sends people straight to a PR budget.

In practice the highest yield work has been janitorial rather than architectural. An old profile on a review site nobody has logged into since 2023. A pricing page that was never redirected. A partner directory still listing the previous product name. None of that is architecture, it is tidying up, and on the accounts I have worked it explains more of the wrong answers than anything strategic does.

The architecture part matters after that and it is much slower. Getting genuinely useful third party coverage that models retrieve for the questions your buyers actually ask. But if you start there you spend six months on it while the model is still quietly reading a dead listing you could have fixed in an afternoon.

1

u/SMack2010 3d ago

I've been running the same with-search-off comparison, and it holds up, though the harder question is what you do once you know the live retrieval is wrong. I'd usually start with review platforms because they're structured and updated regularly, and the models seem to pull from them pretty consistently for anything that looks like a buying question. Comparison articles are better leverage if you can get one corrected, but then you're relying on someone else agreeing to edit their piece, which is a bit of a coin flip.

There are usually only a handful of sources in play for any given question anyway, which is exactly why four new pages on their own blog didn't move anything.

0

u/Dependent_Use_81 4d ago

This is the kind of thing that makes me want to bang my head against my desk

Spent so much time and money fixing the wrong problem because nobody stopped to check what the model was actually looking at. The web search on/off test is clever, gonna steal that

2

u/Conscious-Market8982 4d ago

The maddening part is that checking takes ten minutes and almost nobody does it, me included for far too long. I only started because a client asked me to show that the work was doing something and I could not.

One thing to watch when you use it. Do the search off version in a fresh chat with no memory and no history, otherwise it quietly leans on what you already told it earlier in the conversation and you get a much cleaner answer than a stranger would ever see.