---
title: "The Week an AI Assistant Started Describing Our Client's Product Wrong and We Had No Way To Tell"
url: "https://techmagazine.io/insight/the-week-an-ai-assistant-started-describing-our-clients-product-wrong-and-we-had-no-way-to-tell/"
author: "Kartik Chugh"
published: "2026-09-25"
updated: "2026-09-25"
---

# The Week an AI Assistant Started Describing Our Client's Product Wrong and We Had No Way To Tell

In Q1 2026 a client forwarded us a screenshot. A prospect had asked ChatGPT what their product did, and the answer was confidently wrong about a core capability, describing a limitation the product had removed 14 months earlier. The client's question was reasonable and we could not answer it: how long has it been saying that, and how many buyers have seen it?

We had no idea. We had no instrument that could have told us, and neither did they.

### What we could and could not see

The client's marketing site was current. Their documentation was current. Their pricing page was current. Everything a person would read said the right thing, and had for over 14 months.

What the assistant appeared to be drawing on was older material that still existed in the world: a comparison article from a third-party site published before the change, a forum thread, and a review page that had never been updated. None of those were under the client's control. All of them were indexed, plausible, and wrong.

The uncomfortable realisation was that we had spent 3 years building the discipline of keeping owned surfaces accurate, and the surface that had just misinformed a buyer was assembled from sources we did not own and had never inventoried.

### What we did about it

We started by making the problem measurable, because a complaint you cannot count is a complaint you cannot manage.

We built a fixed prompt set of 30 questions a real buyer might ask, covering capability, pricing shape, integrations and comparisons. We ran it across ChatGPT, Perplexity and Google AI Overviews, and recorded every factual claim made about the client, whether it was accurate, and where it appeared to come from when a source was cited.

The first run was sobering. Roughly a fifth of the factual claims about the client were wrong or badly out of date. Almost all of the wrong ones traced back to third-party pages published before a product change. The client's own current material was rarely the source, because it was written as marketing copy and marketing copy does not state checkable facts in liftable form.

We then ran the same 30 prompts against 4 competitors over 3 weeks, partly to check our method and partly because the client asked. Two of the four had a cleaner factual record than our client did. The one with the cleanest record was not the largest or the best funded; it was the one whose changelog was public, dated, and written in plain sentences. That is a cheap advantage and it was available to everyone in the category.

### What we did not understand at the time

We had been treating accuracy as a property of our own pages. It is not. It is a property of the corpus, and the corpus includes everything anyone has written about you, weighted by how retrievable and how quotable it is.

In retrospect the client had been paying for the wrong kind of correction for years. When a review site published something inaccurate, the standard response was to email and ask for a fix, which sometimes worked and often did not. Nobody had asked the prior question, which is which inaccurate sources actually get quoted, because until recently there was no way to know.

The second thing I got wrong was assuming recency would sort itself out. It does not, at least not quickly. A well-linked article from 2 years ago can outrank and outweigh a current page indefinitely, and the systems assembling answers have no strong prior that newer is more correct unless the newer source says something specific enough to be preferred.

### What changed

Three things, and only the first is unusual.

First, we inventoried the corpus. Over 4 weeks we listed every third-party page that made a factual claim about the client and appeared in a citation or ranked for a relevant query. There were 38. We ranked them by how often they showed up as a source, and worked the top 9. Some were corrected on request. Two were not, and for those the only remedy was to publish something more specific and more current that answered the same question better.

Second, we rewrote the client's own pages to be quotable rather than persuasive. One claim per section, a number attached, a date visible. That sounds like a style change and it is really a retrievability change: a system assembling an answer needs a span it can lift, and marketing prose does not offer one.

Third, we made the prompt run a standing monthly measurement, tracked in Notion alongside the usual reporting. It takes under an hour. It is now the first thing the client asks about, ahead of rankings, because it is the only number that describes what a buyer using an assistant will actually be told.

### What I would tell someone in this position

Start by measuring, because almost nobody has, and the first run will tell you something you did not know. A fixed prompt set, run monthly, recorded. It is cheap and it converts an anxiety into a list.

The second thing is to accept that your correction surface is larger than your website. The pages that misinform your buyers are frequently pages you do not control, and the leverage is not in emailing all of them. It is in working out which ones are actually being quoted, which is a much shorter list than the set of pages that are wrong.

The third thing, and the one I would emphasise to anyone with a product that changes: every capability change creates a decay problem. The moment you remove a limitation, every existing article describing that limitation becomes a source of future misinformation with a long tail. Treating a product change as a content event, with an explicit pass over what the world already says, is now part of how we run launches. Our approach to that measurement is in our [share of AI citations piece](https://forkoff.xyz/blog/ai-seo/measure-share-of-ai-citations).

The client's answer is accurate now, on that question, in those systems. I would not describe it as solved. I would describe it as monitored, which is the honest state and a considerable improvement on not knowing.

---

Kartik Chugh (Simba) is a founder-operator at the intersection of distribution, culture, and narrative control in Web3.

Cofounder of [FORKOFF](https://forkoff.xyz), a culture and distribution studio that designs IP-driven campaigns, event systems, and narrative loops for protocols, funds, and builder ecosystems. FORKOFF treats events as content factories, founders as distribution engines, and culture as infrastructure — not aesthetics. 3,085+ short-form clips every 13 days for clients. $5M+ in ecosystem activations across 14 countries.

Previously CMO at QuillAudits, the Web3 security pioneer, where he scaled security products to 100K+ users, built 150+ ecosystem partnerships, generated $3M+ qualified pipeline, and drove 1Bn+ views across campaigns. Co-founded EdSquare (acquired). Five years across the AI, Web3, and B2B SaaS playbook.

Hosted and partnered on 100+ global events across ETHDenver, Token2049, Consensus, Devcon, and KBW in 20+ countries. Leads Misfits Dubai, a founder-first community built around curated rooms rather than mass communities. Builder at Seedrail (the distribution stack for tech and VCs). Active investor in 12+ early-stage startups across crypto and AI.

Frequent contributor to CoinDesk, CoinTelegraph, The Defiant, and Block Telegraph. Speaker at Token2049 Singapore and QuillCon. Advisor at TiE Global and ADSME HUB.

Speaks on: founder-led distribution, events as content factories, rooms > reach, culture > campaigns, narrative control in Web3, and creator-led distribution.

Available for commentary on: AI agency growth, Web3 marketing, podcast clipping ROI, founder-led GTM, KOL marketing, and ecosystem activation strategy. Based in Dubai.
