Ayan
← All resources
Insights

Everyone Is Arguing About the Thermometer

Two companies that sell the same thing spent last week arguing in public about whether the numbers they sell are real. Both are partly right. But the fight is about the accuracy of a thermometer, and a thermometer was never the job.

Pablo CHARRON, Co-founder, AYANJul 27, 20266 min read
Everyone Is Arguing About the Thermometer

Two companies that sell the same thing spent last week arguing in public about whether the numbers they sell are real. I read the whole exchange twice. What stayed with me was not who won.

The short version, for anyone who missed it. One vendor published charts arguing that a competitor's method of asking a model the same question once a day for thirty days carries a margin of error somewhere around twelve to twenty seven points on a single question. Their conclusion: a brand that improves its score by ten points has no way of knowing whether it improved anything at all. The competitor answered with its own data, and the answer was credible. Across a portfolio of hundreds of questions and seven platforms, asking once a day and asking ten times a day produced nearly identical readings. Repetition matters when you care about one exact question. Breadth matters when you care about a category. Which questions you choose to ask moves the number far more than how often you ask them.

Both sides are partly right, which is what makes the whole thing worth reading rather than dismissing as a vendor knife fight. But buried in the rebuttal is a sentence that neither side can engineer its way out of, and it is the most honest line published in this category all month: you cannot measure your way past drift.


## The floor under the noise

Drift means the platforms quietly change. Models get retuned. Retrieval gets rewired. Ranking behaviour shifts on a Tuesday afternoon with no changelog, no notice, and no obligation to anyone. Run your measurement a thousand times a day and you will get a very precise reading of a target that moved while you were reading it.

Then there is the second floor, which is harder. A well known analyst made the point in the comments and I have not stopped thinking about it: these systems personalize heavily on even a little user history. A measurement tool asking questions from a clean, anonymous session is not seeing what your actual customer sees. It is seeing what nobody sees.

I know this one is true because I ran the experiment myself, badly and by accident, ten days ago. I asked my own logged-in account where to watch a football match in Dubai and which television to buy for it. Then I asked a clean session the same two questions. Different venues. Different television. The logged-in answer had read my own research history from June and handed it back to me with confidence. The clean session, with no memory and nobody to please, recommended a different product at a higher price and called it the better buy.

One question, one evening, one city, two answers. Now imagine a report telling you your score in that category is 43.


## What the number is actually measuring

Put those two floors together and you get an uncomfortable description of what a score in this market currently is. It is a noisy reading, taken from a session that resembles no real buyer, of a system that changes underneath you, with no demonstrated line to revenue.

I want to be careful here, because the easy move is cynicism and cynicism is wrong. Measurement is not worthless. Directionally, over a broad enough set of questions and a long enough window, these readings tell you something real: whether a category has heard of you, which competitors keep appearing beside you, what the machine believes to be true about your product. That is genuinely useful information and no brand should be flying without it.

What worries me is the shape of the spend that is forming around it. I have watched a version of this before. A category discovers a number, the number becomes a target, budgets get built on the target, agencies get hired against the target, and four years later everyone is optimizing a metric whose relationship to money nobody ever established. Search did this. A generation of work went into a rank position that was, for long stretches, a proxy for a proxy.

The AI visibility market is eighteen months old and already selling six figure annual contracts against a number with a confidence interval nobody puts on the slide. That is not a reason to stay out. It is a reason to be specific about what you are buying.


## The question underneath the question

So here is the reframe I keep landing on, and I hold it with some conviction.

The argument about accuracy is an argument about the thermometer. It is a real argument and I am glad serious people are having it in public rather than in sales decks. But a thermometer was never the job. Nobody ever got healthy by buying a better thermometer.

The job is that a second version of your brand now exists inside these systems. It was assembled from whatever the machine could read about you, and it is the version most of your future customers will meet first. It has opinions about your durability, your price position, who you compete with, and who you are for. Some of those opinions are wrong. Most brands have never read it.

Reading that version, finding where it drifts from the truth, and making sure the machine learns the accurate one in your own words: that work does not become pointless when the score wobbles twelve points. The score wobbling is a fact about the instrument. The wrong sentence about your product is a fact about your business, and it stays wrong on the days you are not measuring.

This is the difference between chasing a number and defending the revenue attached to it. If the machine tells a buyer in Riyadh that your product does not do something it plainly does, you do not need a confidence interval to know that costs you money. You need to know it is happening, you need to know why the machine believes it, and you need a way to correct it that does not depend on getting lucky with an algorithm update.


## What to actually ask a vendor

If you are being sold something in this category this quarter, three questions separate the useful from the expensive.

What is the confidence interval on this number, and will you put it on the chart. Anyone unwilling to show you the error bars is selling certainty they do not have, and certainty is the one thing this market genuinely cannot supply yet.

What does this measure that I could act on tomorrow. A movement in a score is not an action. A specific false claim about your product, traced to the source the machine trusted, is an action.

And what happens when the platform changes next month. The honest answer is that everyone's number moves and nobody gets a warning. A vendor who tells you that plainly is worth more than one who does not.


## Where I have landed

I do not think the accuracy debate resolves. I think it recedes, the way arguments about panel methodology receded once people accepted that imperfect measurement of something that matters beats perfect measurement of something that does not.

The brands that will look smart in three years are not the ones who picked the vendor with the tightest error bars. They are the ones who read what the machine says about them, noticed the sentences that were wrong, and did the unglamorous work of making the machine learn the true ones, while everyone else was arguing about the decimal places.

The number will get better. It will still be noisy, and it will still be personalized, and it will still be measuring a moving target. Meanwhile the machine is answering questions about your brand today, to people you will never meet, in words you have never read.

That part is not uncertain at all.


Your data, your brand, your control.

Your data is stored in Europe. Privacy by design, GDPR aligned. Your brand content is never used to train other AI systems. Your Brand Base, briefs, scores and board reads belong to you.

Request a demo

Keep reading