HubSpot shows up in 78 percent of problem framed prompts on ChatGPT. On Perplexity, 34 percent. Same brand, same question set, same analyst, published this year by Kevin Indig with the error bars attached: plus or minus 6 points on one, plus or minus 9 on the other. Average the two and you get 56. There is no engine anywhere in which HubSpot scores 56.
That midpoint is what an AI visibility dashboard prints. One number, a green arrow, a month on month delta. It is not a bad dashboard. It is an index with no published methodology, sold as a KPI.
The thesis
A single AI visibility score is an index. Every index that anyone allocates real money against publishes three things: its constituents, its weights, and its rebalancing rule. AI visibility vendors publish none of the three. The four engines they blend share 3.8 percent of their sources, so almost the entire value of the composite is determined by a weighting the vendor will not show you. More data does not fix this. It makes the number smoother and no more meaningful.
What the engines actually agree on
Two large studies landed in 2026 and they agree with each other, which is more than the engines do.
Writesonic analysed 161,286 prompts across ChatGPT, Gemini, Perplexity and Google AI Overviews in May and June 2026, with 70,879 of those returning citations from all four. Sources cited by all four engines: 3.8 percent. Domains appearing on exactly one engine: 72 to 73 percent. The pairwise Jaccard similarity between any two engines tops out at 0.237, for Perplexity and AI Overviews. ChatGPT against Gemini is 0.119. As the authors put it, no pair clears 0.25, so being cited on one engine does not predict being cited on another.
Wellows ran a larger set over January to June 2026: 22.7 million citations, 1.15 million questions, five engines, 531,889 questions where all five returned citations. Sources appearing on exactly one engine: 79.6 percent. Sources appearing on all five: 0.31 percent. Perplexity never touches 89.1 percent of the pages ChatGPT cites on the same question. ChatGPT is the most isolated of the five, with 76.3 percent of its citations untouched by any other engine.
Now put that next to Ahrefs, who looked at 863,000 keywords and 4 million AI Overview URLs and found that 38 percent of AI Overview citations come from pages ranking in Google’s top 10, down from 76 percent in July 2025. Read those two facts together. The strongest agreement between any two answer engines is 23.7 percent. Google’s own AI agrees with Google’s own top 10 on 38 percent. The engines disagree with each other more than the most Google-shaped engine disagrees with Google. A composite is being built out of constituents that are close to disjoint.
The weighting nobody prints
Here is the part that turns a measurement problem into a capital allocation problem.
BrightEdge put ChatGPT at 95.1 percent of all AI referral traffic in August 2026. Gemini 2.4 percent, Perplexity 1.1 percent, Claude 0.8 percent. ChatGPT is also the engine that agrees least with everything else.
So an equal weighted score across four engines gives ChatGPT 25 percent of the number. Three quarters of what a board sees is describing engines that between them send 4.9 percent of the clicks, and those engines cite a near disjoint set of sources, which means the three quarters is not even a noisy version of the one quarter. It is a different measurement of a different thing.
An audit of nine visibility platforms published on 10 September 2026 found exactly what you would expect. Denominators vary: the same underlying data gives 50 percent on an all response denominator and 83.3 percent on a brand only denominator, a 33 point swing with no change in the world. Run frequency ranges from once a day to 100 runs per prompt. Some vendors pool engines, some average them. Of nine, one discloses a weighting at all, and that one weights by Google search volume, which is a Google proxy applied to engines that are not Google.
Then add sampling. Kevin Indig’s work on prompt tracking puts within engine variance at 10 to 34 percent on identical prompts, with AirOps reaching 815,000 prompt page pairs to establish the baseline. After three runs in ChatGPT, 2.2 percent of citations remain. High versus low reasoning settings move the citation rate 18 points. Five runs per prompt per platform per week is the floor for a confidence interval worth printing.
Stack it up. Four constituents that share 3.8 percent of their contents, each measured with a sampling band of 10 to 34 percent, combined with an undisclosed weight and an undisclosed denominator, reported to one decimal place with a month on month arrow.
Where this breaks
Four places, and two of them cut at my own argument.
The brand layer is not the page layer, and this is the strongest objection. Wellows found engines name the same companies 30.3 percent of the time while citing the same exact page only 6.8 percent of the time. Agreement at the entity level is roughly four times agreement at the document level. If your score counts brand mentions rather than citations, the disjointness I have just spent this section on shrinks by a factor of four. It does not vanish, 30.3 percent is still low, but the argument is weaker there and I am not going to pretend otherwise.
Referral traffic is the wrong denominator for reading. BrightEdge counts clicks. A large share of AI answers produce no click at all, which is the whole zero click argument, so 95.1 percent of referrals is not 95.1 percent of impressions. If Gemini answers ten times as often as it refers, my weighting arithmetic overstates ChatGPT. I have not found a credible cross engine impression share for 2026 and I would not build the number without one.
A consistent bad index still has a usable first derivative. If the vendor never changes the method, the level is meaningless but the trend may be real. That is a fair defence and it holds right up until the vendor ships a model update, adds an engine, or silently changes the denominator, none of which they announce.
And against my own side: an operator who has never written down the prompt set has no standing to complain about the vendor’s weighting. Most of the teams I see have not specified the 30 questions their buyers actually ask. Until you have, the vendor’s defaults are your fault, not theirs.
What I would do on Monday
Kill the composite in the board pack. Report per engine or report nothing. One row per engine, with the run count next to it.
Pick the one engine that matters for your buyer and say out loud why. For most B2B in 2026 that is ChatGPT on referral evidence. Optimise against that one, measure the others quarterly for drift.
Ask the vendor three questions in writing before renewal: what is the denominator, how many runs per prompt per week, and how are engines weighted. If any answer is missing, the number is not a KPI.
Write the prompt set yourself. Thirty questions your buyers ask in their own words, versioned in a file, changed deliberately and dated. That file is the asset, not the dashboard.
Report every figure with a band. 78 plus or minus 6 on one named engine is an honest number. A single 56 with a green arrow is not.
The close
I have spent thirty years building software and a good part of the last few years reading index methodology documents, because when you work against listed company benchmarks you learn quickly that the rulebook matters more than the level. S&P publishes its constituents, its float adjustments and its rebalancing dates. The ASX benchmarks publish theirs. Anybody can reconstruct the number from the method. That transparency is not a courtesy, it is the only reason an index is allowed to be a decision input.
I have never once been handed the prompt list behind an AI visibility number without asking for it twice.
Seventy eight and thirty four are both real. Fifty six is not. The arrow on top of it is the only part anyone ever reads.
Thirty years of building software has mostly taught me to distrust a number I cannot take apart. If an AI visibility score is about to go into your board pack, book a call and we will decompose it first.
