How we measure what AI recommends.
Everything needed to check our numbers or disagree with them: the prompts, the models, the sample size, the score formula, and what the data cannot tell you. If you recompute this and get something different, we want to hear about it.
The same questions, three assistants, web search enabled.
For each trade we wrote six questions a real buyer would ask, mixing straightforward commercial phrasing with the messier situational kind — a burst pipe, a hail claim, a cracked tooth. Each question is asked once per city per assistant, and every answer is stored verbatim with its citations. That is 6 questions × 3 assistants = 18 recorded answers per city and trade, and 540 in total across the study.
- ChatGPT gpt-5-mini · web search enabled
- Google Gemini gemini-flash-latest · web search enabled
- Perplexity sonar · web search enabled
Why require web search? We wanted to measure the kind of answer people receive from current AI products, which can search the web before recommending a business. In a preliminary test we turned search off. The assistants often declined to name anyone because they had no current local information. That measured the test setting, not the market. We discarded that test. Every result published here comes from a new run with web search enabled, and we saved the sources returned with each answer.
One company spelled three ways counts once.
Assistants do not always spell a business name the same way. One answer may include “LLC,” another may add the city, and another may spell out an ampersand. We remove punctuation, company endings, and the city name, then merge clear variants before counting. Otherwise one business could look like two or three different businesses.
We merge only strong matches. That avoids combining unrelated companies, but it can miss a rebrand or a badly misspelled name. Every merge is recorded in the source data so it can be checked.
What the 0–100 number is, exactly.
AI recommendation openness is our 0–100 summary of how dispersed the business names were across 18 sampled answers. A higher score means names were spread across more businesses and repeated less consistently. A lower score means a smaller group of businesses appeared more often. It is not the same as market saturation: it does not measure how many businesses operate there, customer demand, search volume, or how difficult the market is to enter.
The score combines three observable patterns, weighted equally:
- No dominant leader Higher when the most-named business appears in fewer of the 18 answers. Formula: 1 − leader concentration.
- Low assistant consensus Higher when ChatGPT, Gemini, and Perplexity share fewer of the same repeat names. Formula: 1 − consensus.
- Many one-off names Higher when a larger share of businesses appears exactly once. Formula: single-mention churn.
We weighted the three parts equally because this small study does not justify claiming that one matters more than another. Equal thirds also keeps the formula easy to reproduce and prevents us from changing the weights after seeing the results.
We show two other useful numbers separately. The disagreement rate counts questions where the assistants shared no business. The directory share counts how many cited sources came from directories. Neither is part of openness: disagreement is based on only six question comparisons, and source type describes evidence, not how concentrated the names were.
The score is not a percentile. A score of 70 uses the same formula now and next quarter, and adding a new city does not change an old score. The results do not need to fill the entire 0–100 range. In this study they happen to cluster between roughly 50 and 90.
- 80 and abovehighly dispersed recommendations
- 65 to 79mixed recommendation pattern
- Below 65more concentrated recommendations
What these numbers cannot tell you.
It is a snapshot, not a ranking. These models are non-deterministic and they re-read the web constantly. The same question asked next month will not return the same list. Every page carries its collection date for that reason, and we re-run markets rather than citing an old run indefinitely.
Eighteen answers per market is a small sample. It is enough to see who the assistants reach for repeatedly and which sources they lean on. It is not enough to rank two businesses one position apart, and we do not.
Being named is not an endorsement, and neither is being absent. The assistants named the businesses whose evidence was easiest to find and safest to repeat. That correlates with being well-documented, not with doing the best work. None of the businesses named in this study are our clients, and nobody paid to appear or to be left out.
Six trades and five cities is not the country. The trade-level patterns hold across all five cities we measured, which is a reason to take them seriously and not a reason to assume they generalise to markets we have not looked at.
Cite it, quote it, argue with it.
The findings are published under CC BY 4.0, so journalists, analysts, and competitors are welcome to reuse them with attribution. If you are checking our work, the collection date and the exact prompts are on every market page, and the study index lists all 30 of them.
Want this measurement for your own business rather than your market? The free AI visibility report runs the same method against your name and your competitors.