How we measure what AI recommends.
Everything needed to check our numbers or disagree with them: the prompts, the models, the sample size, the score formula, and what the data cannot tell you. If you recompute this and get something different, we want to hear about it.
Identical questions, three assistants, web search on.
For each trade we wrote six questions a real buyer would ask, mixing straightforward commercial phrasing with the messier situational kind — a burst pipe, a hail claim, a cracked tooth. Each question is asked once per city per assistant, and every answer is stored verbatim with its citations. That is 6 questions × 3 assistants = 18 recorded answers per city and trade, and 540 in total across the study.
- ChatGPT gpt-5-mini · web search enabled
- Google Gemini gemini-flash-latest · web search enabled
- Perplexity sonar · web search enabled
Web search is not optional, and this is the part most similar studies get wrong. Our own first run queried these models through plain completion endpoints with no search tool, and measured them refusing to name any business in up to 100% of answers. That number was real and completely misleading: it described an API configuration, not anyone's market. A person using ChatGPT gets an assistant that searches before it answers. Asked the identical question with search enabled, the same model named six Tampa roofers with citations. We threw the first run away and re-collected everything.
One company spelled three ways counts once.
Assistants name the same business inconsistently — with and without the LLC, with and without the city, with the ampersand spelled out. Counting those as separate businesses would inflate how many companies get named and understate how much the assistants agree, so names are resolved before anything is counted: punctuation is normalised, trade words and the city's own name are treated as carrying no identity, and what remains is compared. Where a shorter name is contained entirely within a longer one, the two are merged.
The resolution is deliberately conservative and it is not perfect. It will not catch a genuine rebrand or a misspelling of a surname, and it will not merge two names that share no tokens. Every merge it performs is recorded in the scan file so it can be audited rather than trusted.
What the 0–100 number is, exactly.
The openness score describes a market, not a business in it. It answers one question: is there room to become the named company here, or does someone already own the answer? Higher means less settled. It is the average of three measured shares, weighted equally:
- Leader concentration 1 − (answers naming the top business ÷ all answers)
- Consensus 1 − (shortlist businesses named by all three ÷ shortlist size)
- Churn businesses named exactly once ÷ all businesses named
Equal weights are a choice, and an honest one. Tuned weights would imply we know the relative importance of these three signals from eighteen answers per market, and we do not. Equal thirds is a rule we can state in a sentence, and it means the score cannot be quietly reshaped to make a market look better than it is.
Two figures we publish but deliberately keep out of the score. The disagreement rate — how often the three assistants returned no shared business for the same question — comes from six comparisons per market, so it moves in seventeen-point steps; folding it in would add noise dressed up as signal. And the directory citation share is a different axis altogether: it tells you which playbook applies in a market, not how contested it is. Both sit beside the score on every page instead of inside it.
Scores use absolute scales and are never ranked against each other. A percentile would shift every published number the moment a new city is scanned, and would make this quarter incomparable with the last — which is the one thing an index has to get right. The consequence is that scores cluster between roughly 50 and 90 rather than using the full range: a market scoring near zero would mean one business owned every answer, and none do.
- 80 and abovewide open
- 65 to 79contested
- Below 65consolidating
What these numbers cannot tell you.
It is a snapshot, not a ranking. These models are non-deterministic and they re-read the web constantly. The same question asked next month will not return the same list. Every page carries its collection date for that reason, and we re-run markets rather than citing an old run indefinitely.
Eighteen answers per market is a small sample. It is enough to see who the assistants reach for repeatedly and which sources they lean on. It is not enough to rank two businesses one position apart, and we do not.
Being named is not an endorsement, and neither is being absent. The assistants named the businesses whose evidence was easiest to find and safest to repeat. That correlates with being well-documented, not with doing the best work. None of the businesses named in this study are our clients, and nobody paid to appear or to be left out.
Six trades and five cities is not the country. The trade-level patterns hold across all five cities we measured, which is a reason to take them seriously and not a reason to assume they generalise to markets we have not looked at.
Cite it, quote it, argue with it.
The findings are published under CC BY 4.0, so journalists, analysts, and competitors are welcome to reuse them with attribution. If you are checking our work, the collection date and the exact prompts are on every market page, and the study index lists all 30 of them.
Want this measurement for your own business rather than your market? The free AI visibility report runs the same method against your name and your competitors.