Methodology
We asked ChatGPT to recommend a local business the way a real buyer would. We left web search un-forced, and we recorded the full search chain, including the searches that never produced a citation.
Scale
LLM output changes from run to run, so we built the study to hold up anyway. There's a core measured set, then repeat runs and intent tests on top, then a higher-volume pass for the parts that need a lot of named businesses.
| Measure | Count |
|---|---|
| Core recommendation sessions (GPT-5.5, un-forced) | 200 |
| Industries × states × core prompts | 10 × 10 × 2 |
| Variance re-runs (repeated measurement) | 120 (4 passes over 10 industries × 3 states) |
| Sub-service sessions (intent-robustness) | 30 |
| Higher-volume pass for named-business data | about 1,000 |
| Total web searches logged | 5,933 (about 30 per session, max 58) |
| Reasoning tokens per session (avg) | about 5,059 |
| Citations captured | 1,982 |
| Distinct sources cited | 858 |
| Businesses named per answer (avg) | about 8 |
| SERP rank checks (both directions) | 3,431 plus reverse-direction |
Process
Un-forced, full-chain capture. Most AI search studies force a single search and read the citation. That's clean, but it isn't what really happens. We let the model run its real multi-step process, so the data reflects what a paying customer actually triggers.
And we logged every query in the chain, including the ones that never produced a visible citation.
One assistant, two settings. Everything here is ChatGPT. The core study runs it the way a buyer gets it today, which is GPT-5.5 with reasoning and live web search.
One part of the analysis needs a lot more named businesses than the deep reasoning runs can cheaply produce. That's the rank tests (Finding 1) and the count of distinct winners per market (Finding 5.2). So for that part we ran a higher-throughput ChatGPT pass (gpt-5-chat-latest) to collect business names at volume.
Every behavioral finding (the audit, the credential checks, the citation and source mix) comes from the reasoning runs. We point out which is which where it matters.
Date (for reproducibility). Runs were on June 16 and 17, 2026, through the Responses API web_search tool, with reasoning effort set to high.
The prompts, verbatim ({X} = trade, {sub-service} = a specific job, {city, state} = location):
- "I'm looking for a reputable {X} in {city, state}. Help me find a good one I can hire."
- "Who are the best, most reputable {X}s in {city, state} that I can hire?"
- "I need {sub-service} in {city, state}. Who should I hire?" (a specific job with no "reputable" cue: car accident lawyer, water heater repair, dental implants, mortgage refinance, couples therapist, Botox, managed cybersecurity, tree removal, homeowners insurance, kitchen remodel). We used this one to test whether the audit only happens because we asked for "reputable."
Variance-tested. LLM output changes from run to run, so we re-ran cells several times and report the key behaviors as ranges, not single numbers.
Scope
This is ChatGPT specifically, on locally hired service businesses. We did not test other assistants or other business types (SaaS, ecommerce, retail). See Limitations.
Finding 1: Search rankings are decoupled from recommendations
Ranking isn't necessary and it isn't sufficient. 90% of recommended businesses don't rank top 20, and 70% of top-ranked businesses aren't recommended. This is the most misunderstood part of AI recommendations, so we're starting with it.
1.1 The mechanism: it never runs the query a ranking would answer
ChatGPT never runs the buyer keyword, so your rank for it can't be a direct input. Its searches average 8.4 words and go after named credibility sources. It doesn't run the short query a buyer types, only shortlist and verification queries.
Take "plumber in Dallas." A buyer types that into Google and a ranking decides what they see. But ChatGPT doesn't type it at all.
Its real logged queries for the same job look like the examples in Finding 6.3. There's a ranking lookup like Dallas plumbers Best Pick Reports top rated, a trust check like site:bbb.org/us/tx/dallas plumber accredited A+, then a license check like site:tsbpe.texas.gov "[company] Plumbing" responsible master plumber.
By then it has already picked candidates from the authority tier and is just confirming them. So your rank for "plumber in Dallas" never gets read.
(We observed the mechanism on the reasoning runs. The rank distributions in 1.2 and 1.3 use the higher-volume pass, so there are enough named businesses to check against. See Methodology.)
1.2 Recommended businesses don't rank, and the few that do sit on page 3
80% of recommended businesses don't rank anywhere in the top 50 of Google or Bing.
We took the businesses ChatGPT recommended and looked up where each one actually ranks for the main buyer keyword a customer would type ("personal injury lawyer Houston," "plumber Dallas," "managed IT services Chicago," "tree service Phoenix"). For the most part, the answer is nowhere:
| Organic position | Bing | |
|---|---|---|
| Top 3 | 0% | 0.3% |
| 4 to 10 | 4.1% | 0.9% |
| 11 to 20 | 5.5% | 0% |
| 21 to 50 | 10.3% | 0% |
| Nowhere (>50) | 80.1% | 98.7% |
Only 5% rank top 10 in either engine. The ones that do rank average position 23, which is page three.
1.3 And ranking isn't enough either (the reverse direction)
Of businesses that rank in Google's top 10 for the buyer keyword, only 30% get recommended.
We ran it the other way too. For each of those same buyer keywords ("dentist Phoenix," "mortgage broker Dallas," and so on), we took the businesses Google ranks in its top 10 and checked how often ChatGPT recommends them back:
| Industry | Top-10-ranked → recommended |
|---|---|
| Medical spa | 39.8% |
| Therapist | 36.8% |
| Home renovation | 36.5% |
| Plumber | 36.2% |
| Insurance | 32.9% |
| Mortgage | 28.1% |
| Arborist | 27.7% |
| IT services | 26.0% |
| Lawyer | 21.7% |
| Dentist | 17.2% |
| Overall | 29.7% |
And inside that top 10, rank position barely matters. If you pool all ranked businesses by position, the top 3 get recommended about 26% of the time and positions 4 to 10 about 24%, which is basically flat. Ranking first (31%) is no better than ranking eighth (30%).
So it isn't even true that ranking higher helps once you're on page one.
That means 70% of top-10-ranked businesses don't get recommended. That 30% is higher than chance alone would give you, because the model only names a handful of businesses out of the many in each market. But it's nowhere near the close to 100% you'd expect if rank drove recommendations.
The simplest explanation is that rank and recommendation share a cause further up, rather than rank being an input itself. A prominent business tends to both rank and land in the authority rankings. So ranking isn't a direct input to recommendations, and at best it's weakly correlated through prominence.
1.4 "Search rank" isn't even one thing: Google and Bing barely agree
Of businesses that rank top 50 on at least one engine, only 4% rank on both Google and Bing.
| Where a top-50-ranked business appears | Share |
|---|---|
| On both Google and Bing | 4% |
| On one engine only | 96% |
So the "ChatGPT runs on Bing, so optimize for Bing" idea falls apart here. The two engines show almost totally different businesses for the same query.
We pulled both engines through the same SERP API under matched conditions (same location, language, device and depth, de-personalized). So the gap reflects real differences between the engines, not session or geo personalization.
There's no single "ranking" to win. That fits with the model not using either engine's results page to choose who to recommend.
1.5 The contrast with citation studies
Search rank decides citations, but it doesn't decide recommendations. Ranking and recommendation overlap on 30% of cases at most, and only 5% of recommended businesses rank in any top 10.
Citation-focused research (like the widely cited query fan-out work) finds that retrieval rank strongly predicts whether a page gets cited in an informational answer. The "rank" in that work is a page's position inside ChatGPT's own retrieved results, not its Google or Bing ranking.
This study measures traditional Google and Bing organic rank, which is the thing businesses actually optimize. They're two different ladders, so that research doesn't conflict with this study, because it measures a different outcome.
Our data is about recommendation, and there Google and Bing rank doesn't predict who gets named. Measured in both directions, the overlap is weak:
| Overlap between ranking and recommendation | Rate |
|---|---|
| Top-10 ranked → recommended | 30% |
| Recommended → ranks anywhere in top 50 | 20% |
| Recommended → ranks in top 10 | 5% |
Traditional rankings and recommendations are pretty close to decoupled. Rank can help your content get quoted, but it won't get your business handed to a buyer.
Key takeaways (Finding 1): the model never runs the buyer keyword. Recommended businesses mostly don't rank, ranked businesses mostly aren't recommended, and "search rank" splits across engines (4% Google and Bing overlap). So rank isn't necessary and it isn't sufficient.
Finding 2: Your own site gets you into the running, not recommended
About half of everything ChatGPT cites in the real audit is a business's own site (roughly 52% of the 1,982 citations, classifier-estimated). This is where the data corrected our first guess. We expected business websites to be marginal in the recommendation path, and they aren't.
The model really does read your pages. What matters is what it reads them for, and which pages.
2.1 Your site is cited about half the time, but the rate swings a lot by industry
Your own site is about half of every citation the model makes. But how much it leans on your pages swings by industry, from 10% (plumbers) to 74% (med spas).
| Industry | Citations that are the business's own site |
|---|---|
| Medical spa | 74% |
| Therapist | 73% |
| Arborist | 70% |
| Lawyer | 68% |
| IT services | 54% |
| Dentist | 54% |
| Insurance | 44% |
| Mortgage | 34% |
| Home renovation | 23% |
| Plumber | 10% |
| Overall | about 52% |
The model leans on your own site where third-party records are thin. Therapists, med spas and arborists have few directories, so it reads the businesses' sites directly.
And it barely touches your site where strong third-party records exist. A plumber recommendation is about 90% third-party: BBB, contractor boards, BuildZoom.
So your own site is a fallback the model uses when it has nothing better, rather than the source that decides it. (The business versus third-party split is classifier-estimated, so read these as directional.)
2.2 And it's mostly your homepage, used to confirm you exist and offer the service
When the model reads your site, it's mostly reading your homepage. Of the business pages we have full URLs for, about 70% were the homepage and 30% were a deeper service or location page.
| Industry | Homepage (/) | Deeper page (service or location) |
|---|---|---|
| Insurance | 80% | 20% |
| Home renovation | 78% | 22% |
| Arborist | 74% | 26% |
| Lawyer | 72% | 28% |
| Mortgage | 72% | 28% |
| Medical spa | 70% | 30% |
| Dentist | 69% | 31% |
| IT services | 68% | 32% |
| Therapist | 57% | 43% |
| Plumber | 55% | 45% |
| Overall | about 70% | about 30% |
The deeper pages are what you'd expect, service and location landing pages like rotorooter.com/newyork/, nycarborist.com/manhattan, blondiestreehouse.com/arborist-services/ and emeraldtreecare.com/about/meet-our-arborists/.
Trades that organize by service or city (plumbers and therapists) get their deeper pages read most often, up to 45% of the time. And the buyer's intent shifts it too.
A specific, urgent question ("my water heater just burst, who do I call?") pulled deeper service pages 39% of the time, against 15% for a generic "I need a plumber." So the more specific the job, the more the model reaches past the homepage.
But even then the homepage leads. Service and location pages do matter, because they're how the model confirms you serve this job in this place. The homepage is still the main check though.
2.3 But never on its own: every recommendation was credibility-checked
Your site alone is never enough, because every recommendation was also credibility-checked. Across the 196 sessions that produced a recommendation, all 196 ran a credibility check. Zero were standalone, and each one came with about 17 verification searches.
We checked every session for a credibility check (a verification search, or a citation to a third-party credibility source) in the same session where it read the business's pages. It always did.
| Across the 196 sessions that produced a recommendation | Result |
|---|---|
| Recommendations that also ran a credibility check | 196 of 196 (100%) |
| Standalone recommendations (named on own pages, no credibility check) | 0 |
| Credibility-verification searches per recommendation (avg) | about 17 |
Not one business got recommended on the strength of its own website alone. The homepage that gets cited 70% of the time is read in the same session as an average of about 17 credential and trust verification searches.
Your pages confirm you're real and relevant, but they never replace the audit. And no single business site recurs across recommendations the way the BBB or a state board does (Finding 5.1). So your pages confirm you, but they don't set you apart.
Getting your site right makes you eligible. It's still a long way from getting you recommended.
Key takeaways (Finding 2): your own site is about half of all citations, and the model really does read it, mostly your homepage, but never on its own. Every recommendation was also credibility-checked (0 standalone, about 17 verification searches each). Your pages make you eligible, and third-party credibility is what gets you recommended.
Finding 3: Reddit and user-generated content are not used
Reddit was cited zero times in the 1,982-citation audit. "Get big on Reddit" has become almost as common a piece of AI visibility advice as "get more brand mentions." But in the mode that actually recommends businesses, it does nothing.
3.1 Zero Reddit citations in the real audit
Reddit was cited zero times in the real audit, out of 1,982 citations. In a forced single-search pass it shows up 171 times.
| Pass | Reddit citations | Total citations | Reddit share |
|---|---|---|---|
| Real reasoning audit | 0 | 1,982 | 0% |
| Forced single-search pass | 171 | 13,655 | 1.3% |
In the reasoning audit, Reddit was cited zero times. The only user-generated or social citations of any kind were two LinkedIn business profiles. Quora, Facebook and forums got zero.
3.2 Reddit only shows up when the model searches shallow
Reddit only shows up when the model searches shallow: 171 citations in the forced single-search pass, against 0 in the real audit. A one-shot search lands on whatever ranks for the query. Reddit threads rank well, so they get pulled in.
The real reasoning audit never works like that. It sends targeted queries to named credibility sources (boards, associations, the BBB), and a Reddit thread isn't any of those. So Reddit citations come from shallow retrieval, and the model doesn't use them to recommend.
This matters because simple checks and a lot of AI visibility tools rely on exactly that kind of one-shot call. That's why they over-report how much Reddit matters.
Key takeaways (Finding 3): the recommendation audit cited Reddit zero times. It only shows up (171 times) in a shallow single search, and that isn't how recommendations get made. Being active on forums doesn't move a recommendation.
Finding 4: The platforms you pay for are bystanders
The five platforms businesses most often pay for add up to just 8 of 1,982 citations. The directories you budget for barely register in the audit.
4.1 Paid platforms in the real audit
The big five paid platforms (Yelp, Angi, Avvo, Thumbtack, HomeAdvisor) add up to just 8 of 1,982 citations.
| Platform | Citations (of 1,982) |
|---|---|
| Healthgrades | 9 |
| Yelp | 6 |
| Angi | 2 |
| Trustpilot | 2 |
| Thumbtack / HomeAdvisor / Porch | 0 |
| Avvo / Justia / FindLaw | 0 |
| Nextdoor / Bark / Networx | 0 |
Healthgrades, a medical directory, is the most-cited paid platform at 9, and that's still negligible. Vitals, Martindale and Lawyers.com are at or near zero.
4.2 They show up in shallow searches, like Reddit
Like Reddit, paid platforms mostly show up in shallow searches. HomeAdvisor goes from 0 in the real audit to 22 in the forced pass, and Yelp goes from 6 to 12.
| Platform | Real audit | Forced single-search pass |
|---|---|---|
| Healthgrades | 9 | 26 |
| HomeAdvisor | 0 | 22 |
| Yelp | 6 | 12 |
| Angi | 2 | 6 |
| Avvo | 0 | 3 |
Paid directories show up several times more often in the shallow forced pass, where the model grabs whatever ranks, than in the real audit, where it's checking credibility. These are advertising marketplaces, and the audit isn't looking for ads.
So paying to appear on a platform the audit never checks doesn't touch the recommendation. It might still drive direct traffic, but that's a separate question.
Key takeaways (Finding 4): the paid directories businesses budget for are almost entirely missing from the audit (8 of 1,982 for the big five), and what little shows up comes from shallow searches. Spending on these platforms doesn't buy you a recommendation.
Finding 5: There is no universal playbook
ChatGPT cited 858 distinct sources. 95% show up in only one industry, and 68% were cited only once. The sources that decide recommendations are specific and fragmented, and mostly nobody is optimizing for them.
5.1 The sources are very industry-specific, and mostly one-off
68% of all cited sources showed up exactly once, and the average source spans just 1.1 of 10 industries. The only sources that cross industries are generic trust and government sites, nothing you'd "market into":
| Cross-industry source | Industries (of 10) | Type |
|---|---|---|
| mass.gov | 8 | govt umbrella portal |
| bbb.org | 7 | trust directory |
| pa.gov | 6 | govt umbrella portal |
| content.boston.gov | 5 | govt portal |
| expertise.com | 5 | listicle |
| (everything else) | 4 or fewer. 95% appear in exactly 1, 68% cited only once | industry-specific |
The top 20 sources are only 34% of all citations, so there's a long, fragmented tail. That doesn't contradict the "same stack everywhere" idea in Finding 6.
The structure is universal: discover, verify, trust, the same in every market. What fills each tier isn't. The exact sources change by industry, by state, and even by service inside an industry.
So the shape stays the same and the contents are local. That's why there's no list to copy, only a process you run for each market.
5.2 The winners are hyper-local too: it's 100 separate local races
Pool about 100 recommendations per industry and you get 203 to 425 distinct businesses, so there's no national winner.
We also looked at who gets recommended as well as the sources, and the winners are as fragmented as the sources are. This winner count uses the higher-volume pass (about 100 recommendations per industry), because counting distinct winners needs a lot more named businesses than the deep reasoning runs produce.
| Signal | Result |
|---|---|
| Distinct businesses recommended per industry (about 100 recommendations) | 203 to 425 |
| Most-recommended business's share of an industry | 20% at most (no dominant national winner) |
| Within a single city, how often the leading business recurs across differently-worded queries | 78% (sticky per market) |
| Exception, a national brand winning across markets | Roto-Rooter (plumbers), 20 of 100 |
So AI recommendation is basically about 100 separate local races. Within one market the leader is locked in (78% consensus across phrasings). But across markets the winners are all different, hundreds of them, because the audit resolves locally.
National brands can sometimes win broadly (Roto-Rooter), but most verticals have purely local winners. So the winners don't fit a generic playbook any better than the sources do.
5.3 BBB is the one near-universal trust anchor
One source, the BBB, is 13% of every citation in the study (261 of 1,982). It dominates where a profession has no prestige system of its own, and it disappears where one exists:
| Industry | BBB citations | Has its own prestige ranking? |
|---|---|---|
| Plumber | 81 | no |
| Insurance | 62 | no |
| Mortgage | 56 | no |
| Home renovation | 51 | no |
| Arborist | 8 | no |
| IT services | 2 | yes |
| Lawyer / Dentist / Therapist | 0 | yes |
It checks the BBB by exact profile URL, by name (site:bbb.org/us/tx/houston/profile/plumber … A+). So it uses the profession's own authority where one exists, and the BBB where it doesn't.
5.4 Listicles, editorial press and media coverage: a top lever in some industries, irrelevant in others
Editorial press (city magazines, "best of" listicles, trade press) is 13.5% of citations overall, but that average hides a split. It shows up in 85% of lawyer and IT sessions and 0% of therapist and arborist sessions.
| Industry | Sessions with an editorial source (of 20) | Dominant editorial sources |
|---|---|---|
| Lawyer | 17 | Best Law Firms, Chambers, Best Lawyers, Forbes |
| IT services | 17 | Clutch, CRN, Channel Futures |
| Dentist | 13 | Seattle Met, Philly Mag, Houstonia, Castle Connolly |
| Plumber | 5 | Best Pick Reports, Expertise |
| Medical spa | 4 | city magazines, RealSelf |
| Mortgage | 4 | Expertise, WalletHub |
| Insurance | 2 | n/a |
| Home renovation | 2 | NY Magazine, Houzz |
| Therapist | 0 | (runs on associations instead) |
| Arborist | 0 | (runs on certifications instead) |
So why the split? Editorial press is how ChatGPT fills the authority-ranking tier (tier two).
Fields with a strong third-party ranking culture give it plenty to cite. Law has Best Law Firms and Chambers, IT consulting has Clutch and CRN, and dentistry has city "Top Dentists" features.
Fields that run on membership and certification instead (therapy on APA and Gottman, arboriculture on ISA and TCIA) don't give it anything like that to cite. So it falls back to those directories.
Unregulated IT shows a second effect on top of that. With no licensing board to check, listicles like Clutch become the stand-in for the credibility the license tier would normally provide.
So it isn't simply "no regulator means more press," because law has a strong regulator and heavy press. What drives it is whether the field has a public ranking and press culture at all.
The authority tier is really two separate things, and in SEO they aren't interchangeable. There are review aggregators (Clutch, RealSelf, Zillow), which rank businesses by collected ratings and reviews. And there are editorial "best of" rankings and listicles (Best Law Firms, Chambers, Expertise.com, a city magazine's "Top Dentists"), where a publication curates the list using its own methodology.
Which one the model leans on depends entirely on the field:
| Source | Type | Citations | Where it carries |
|---|---|---|---|
| Clutch | review aggregator | 67 | IT services (almost only) |
| Best Law Firms | editorial ranking | 66 | Law |
| City magazines ("Top X") | editorial listicle | 47 | Dentist, med spa |
| Expertise.com | editorial listicle | 18 | Plumber, mortgage |
| Best Lawyers | editorial ranking | 15 | Law |
| Chambers | editorial ranking | 12 | Law |
| Forbes | editorial media | 8 | Law |
| RealSelf / Zillow | review aggregator | 1 each | med spa, mortgage |
Editorial rankings and listicles carry most of the citations (about 170), and review aggregators get far fewer (about 69, almost all Clutch). Aggregators own exactly one field here, IT services, where Clutch is basically the category ranking. Everywhere else, editorial rankings and listicles dominate.
So the play depends on your vertical. Go after the aggregator if you're in IT (Clutch), and go after the editorial ranking or listicle otherwise (Best Law Firms, a city "Top X" feature).
A placement in the right one for your field is a direct recommendation lever. In therapy or tree care, it's close to wasted effort.
5.5 Professional associations are the biggest hidden layer
Associations are the main authority tier for fields driven by certification and membership, like arborists, therapists and remodelers, and almost no one markets to them. For arborists, the association directories (treesaregood and TCIA, 21 citations) outrank the firms themselves.
| Association (citations) | Industry it carries |
|---|---|
| NARI: nariatlanta (18), greaterphoenixnari (9), nari.org | Home renovation |
| ISA / TCIA: treesaregood (11), treecareindustryassociation (10) | Arborist (top sources, above the firms) |
| TrustedChoice / Big "I" (8) | Insurance |
| ADA: findadentist (4) | Dentist |
| Gottman (3), ABCT (3), APA locator | Therapist |
Recognized membership is a credibility signal the model can verify, and it treats the association directory as a trusted shortlist. So active, listed membership does a lot of work here, and anyone optimizing for Google or Yelp won't even see it.
Key takeaways (Finding 5): sources and winners are both fragmented (95% single-industry sources, 68% one-off, 203 to 425 distinct winners per vertical). BBB is the one near-universal lever, editorial press is a big but uneven PR lever, and associations are the biggest hidden layer. The stack is different for every industry and state.
Finding 6: How it really works, the credibility audit
Every recommendation is roughly a 30-search investigation. This is the credibility layer in full, and it's what everything above points to.
When a buyer asks ChatGPT to recommend a provider, it doesn't read a rankings page and copy the top results. It runs a credibility audit. Across 200 sessions it averaged around 30 web searches and about 5,059 tokens of reasoning per recommendation, and it checked roughly seven distinct sources.
It works in a consistent order: find candidates, verify their credentials, cross-check a trust directory, then name eight or so businesses. The shape was the same in every industry.
6.1 A recommendation is a 30-step search process
It runs about 30 searches per recommendation, and two-thirds of them are hidden verification that never turns into a visible citation.
| Metric | Value |
|---|---|
| Searches per session | avg about 30, max 58 |
| Internal reasoning per session | about 5,059 tokens |
| Distinct sources consulted per session | about 7 |
| Businesses named in the final answer | avg 8.2, max 13 |
| Searches per citation produced | about 3 to 1 |
Roughly two-thirds of the searches, nearly all credential checks, never show up as citations. And about half of what it does cite is a regulator, ranking or directory rather than a business (about 50%, classifier-estimated).
There's a mild tendency in the ordering. Discovery and ranking queries lean slightly earlier in the chain than credential and trust queries (mean normalized position 0.41 versus 0.49). It's a small gap, so read it as a lean toward "discover first," not a clean two-phase sequence.
Either way, you learn more from the search chain than from the citation list.
6.2 Every industry runs the same three-tier stack
It's the same three tiers in all 10 industries. Only what's in them changes.
| Industry | ① Licensing / credential | ② Authority ranking | ③ Trust directory |
|---|---|---|---|
| Lawyer | State Bars (calbar, texasbar…) | Best Law Firms, Chambers, Super Lawyers | bar referral |
| Plumber | Contractor boards (CSLB, AZ ROC, WA L&I) + permits | Expertise, Best Pick Reports | BBB |
| Dentist | Dental boards + ADA find-a-dentist | City "Top Dentists" magazines, Castle Connolly | Zocdoc |
| Mortgage broker | NMLS + CFPB + state DFI/DRE/DFS | Expertise, Zillow | BBB |
| Therapist | State licensing + APA/ADAA locators | n/a | Psychology Today |
| Medical spa | FDA + state medical/cosmetology boards | RealSelf, city magazines | BBB |
| IT services | n/a (unregulated) | Clutch, CRN, Channel Futures | Microsoft Partner |
| Arborist | ISA / TCIA / ASCA certification | n/a | BBB |
| Insurance agency | State Insurance Depts (CA DOI, TX TDI) | n/a | BBB, TrustedChoice |
| Home renovation | Contractor boards + NARI | Houzz, NY Magazine "best contractors" | BBB, BuildZoom |
The tiers adapt. Unregulated IT drops the licensing tier, and therapy and insurance, which have no prestige ranking, lean on associations and trust directories instead.
6.3 The model's queries are nothing a human would type
Its searches average 8.4 words (a human's is about 4), and 18% use a site: operator. These are machine-shaped queries, and no human would type them.
| Query trait | Value |
|---|---|
| Average length | 8.4 words |
| Longer than 6 words | 77% |
Uses a site: operator | 18% |
| Verifies a name in quotes | 14% |
| Carries a recency year (2023 to 2026) | 9% |
| Broad "brand buzz" sweeps | 0 |
Verbatim examples:
site:mqa-internet.doh.state.fl.us "Leslie Haller" "Dentist" "Coral Gables"
Abraham Watkins Houston personal injury attorneys board certified Best Law Firms Tier 1
https://data.wa.gov/resource/m8qx-ubtq.json?$select=businessname (a raw state open-data API endpoint)
To check a Washington contractor, ChatGPT queried a government open-data API directly. It makes these queries up on the fly, and there's basically no limit to how varied they get.
You can't optimize for a query you can't see. And you can't see these unless you capture them from real ChatGPT calls.
6.4 The audit holds without asking for "reputable"
Take out the "reputable/best" cue and credential verification still fired in 93% of sub-service sessions.
| Sub-service query | Credential-verified | What it specialized to |
|---|---|---|
| dental implants | ✓ | TX dental board, UCLA dentistry, Mount Sinai |
| Botox | ✓ | Mount Sinai, Weill Cornell, CDC |
| couples therapist | ✓ | Gottman Referral Network |
| mortgage refinance | ✓ | CFPB, Freddie Mac |
| tree removal | ✓ | TCIA, city forestry |
| homeowners insurance | ✓ | NY DFS, CA DOI, TX TDI |
| car accident lawyer | ✓ | State Bar, courts |
| water heater repair | ✓ | city permits, BBB |
93% of sub-service sessions (28 of 30) verified credentials, at 24.4 searches each. So the audit holds up whatever the buyer's intent is, and it isn't something our wording caused.
6.5 It verifies against the specific licensing body for each state
Credential verification is the main behavior across the study, firing in 74 to 96% of regulated-industry sessions. It checks about six candidates by name per session, and it checked a license by name in 40% of sessions.
| Measure | Result |
|---|---|
| Sessions touching a government/regulator source | 90% (98% of regulated industries) |
Named-individual license checks (site:[board] + a name) | 40% |
| Distinct candidates verified by name per session | avg 5.5, max 51 |
| Regulated-industry verification (variance-tested) | 74 to 96% (mean 87%) |
| Regulated markets verifying at least once | 100% |
The credential check is the main behavior, and it keeps coming back. Every regulated market triggers it, but it doesn't happen 100% of the time. On any given run, roughly 1 in 5 regulated recommendations skip the explicit license step.
So the behavior is reliable but the rate moves around, which is why we report a range. For health verticals it also reaches academic and clinical institutions (about 39 .edu citations: UCLA/USC dentistry, Mount Sinai for med-spa).
What we measured is a query pattern. Lawyers route to the State Bar, mortgage brokers to NMLS, and contractors to the state registry. The data supports the functional claim, but the cognitive claim ("it knows") is our interpretation.
6.6 The right board appears for each state
Each industry surfaced 16 to 31 distinct regulators, and the right state's board showed up in most runs.
| Add this state… | …and the model verifies against |
|---|---|
| Arizona | azbar.org, AZ Registrar of Contractors, azdentalboard.us |
| Florida | floridabar.org, myfloridalicense.com, FL Dept. of Health |
| Washington | wsba.org, secure.lni.wa.gov, doh.wa.gov |
| Texas | texasbar.com, Texas Board of Legal Specialization, tsbpe.texas.gov |
In some states one portal covers lots of trades. mass.gov was the verification source across 8 different industries.
6.7 Precision beats volume
In 5,933 searches, ChatGPT ran zero broad "what's-being-said-about-this-brand" sweeps. Every search is a targeted query against a named source.
And when it does search for a specific business by name, it never searches the name on its own. Every branded query in the sample paired the name with a verifier: a site: check against a licensing board, a BBB profile lookup, or a qualifier like "Best Law Firms Tier 1" or "complaints."
| Query behavior | Of 5,933 searches |
|---|---|
| Broad brand-buzz sweeps ("what's being said about X") | 0 |
| Bare-brand exploration (a name with no verifier) | 0 |
| Names a specific business, always with a verifier attached | 14% |
So "get more brand mentions everywhere" misses the point. A mention on fifty mid-authority sites is on zero of the sources the audit checks, while the one right source (your board, association or category ranking) gets checked directly.
The only way the shotgun approach works is by seeping into a future model's training data. That's a bet that pays off at the next model cycle (months to years out). It doesn't help with today's buyer, who gets served by a live audit that mentions don't touch.
Key takeaways (Finding 6): each recommendation is a 30-search investigation, and two-thirds of it is hidden verification. It follows a universal three-tier stack, verifies against the specific state licensing body, and uses precise, machine-shaped queries that never search a brand name on its own.
How to actually get recommended
Everything above covers what doesn't move a recommendation. Here's what does, in the order the audit cares about.
The deciding sources are specific, fragmented and per-market. So there's no list you can copy, and you have to run this as a process for each market you serve.
- Cover the basics so you're in the running. ChatGPT reads your own pages to confirm you exist, offer the service and work in the market, and that's about half of what it cites (Finding 2). A homepage, a page per service and a page per location clear this bar, and most sites already have them. This won't get you recommended, but it makes you eligible.
- Win the authority tier. This is where the work pays off most, and it's per-market. It's a mix of review aggregators where they carry weight (Clutch for IT services), curated "best of" listicles (Expertise, a city magazine's "Top X"), and the specific editorial rankings worth going after in your field (Best Law Firms, Chambers). Capture which of these the model actually pulls for your category and city (Finding 5.4), then earn a spot on them.
- Be in good standing on the credential tier. The board changes by industry and state (Finding 6.6). Make sure you're correctly licensed, listed and clean, by name, on the exact regulator the model checks for your market.
- Confirm the trust tier. That's the BBB by exact profile URL in trades and financial services, and the relevant association directory in the professions (Findings 5.3 and 5.5). Make that profile accurate, accredited and well rated.
- Repeat for every service and every location. Every service and every city resolves to a different source list and a different set of winners (Findings 4 and 5). What gets a Dallas plumber recommended isn't what gets a Houston one recommended, and tree removal isn't tree trimming. Each one is its own job.
- Re-run on a schedule. The behavior changes from run to run and drifts with model versions. So re-capture every so often to catch new opportunities, hold your place on the shortlist, and stay current as new versions of ChatGPT ship.
Why you can't do this by hand
One ChatGPT answer isn't evidence. The output changes from run to run, so the same prompt run twice gives you different names, different sources and a different search chain.
We measured this directly. Regulated-industry verification swings between 74% and 96% across re-runs, roughly 1 in 5 regulated recommendations skip the license check on any given pass, and within a single market the leading business only recurs 78% of the time across reworded queries.
To get a stable read on even one market, you have to ask the same question many times, across phrasings and across the related sub-services. Then you have to do it again as the model updates. That's thousands of calls per market.
Looking at a few answers won't tell you anything reliable. And that's exactly the trap most "I checked ChatGPT and we're not there" reactions fall into.
Why anyone can do this now
The good news is you don't need special access for this, and you don't need a premium AI visibility tracking suite either.
The whole method is the model's own API, plus a capable general AI (Claude, for instance) to run the sampling, pull the full search chain out of each response, and turn the patterns into a per-market source list. The API returns the queries and citations, and the AI does the orchestration and the reading. Everything in this report was produced that way.
So anyone who's willing to capture it properly can see what the audit does. It's hidden from the answer itself, but it isn't secret.
Limitations
- Local-service businesses only. All ten industries are locally hired services. We didn't test SaaS, ecommerce, retail or other categories. The machinery should extend to them with different tiers (G2 and Capterra and security review for SaaS, marketplace ratings for ecommerce), but that's untested here. The unregulated IT vertical, which dropped the licensing tier for Clutch and CRN, is an early hint.
- One named engine, dated. This is ChatGPT (GPT-5.5) in June 2026, not "AI" in general. Other assistants weight sources differently, and behavior shifts across versions. ChatGPT is the dominant consumer assistant, on the order of roughly 900M weekly users, which is why it's the right single focus.
- Stochastic, and moving. We report key rates as ranges to account for run-to-run variation, and we didn't measure day-to-day index drift over weeks. The model and the specific businesses it cites will both keep changing as ChatGPT updates. So treat every named source and business here as a snapshot, because the method is the part that lasts, not any single name.
- Two figures are automated estimates, not hand counts, so treat them as approximate. One is the roughly 50/50 split between credibility sources and businesses in the citations. The other is the distinct-winner tallies in Finding 5.2, which may include a few directory pages that look like business sites.
Scope: 200 reasoning sessions on GPT-5.5 (plus sub-service and variance runs), plus a higher-volume ChatGPT pass for named-business, rank and winner data. 10 local-service industries, 10 states, 5,933 logged searches, 1,982 citations, 3,431 plus reverse-direction rank checks. ChatGPT, June 16 and 17, 2026.
