Skip to content

The Credibility LayerHow ChatGPT Recommends Local Businesses

We analyzed more than 1,300 ChatGPT recommendations across 10 industries and 10 US states to work out the off-site audit it runs before it names a business. It runs about 30 web searches across a three-tier credibility stack. Here’s what moves a recommendation, and what doesn’t.

200
Reasoning sessions
5,933
Logged searches
1,982
Citations captured
858
Distinct sources

Key Takeaways

  • Search rankings barely touch recommendations. 80% of recommended businesses rank nowhere in the top 50 of Google or Bing, only 30% of top-ranked businesses get recommended back, and Google and Bing agree on just 4% of ranked businesses. So you can get recommended without ranking, and ranking on its own isn’t enough to get you recommended.
  • Your own site gets you into the running, and third parties decide the rest. ChatGPT does read your pages. About half of all citations are businesses' own sites, mostly the homepage, and it uses them to confirm you exist and offer the service. But what separates the businesses it recommends from the rest is credibility on other people’s sites.
  • Reddit and user-generated content were cited zero times. Across 1,982 citations in the reasoning audit, Reddit didn’t appear once. So getting big on Reddit doesn’t do anything for the recommendation.
  • The platforms you pay for barely show up. Yelp, Angi, Thumbtack, HomeAdvisor and Avvo together account for 8 of 1,982 citations. The audit is looking for credibility it can check, and paid listings don’t give it that.
  • There is no universal playbook. ChatGPT cited 858 distinct sources. 95% appear in only one industry and 68% only once, and about 100 recommendations per industry produce 203 to 425 distinct winners. The sources that decide it are local, and specific to your industry.
  • What it actually does is run a credibility audit. About 30 web searches and seven sources go into a single recommendation, and they follow a universal three-tier stack: licensing, then authority ranking, then a trust directory. In regulated industries it checks credentials in 74 to 96% of sessions.

The report goes through these in order. It starts with the expensive habits that don’t move a recommendation, then the audit that does, with the full method and limitations behind every number.

Watch the video

ChatGPT For Local Businesses: How To Show Up

Methodology

We asked ChatGPT to recommend a local business the way a real buyer would. We left web search un-forced, and we recorded the full search chain, including the searches that never produced a citation.

Scale

LLM output changes from run to run, so we built the study to hold up anyway. There's a core measured set, then repeat runs and intent tests on top, then a higher-volume pass for the parts that need a lot of named businesses.

MeasureCount
Core recommendation sessions (GPT-5.5, un-forced)200
Industries × states × core prompts10 × 10 × 2
Variance re-runs (repeated measurement)120 (4 passes over 10 industries × 3 states)
Sub-service sessions (intent-robustness)30
Higher-volume pass for named-business dataabout 1,000
Total web searches logged5,933 (about 30 per session, max 58)
Reasoning tokens per session (avg)about 5,059
Citations captured1,982
Distinct sources cited858
Businesses named per answer (avg)about 8
SERP rank checks (both directions)3,431 plus reverse-direction

Process

Un-forced, full-chain capture. Most AI search studies force a single search and read the citation. That's clean, but it isn't what really happens. We let the model run its real multi-step process, so the data reflects what a paying customer actually triggers.

And we logged every query in the chain, including the ones that never produced a visible citation.

One assistant, two settings. Everything here is ChatGPT. The core study runs it the way a buyer gets it today, which is GPT-5.5 with reasoning and live web search.

One part of the analysis needs a lot more named businesses than the deep reasoning runs can cheaply produce. That's the rank tests (Finding 1) and the count of distinct winners per market (Finding 5.2). So for that part we ran a higher-throughput ChatGPT pass (gpt-5-chat-latest) to collect business names at volume.

Every behavioral finding (the audit, the credential checks, the citation and source mix) comes from the reasoning runs. We point out which is which where it matters.

Date (for reproducibility). Runs were on June 16 and 17, 2026, through the Responses API web_search tool, with reasoning effort set to high.

The prompts, verbatim ({X} = trade, {sub-service} = a specific job, {city, state} = location):

  1. "I'm looking for a reputable {X} in {city, state}. Help me find a good one I can hire."
  2. "Who are the best, most reputable {X}s in {city, state} that I can hire?"
  3. "I need {sub-service} in {city, state}. Who should I hire?" (a specific job with no "reputable" cue: car accident lawyer, water heater repair, dental implants, mortgage refinance, couples therapist, Botox, managed cybersecurity, tree removal, homeowners insurance, kitchen remodel). We used this one to test whether the audit only happens because we asked for "reputable."

Variance-tested. LLM output changes from run to run, so we re-ran cells several times and report the key behaviors as ranges, not single numbers.

Scope

This is ChatGPT specifically, on locally hired service businesses. We did not test other assistants or other business types (SaaS, ecommerce, retail). See Limitations.


Finding 1: Search rankings are decoupled from recommendations

Ranking isn't necessary and it isn't sufficient. 90% of recommended businesses don't rank top 20, and 70% of top-ranked businesses aren't recommended. This is the most misunderstood part of AI recommendations, so we're starting with it.

1.1 The mechanism: it never runs the query a ranking would answer

ChatGPT never runs the buyer keyword, so your rank for it can't be a direct input. Its searches average 8.4 words and go after named credibility sources. It doesn't run the short query a buyer types, only shortlist and verification queries.

Take "plumber in Dallas." A buyer types that into Google and a ranking decides what they see. But ChatGPT doesn't type it at all.

Its real logged queries for the same job look like the examples in Finding 6.3. There's a ranking lookup like Dallas plumbers Best Pick Reports top rated, a trust check like site:bbb.org/us/tx/dallas plumber accredited A+, then a license check like site:tsbpe.texas.gov "[company] Plumbing" responsible master plumber.

By then it has already picked candidates from the authority tier and is just confirming them. So your rank for "plumber in Dallas" never gets read.

(We observed the mechanism on the reasoning runs. The rank distributions in 1.2 and 1.3 use the higher-volume pass, so there are enough named businesses to check against. See Methodology.)

1.2 Recommended businesses don't rank, and the few that do sit on page 3

80% of recommended businesses don't rank anywhere in the top 50 of Google or Bing.

We took the businesses ChatGPT recommended and looked up where each one actually ranks for the main buyer keyword a customer would type ("personal injury lawyer Houston," "plumber Dallas," "managed IT services Chicago," "tree service Phoenix"). For the most part, the answer is nowhere:

Organic positionGoogleBing
Top 30%0.3%
4 to 104.1%0.9%
11 to 205.5%0%
21 to 5010.3%0%
Nowhere (>50)80.1%98.7%

Only 5% rank top 10 in either engine. The ones that do rank average position 23, which is page three.

1.3 And ranking isn't enough either (the reverse direction)

Of businesses that rank in Google's top 10 for the buyer keyword, only 30% get recommended.

We ran it the other way too. For each of those same buyer keywords ("dentist Phoenix," "mortgage broker Dallas," and so on), we took the businesses Google ranks in its top 10 and checked how often ChatGPT recommends them back:

IndustryTop-10-ranked → recommended
Medical spa39.8%
Therapist36.8%
Home renovation36.5%
Plumber36.2%
Insurance32.9%
Mortgage28.1%
Arborist27.7%
IT services26.0%
Lawyer21.7%
Dentist17.2%
Overall29.7%

And inside that top 10, rank position barely matters. If you pool all ranked businesses by position, the top 3 get recommended about 26% of the time and positions 4 to 10 about 24%, which is basically flat. Ranking first (31%) is no better than ranking eighth (30%).

So it isn't even true that ranking higher helps once you're on page one.

That means 70% of top-10-ranked businesses don't get recommended. That 30% is higher than chance alone would give you, because the model only names a handful of businesses out of the many in each market. But it's nowhere near the close to 100% you'd expect if rank drove recommendations.

The simplest explanation is that rank and recommendation share a cause further up, rather than rank being an input itself. A prominent business tends to both rank and land in the authority rankings. So ranking isn't a direct input to recommendations, and at best it's weakly correlated through prominence.

1.4 "Search rank" isn't even one thing: Google and Bing barely agree

Of businesses that rank top 50 on at least one engine, only 4% rank on both Google and Bing.

Where a top-50-ranked business appearsShare
On both Google and Bing4%
On one engine only96%

So the "ChatGPT runs on Bing, so optimize for Bing" idea falls apart here. The two engines show almost totally different businesses for the same query.

We pulled both engines through the same SERP API under matched conditions (same location, language, device and depth, de-personalized). So the gap reflects real differences between the engines, not session or geo personalization.

There's no single "ranking" to win. That fits with the model not using either engine's results page to choose who to recommend.

1.5 The contrast with citation studies

Search rank decides citations, but it doesn't decide recommendations. Ranking and recommendation overlap on 30% of cases at most, and only 5% of recommended businesses rank in any top 10.

Citation-focused research (like the widely cited query fan-out work) finds that retrieval rank strongly predicts whether a page gets cited in an informational answer. The "rank" in that work is a page's position inside ChatGPT's own retrieved results, not its Google or Bing ranking.

This study measures traditional Google and Bing organic rank, which is the thing businesses actually optimize. They're two different ladders, so that research doesn't conflict with this study, because it measures a different outcome.

Our data is about recommendation, and there Google and Bing rank doesn't predict who gets named. Measured in both directions, the overlap is weak:

Overlap between ranking and recommendationRate
Top-10 ranked → recommended30%
Recommended → ranks anywhere in top 5020%
Recommended → ranks in top 105%

Traditional rankings and recommendations are pretty close to decoupled. Rank can help your content get quoted, but it won't get your business handed to a buyer.

Key takeaways (Finding 1): the model never runs the buyer keyword. Recommended businesses mostly don't rank, ranked businesses mostly aren't recommended, and "search rank" splits across engines (4% Google and Bing overlap). So rank isn't necessary and it isn't sufficient.


Finding 2: Your own site gets you into the running, not recommended

About half of everything ChatGPT cites in the real audit is a business's own site (roughly 52% of the 1,982 citations, classifier-estimated). This is where the data corrected our first guess. We expected business websites to be marginal in the recommendation path, and they aren't.

The model really does read your pages. What matters is what it reads them for, and which pages.

2.1 Your site is cited about half the time, but the rate swings a lot by industry

Your own site is about half of every citation the model makes. But how much it leans on your pages swings by industry, from 10% (plumbers) to 74% (med spas).

IndustryCitations that are the business's own site
Medical spa74%
Therapist73%
Arborist70%
Lawyer68%
IT services54%
Dentist54%
Insurance44%
Mortgage34%
Home renovation23%
Plumber10%
Overallabout 52%

The model leans on your own site where third-party records are thin. Therapists, med spas and arborists have few directories, so it reads the businesses' sites directly.

And it barely touches your site where strong third-party records exist. A plumber recommendation is about 90% third-party: BBB, contractor boards, BuildZoom.

So your own site is a fallback the model uses when it has nothing better, rather than the source that decides it. (The business versus third-party split is classifier-estimated, so read these as directional.)

2.2 And it's mostly your homepage, used to confirm you exist and offer the service

When the model reads your site, it's mostly reading your homepage. Of the business pages we have full URLs for, about 70% were the homepage and 30% were a deeper service or location page.

IndustryHomepage (/)Deeper page (service or location)
Insurance80%20%
Home renovation78%22%
Arborist74%26%
Lawyer72%28%
Mortgage72%28%
Medical spa70%30%
Dentist69%31%
IT services68%32%
Therapist57%43%
Plumber55%45%
Overallabout 70%about 30%

The deeper pages are what you'd expect, service and location landing pages like rotorooter.com/newyork/, nycarborist.com/manhattan, blondiestreehouse.com/arborist-services/ and emeraldtreecare.com/about/meet-our-arborists/.

Trades that organize by service or city (plumbers and therapists) get their deeper pages read most often, up to 45% of the time. And the buyer's intent shifts it too.

A specific, urgent question ("my water heater just burst, who do I call?") pulled deeper service pages 39% of the time, against 15% for a generic "I need a plumber." So the more specific the job, the more the model reaches past the homepage.

But even then the homepage leads. Service and location pages do matter, because they're how the model confirms you serve this job in this place. The homepage is still the main check though.

2.3 But never on its own: every recommendation was credibility-checked

Your site alone is never enough, because every recommendation was also credibility-checked. Across the 196 sessions that produced a recommendation, all 196 ran a credibility check. Zero were standalone, and each one came with about 17 verification searches.

We checked every session for a credibility check (a verification search, or a citation to a third-party credibility source) in the same session where it read the business's pages. It always did.

Across the 196 sessions that produced a recommendationResult
Recommendations that also ran a credibility check196 of 196 (100%)
Standalone recommendations (named on own pages, no credibility check)0
Credibility-verification searches per recommendation (avg)about 17

Not one business got recommended on the strength of its own website alone. The homepage that gets cited 70% of the time is read in the same session as an average of about 17 credential and trust verification searches.

Your pages confirm you're real and relevant, but they never replace the audit. And no single business site recurs across recommendations the way the BBB or a state board does (Finding 5.1). So your pages confirm you, but they don't set you apart.

Getting your site right makes you eligible. It's still a long way from getting you recommended.

Key takeaways (Finding 2): your own site is about half of all citations, and the model really does read it, mostly your homepage, but never on its own. Every recommendation was also credibility-checked (0 standalone, about 17 verification searches each). Your pages make you eligible, and third-party credibility is what gets you recommended.


Finding 3: Reddit and user-generated content are not used

Reddit was cited zero times in the 1,982-citation audit. "Get big on Reddit" has become almost as common a piece of AI visibility advice as "get more brand mentions." But in the mode that actually recommends businesses, it does nothing.

3.1 Zero Reddit citations in the real audit

Reddit was cited zero times in the real audit, out of 1,982 citations. In a forced single-search pass it shows up 171 times.

PassReddit citationsTotal citationsReddit share
Real reasoning audit01,9820%
Forced single-search pass17113,6551.3%

In the reasoning audit, Reddit was cited zero times. The only user-generated or social citations of any kind were two LinkedIn business profiles. Quora, Facebook and forums got zero.

3.2 Reddit only shows up when the model searches shallow

Reddit only shows up when the model searches shallow: 171 citations in the forced single-search pass, against 0 in the real audit. A one-shot search lands on whatever ranks for the query. Reddit threads rank well, so they get pulled in.

The real reasoning audit never works like that. It sends targeted queries to named credibility sources (boards, associations, the BBB), and a Reddit thread isn't any of those. So Reddit citations come from shallow retrieval, and the model doesn't use them to recommend.

This matters because simple checks and a lot of AI visibility tools rely on exactly that kind of one-shot call. That's why they over-report how much Reddit matters.

Key takeaways (Finding 3): the recommendation audit cited Reddit zero times. It only shows up (171 times) in a shallow single search, and that isn't how recommendations get made. Being active on forums doesn't move a recommendation.


Finding 4: The platforms you pay for are bystanders

The five platforms businesses most often pay for add up to just 8 of 1,982 citations. The directories you budget for barely register in the audit.

4.1 Paid platforms in the real audit

The big five paid platforms (Yelp, Angi, Avvo, Thumbtack, HomeAdvisor) add up to just 8 of 1,982 citations.

PlatformCitations (of 1,982)
Healthgrades9
Yelp6
Angi2
Trustpilot2
Thumbtack / HomeAdvisor / Porch0
Avvo / Justia / FindLaw0
Nextdoor / Bark / Networx0

Healthgrades, a medical directory, is the most-cited paid platform at 9, and that's still negligible. Vitals, Martindale and Lawyers.com are at or near zero.

4.2 They show up in shallow searches, like Reddit

Like Reddit, paid platforms mostly show up in shallow searches. HomeAdvisor goes from 0 in the real audit to 22 in the forced pass, and Yelp goes from 6 to 12.

PlatformReal auditForced single-search pass
Healthgrades926
HomeAdvisor022
Yelp612
Angi26
Avvo03

Paid directories show up several times more often in the shallow forced pass, where the model grabs whatever ranks, than in the real audit, where it's checking credibility. These are advertising marketplaces, and the audit isn't looking for ads.

So paying to appear on a platform the audit never checks doesn't touch the recommendation. It might still drive direct traffic, but that's a separate question.

Key takeaways (Finding 4): the paid directories businesses budget for are almost entirely missing from the audit (8 of 1,982 for the big five), and what little shows up comes from shallow searches. Spending on these platforms doesn't buy you a recommendation.


Finding 5: There is no universal playbook

ChatGPT cited 858 distinct sources. 95% show up in only one industry, and 68% were cited only once. The sources that decide recommendations are specific and fragmented, and mostly nobody is optimizing for them.

5.1 The sources are very industry-specific, and mostly one-off

68% of all cited sources showed up exactly once, and the average source spans just 1.1 of 10 industries. The only sources that cross industries are generic trust and government sites, nothing you'd "market into":

Cross-industry sourceIndustries (of 10)Type
mass.gov8govt umbrella portal
bbb.org7trust directory
pa.gov6govt umbrella portal
content.boston.gov5govt portal
expertise.com5listicle
(everything else)4 or fewer. 95% appear in exactly 1, 68% cited only onceindustry-specific

The top 20 sources are only 34% of all citations, so there's a long, fragmented tail. That doesn't contradict the "same stack everywhere" idea in Finding 6.

The structure is universal: discover, verify, trust, the same in every market. What fills each tier isn't. The exact sources change by industry, by state, and even by service inside an industry.

So the shape stays the same and the contents are local. That's why there's no list to copy, only a process you run for each market.

5.2 The winners are hyper-local too: it's 100 separate local races

Pool about 100 recommendations per industry and you get 203 to 425 distinct businesses, so there's no national winner.

We also looked at who gets recommended as well as the sources, and the winners are as fragmented as the sources are. This winner count uses the higher-volume pass (about 100 recommendations per industry), because counting distinct winners needs a lot more named businesses than the deep reasoning runs produce.

SignalResult
Distinct businesses recommended per industry (about 100 recommendations)203 to 425
Most-recommended business's share of an industry20% at most (no dominant national winner)
Within a single city, how often the leading business recurs across differently-worded queries78% (sticky per market)
Exception, a national brand winning across marketsRoto-Rooter (plumbers), 20 of 100

So AI recommendation is basically about 100 separate local races. Within one market the leader is locked in (78% consensus across phrasings). But across markets the winners are all different, hundreds of them, because the audit resolves locally.

National brands can sometimes win broadly (Roto-Rooter), but most verticals have purely local winners. So the winners don't fit a generic playbook any better than the sources do.

5.3 BBB is the one near-universal trust anchor

One source, the BBB, is 13% of every citation in the study (261 of 1,982). It dominates where a profession has no prestige system of its own, and it disappears where one exists:

IndustryBBB citationsHas its own prestige ranking?
Plumber81no
Insurance62no
Mortgage56no
Home renovation51no
Arborist8no
IT services2yes
Lawyer / Dentist / Therapist0yes

It checks the BBB by exact profile URL, by name (site:bbb.org/us/tx/houston/profile/plumber … A+). So it uses the profession's own authority where one exists, and the BBB where it doesn't.

5.4 Listicles, editorial press and media coverage: a top lever in some industries, irrelevant in others

Editorial press (city magazines, "best of" listicles, trade press) is 13.5% of citations overall, but that average hides a split. It shows up in 85% of lawyer and IT sessions and 0% of therapist and arborist sessions.

IndustrySessions with an editorial source (of 20)Dominant editorial sources
Lawyer17Best Law Firms, Chambers, Best Lawyers, Forbes
IT services17Clutch, CRN, Channel Futures
Dentist13Seattle Met, Philly Mag, Houstonia, Castle Connolly
Plumber5Best Pick Reports, Expertise
Medical spa4city magazines, RealSelf
Mortgage4Expertise, WalletHub
Insurance2n/a
Home renovation2NY Magazine, Houzz
Therapist0(runs on associations instead)
Arborist0(runs on certifications instead)

So why the split? Editorial press is how ChatGPT fills the authority-ranking tier (tier two).

Fields with a strong third-party ranking culture give it plenty to cite. Law has Best Law Firms and Chambers, IT consulting has Clutch and CRN, and dentistry has city "Top Dentists" features.

Fields that run on membership and certification instead (therapy on APA and Gottman, arboriculture on ISA and TCIA) don't give it anything like that to cite. So it falls back to those directories.

Unregulated IT shows a second effect on top of that. With no licensing board to check, listicles like Clutch become the stand-in for the credibility the license tier would normally provide.

So it isn't simply "no regulator means more press," because law has a strong regulator and heavy press. What drives it is whether the field has a public ranking and press culture at all.

The authority tier is really two separate things, and in SEO they aren't interchangeable. There are review aggregators (Clutch, RealSelf, Zillow), which rank businesses by collected ratings and reviews. And there are editorial "best of" rankings and listicles (Best Law Firms, Chambers, Expertise.com, a city magazine's "Top Dentists"), where a publication curates the list using its own methodology.

Which one the model leans on depends entirely on the field:

SourceTypeCitationsWhere it carries
Clutchreview aggregator67IT services (almost only)
Best Law Firmseditorial ranking66Law
City magazines ("Top X")editorial listicle47Dentist, med spa
Expertise.comeditorial listicle18Plumber, mortgage
Best Lawyerseditorial ranking15Law
Chamberseditorial ranking12Law
Forbeseditorial media8Law
RealSelf / Zillowreview aggregator1 eachmed spa, mortgage

Editorial rankings and listicles carry most of the citations (about 170), and review aggregators get far fewer (about 69, almost all Clutch). Aggregators own exactly one field here, IT services, where Clutch is basically the category ranking. Everywhere else, editorial rankings and listicles dominate.

So the play depends on your vertical. Go after the aggregator if you're in IT (Clutch), and go after the editorial ranking or listicle otherwise (Best Law Firms, a city "Top X" feature).

A placement in the right one for your field is a direct recommendation lever. In therapy or tree care, it's close to wasted effort.

5.5 Professional associations are the biggest hidden layer

Associations are the main authority tier for fields driven by certification and membership, like arborists, therapists and remodelers, and almost no one markets to them. For arborists, the association directories (treesaregood and TCIA, 21 citations) outrank the firms themselves.

Association (citations)Industry it carries
NARI: nariatlanta (18), greaterphoenixnari (9), nari.orgHome renovation
ISA / TCIA: treesaregood (11), treecareindustryassociation (10)Arborist (top sources, above the firms)
TrustedChoice / Big "I" (8)Insurance
ADA: findadentist (4)Dentist
Gottman (3), ABCT (3), APA locatorTherapist

Recognized membership is a credibility signal the model can verify, and it treats the association directory as a trusted shortlist. So active, listed membership does a lot of work here, and anyone optimizing for Google or Yelp won't even see it.

Key takeaways (Finding 5): sources and winners are both fragmented (95% single-industry sources, 68% one-off, 203 to 425 distinct winners per vertical). BBB is the one near-universal lever, editorial press is a big but uneven PR lever, and associations are the biggest hidden layer. The stack is different for every industry and state.


Finding 6: How it really works, the credibility audit

Every recommendation is roughly a 30-search investigation. This is the credibility layer in full, and it's what everything above points to.

When a buyer asks ChatGPT to recommend a provider, it doesn't read a rankings page and copy the top results. It runs a credibility audit. Across 200 sessions it averaged around 30 web searches and about 5,059 tokens of reasoning per recommendation, and it checked roughly seven distinct sources.

It works in a consistent order: find candidates, verify their credentials, cross-check a trust directory, then name eight or so businesses. The shape was the same in every industry.

6.1 A recommendation is a 30-step search process

It runs about 30 searches per recommendation, and two-thirds of them are hidden verification that never turns into a visible citation.

MetricValue
Searches per sessionavg about 30, max 58
Internal reasoning per sessionabout 5,059 tokens
Distinct sources consulted per sessionabout 7
Businesses named in the final answeravg 8.2, max 13
Searches per citation producedabout 3 to 1

Roughly two-thirds of the searches, nearly all credential checks, never show up as citations. And about half of what it does cite is a regulator, ranking or directory rather than a business (about 50%, classifier-estimated).

There's a mild tendency in the ordering. Discovery and ranking queries lean slightly earlier in the chain than credential and trust queries (mean normalized position 0.41 versus 0.49). It's a small gap, so read it as a lean toward "discover first," not a clean two-phase sequence.

Either way, you learn more from the search chain than from the citation list.

6.2 Every industry runs the same three-tier stack

It's the same three tiers in all 10 industries. Only what's in them changes.

Industry① Licensing / credential② Authority ranking③ Trust directory
LawyerState Bars (calbar, texasbar…)Best Law Firms, Chambers, Super Lawyersbar referral
PlumberContractor boards (CSLB, AZ ROC, WA L&I) + permitsExpertise, Best Pick ReportsBBB
DentistDental boards + ADA find-a-dentistCity "Top Dentists" magazines, Castle ConnollyZocdoc
Mortgage brokerNMLS + CFPB + state DFI/DRE/DFSExpertise, ZillowBBB
TherapistState licensing + APA/ADAA locatorsn/aPsychology Today
Medical spaFDA + state medical/cosmetology boardsRealSelf, city magazinesBBB
IT servicesn/a (unregulated)Clutch, CRN, Channel FuturesMicrosoft Partner
ArboristISA / TCIA / ASCA certificationn/aBBB
Insurance agencyState Insurance Depts (CA DOI, TX TDI)n/aBBB, TrustedChoice
Home renovationContractor boards + NARIHouzz, NY Magazine "best contractors"BBB, BuildZoom

The tiers adapt. Unregulated IT drops the licensing tier, and therapy and insurance, which have no prestige ranking, lean on associations and trust directories instead.

6.3 The model's queries are nothing a human would type

Its searches average 8.4 words (a human's is about 4), and 18% use a site: operator. These are machine-shaped queries, and no human would type them.

Query traitValue
Average length8.4 words
Longer than 6 words77%
Uses a site: operator18%
Verifies a name in quotes14%
Carries a recency year (2023 to 2026)9%
Broad "brand buzz" sweeps0

Verbatim examples:

site:mqa-internet.doh.state.fl.us "Leslie Haller" "Dentist" "Coral Gables"
Abraham Watkins Houston personal injury attorneys board certified Best Law Firms Tier 1
https://data.wa.gov/resource/m8qx-ubtq.json?$select=businessname   (a raw state open-data API endpoint)

To check a Washington contractor, ChatGPT queried a government open-data API directly. It makes these queries up on the fly, and there's basically no limit to how varied they get.

You can't optimize for a query you can't see. And you can't see these unless you capture them from real ChatGPT calls.

6.4 The audit holds without asking for "reputable"

Take out the "reputable/best" cue and credential verification still fired in 93% of sub-service sessions.

Sub-service queryCredential-verifiedWhat it specialized to
dental implants✓TX dental board, UCLA dentistry, Mount Sinai
Botox✓Mount Sinai, Weill Cornell, CDC
couples therapist✓Gottman Referral Network
mortgage refinance✓CFPB, Freddie Mac
tree removal✓TCIA, city forestry
homeowners insurance✓NY DFS, CA DOI, TX TDI
car accident lawyer✓State Bar, courts
water heater repair✓city permits, BBB

93% of sub-service sessions (28 of 30) verified credentials, at 24.4 searches each. So the audit holds up whatever the buyer's intent is, and it isn't something our wording caused.

6.5 It verifies against the specific licensing body for each state

Credential verification is the main behavior across the study, firing in 74 to 96% of regulated-industry sessions. It checks about six candidates by name per session, and it checked a license by name in 40% of sessions.

MeasureResult
Sessions touching a government/regulator source90% (98% of regulated industries)
Named-individual license checks (site:[board] + a name)40%
Distinct candidates verified by name per sessionavg 5.5, max 51
Regulated-industry verification (variance-tested)74 to 96% (mean 87%)
Regulated markets verifying at least once100%

The credential check is the main behavior, and it keeps coming back. Every regulated market triggers it, but it doesn't happen 100% of the time. On any given run, roughly 1 in 5 regulated recommendations skip the explicit license step.

So the behavior is reliable but the rate moves around, which is why we report a range. For health verticals it also reaches academic and clinical institutions (about 39 .edu citations: UCLA/USC dentistry, Mount Sinai for med-spa).

What we measured is a query pattern. Lawyers route to the State Bar, mortgage brokers to NMLS, and contractors to the state registry. The data supports the functional claim, but the cognitive claim ("it knows") is our interpretation.

6.6 The right board appears for each state

Each industry surfaced 16 to 31 distinct regulators, and the right state's board showed up in most runs.

Add this state……and the model verifies against
Arizonaazbar.org, AZ Registrar of Contractors, azdentalboard.us
Floridafloridabar.org, myfloridalicense.com, FL Dept. of Health
Washingtonwsba.org, secure.lni.wa.gov, doh.wa.gov
Texastexasbar.com, Texas Board of Legal Specialization, tsbpe.texas.gov

In some states one portal covers lots of trades. mass.gov was the verification source across 8 different industries.

6.7 Precision beats volume

In 5,933 searches, ChatGPT ran zero broad "what's-being-said-about-this-brand" sweeps. Every search is a targeted query against a named source.

And when it does search for a specific business by name, it never searches the name on its own. Every branded query in the sample paired the name with a verifier: a site: check against a licensing board, a BBB profile lookup, or a qualifier like "Best Law Firms Tier 1" or "complaints."

Query behaviorOf 5,933 searches
Broad brand-buzz sweeps ("what's being said about X")0
Bare-brand exploration (a name with no verifier)0
Names a specific business, always with a verifier attached14%

So "get more brand mentions everywhere" misses the point. A mention on fifty mid-authority sites is on zero of the sources the audit checks, while the one right source (your board, association or category ranking) gets checked directly.

The only way the shotgun approach works is by seeping into a future model's training data. That's a bet that pays off at the next model cycle (months to years out). It doesn't help with today's buyer, who gets served by a live audit that mentions don't touch.

Key takeaways (Finding 6): each recommendation is a 30-search investigation, and two-thirds of it is hidden verification. It follows a universal three-tier stack, verifies against the specific state licensing body, and uses precise, machine-shaped queries that never search a brand name on its own.


Everything above covers what doesn't move a recommendation. Here's what does, in the order the audit cares about.

The deciding sources are specific, fragmented and per-market. So there's no list you can copy, and you have to run this as a process for each market you serve.

  1. Cover the basics so you're in the running. ChatGPT reads your own pages to confirm you exist, offer the service and work in the market, and that's about half of what it cites (Finding 2). A homepage, a page per service and a page per location clear this bar, and most sites already have them. This won't get you recommended, but it makes you eligible.
  2. Win the authority tier. This is where the work pays off most, and it's per-market. It's a mix of review aggregators where they carry weight (Clutch for IT services), curated "best of" listicles (Expertise, a city magazine's "Top X"), and the specific editorial rankings worth going after in your field (Best Law Firms, Chambers). Capture which of these the model actually pulls for your category and city (Finding 5.4), then earn a spot on them.
  3. Be in good standing on the credential tier. The board changes by industry and state (Finding 6.6). Make sure you're correctly licensed, listed and clean, by name, on the exact regulator the model checks for your market.
  4. Confirm the trust tier. That's the BBB by exact profile URL in trades and financial services, and the relevant association directory in the professions (Findings 5.3 and 5.5). Make that profile accurate, accredited and well rated.
  5. Repeat for every service and every location. Every service and every city resolves to a different source list and a different set of winners (Findings 4 and 5). What gets a Dallas plumber recommended isn't what gets a Houston one recommended, and tree removal isn't tree trimming. Each one is its own job.
  6. Re-run on a schedule. The behavior changes from run to run and drifts with model versions. So re-capture every so often to catch new opportunities, hold your place on the shortlist, and stay current as new versions of ChatGPT ship.

Why you can't do this by hand

One ChatGPT answer isn't evidence. The output changes from run to run, so the same prompt run twice gives you different names, different sources and a different search chain.

We measured this directly. Regulated-industry verification swings between 74% and 96% across re-runs, roughly 1 in 5 regulated recommendations skip the license check on any given pass, and within a single market the leading business only recurs 78% of the time across reworded queries.

To get a stable read on even one market, you have to ask the same question many times, across phrasings and across the related sub-services. Then you have to do it again as the model updates. That's thousands of calls per market.

Looking at a few answers won't tell you anything reliable. And that's exactly the trap most "I checked ChatGPT and we're not there" reactions fall into.

Why anyone can do this now

The good news is you don't need special access for this, and you don't need a premium AI visibility tracking suite either.

The whole method is the model's own API, plus a capable general AI (Claude, for instance) to run the sampling, pull the full search chain out of each response, and turn the patterns into a per-market source list. The API returns the queries and citations, and the AI does the orchestration and the reading. Everything in this report was produced that way.

So anyone who's willing to capture it properly can see what the audit does. It's hidden from the answer itself, but it isn't secret.


Limitations

  • Local-service businesses only. All ten industries are locally hired services. We didn't test SaaS, ecommerce, retail or other categories. The machinery should extend to them with different tiers (G2 and Capterra and security review for SaaS, marketplace ratings for ecommerce), but that's untested here. The unregulated IT vertical, which dropped the licensing tier for Clutch and CRN, is an early hint.
  • One named engine, dated. This is ChatGPT (GPT-5.5) in June 2026, not "AI" in general. Other assistants weight sources differently, and behavior shifts across versions. ChatGPT is the dominant consumer assistant, on the order of roughly 900M weekly users, which is why it's the right single focus.
  • Stochastic, and moving. We report key rates as ranges to account for run-to-run variation, and we didn't measure day-to-day index drift over weeks. The model and the specific businesses it cites will both keep changing as ChatGPT updates. So treat every named source and business here as a snapshot, because the method is the part that lasts, not any single name.
  • Two figures are automated estimates, not hand counts, so treat them as approximate. One is the roughly 50/50 split between credibility sources and businesses in the citations. The other is the distinct-winner tallies in Finding 5.2, which may include a few directory pages that look like business sites.

Scope: 200 reasoning sessions on GPT-5.5 (plus sub-service and variance runs), plus a higher-volume ChatGPT pass for named-business, rank and winner data. 10 local-service industries, 10 states, 5,933 logged searches, 1,982 citations, 3,431 plus reverse-direction rank checks. ChatGPT, June 16 and 17, 2026.

Get the tools behind this research.

The workflows and teardowns are inside the members community.

See what is inside

Run credibility audits at scale.

Omnipresence captures the live ChatGPT audit for your client's industry and city, then puts your own agent on the sources it actually reads. Apply to see if your agency is a fit.

Start your 7-day free trial