The Verification Gap

What AI actually says about Chiang Mai businesses, why the schema fix does not work, and where local authority is really built


Ask any assistant for the opening hours of a small business in Chiang Mai and the answer arrives in a confident, complete sentence. Check that sentence against the business itself and the confidence turns out to be unearned more often than the fluency suggests. The systems producing these answers have no mechanism for signalling the difference between a fact retrieved and a fact assembled, so a guess and a certainty leave the model looking identical.

That gap has now been measured, and the measurement contains an uncomfortable finding for every small operator in this city.

Small businesses absorb the error rate

Searchable, an AI visibility platform, ran more than 13,000 prompts through ChatGPT, Perplexity and Gemini about real London companies, comparing small and medium enterprises against firms of 500 employees or more. Every answer was checked against verified records including Companies House filings and official company biographies. Across the full sample, 93% of companies had at least one basic fact hallucinated or missing from an assistant’s answer.

The distribution of that error is the part that matters for anyone running a business here.

Half of the SMEs in the study received at least one outright fabricated fact, against 32% of large companies, a rate 56% higher. Assistants confidently stated false information about SMEs at 5% against 2% for large brands. Key information went missing entirely on 6.3% of SME queries and 4.8% of large company queries. Brand name confusion, where a model attributes one company’s details to another, occurred at 4% for SMEs and 0.7% for large brands, a fivefold difference. Services were described accurately 87% of the time for SMEs and 93% of the time for large brands. Aggregated across every prompt, 11 in 100 questions about SMEs returned false or missing information against seven in 100 for large companies.

The three categories the models handled worst were company size, contact details and founding year. Those are the facts a customer needs before they walk through a door.

The study covered London rather than Chiang Mai, and the specific percentages do not transfer directly to a Thai market with a different language mix and a different indexed footprint. The mechanism transfers completely. Local Falcon’s analysis of the same data puts the cause plainly: even a business with an accurate, well-optimised website struggles to surface reliably against a competitor carrying substantial third-party validation across the web. Models fabricate where the public record is thin, and the public record is thinnest around businesses that nobody other than the business itself has ever written about.

Thailand is further into this than most markets

The assumption that AI search is a Western phenomenon arriving here on a delay does not survive contact with the adoption data.

DataReportal’s Digital 2026: Thailand report counted 67.8 million internet users at the end of 2025, putting online penetration at 94.7%, alongside 56.6 million social media identities equal to 79.1% of the population. Social identities grew by 7.5 million in a single year, a 15.2% increase ranking among the fastest in the Asia-Pacific region. The typical Thai internet user spends around 34 hours and 32 minutes online each week.

Microsoft’s country-level adoption data records Thai AI usage growing 36.2% between the first half of 2025 and the first quarter of 2026, the second largest increase globally behind South Korea and roughly double the United States rate of 19%. The same analysis notes that adoption in low and middle income countries is growing at more than four times the rate of the wealthiest nations.

Platform concentration here is unusual. Statcounter data from December 2025 gives ChatGPT 74% of Thai AI chatbot traffic, above its global average, with Gemini at 16%, Perplexity at 7%, Copilot at 2% and Claude under 1%. Thailand also remains one of the most Google-dominant search markets in the world, with Google holding approximately 99.56% of Thai search traffic in early 2026. A Chiang Mai business therefore faces a narrow and specific exposure: ChatGPT for conversational discovery, and Google AI Overviews for everything else.

Thai businesses have adopted the technology enthusiastically. The UOB Business Outlook Study 2026, surveying 265 Thai business owners and senior decision makers, found more than seven in 10 Thai SMEs implementing AI and more than eight in 10 using digital tools, both above the regional average. Very few of them have audited what the technology says about them.

The Thai language problem nobody is costing

There is a second exposure specific to this market, and it receives almost no attention in the local marketing conversation.

Large language models perform measurably worse in lower-resource languages, a pattern documented well enough in the research literature to have its own name. Work presented at EACL 2026 on hallucination detection across languages found that task accuracy drops sharply when models move from English to lower-resource languages, across factual recall, STEM and humanities domains. A parallel study on factual accuracy in English versus low-resource languages found models frequently perform better in English even on questions rooted in local context, with a higher tendency toward hallucination in the local language.

Thai carries additional structural difficulty. The language has no spaces between words, which complicates tokenisation, and Thai discourse norms soften definitive statements in ways that introduce apparent uncertainty into otherwise factual content. Dedicated Thai-adapted models exist, including Typhoon and the SEA-LION family, yet none of them powers the assistants Thai consumers actually use.

The practical consequence for a Chiang Mai business is a two-sided risk. A business documented only in Thai carries a thinner English record for models that reason best in English. A business documented only in English disappears from Thai-language queries entirely. Businesses serving both a Thai domestic market and an expatriate or tourist market need their core facts stated cleanly in both languages, in visible text, on surfaces that get crawled.

The confidence gap runs in both directions

Executives trust these systems more than the systems have earned. Workiva’s 2026 Midyear Executive Benchmark Survey, published in August, polled 2,272 finance, risk and sustainability professionals including 847 C-level executives, alongside 367 institutional investors. It found that 84% of executives are at least somewhat confident in the accuracy of AI output without human review, broken down as 39% somewhat confident and 45% very confident, while 26% report that internal audits caught AI errors which reached external audiences or board members. Only 11% believe their organisation’s data quality is sufficient for AI use, and 89% of the institutional investors surveyed are concerned about the accuracy of AI-generated information in disclosures, with 47% actively looking for such errors.

Consumers have moved the other way. Fractl, working with Search Engine Land, surveyed 1,008 United States consumers and 150 marketers in the second quarter of 2026, repeating questions asked a year earlier. The share finding AI more helpful than traditional search fell from 82% to 54% in twelve months, a 28-point drop. The group actively rating AI as less helpful grew from 3% to 17%, close to six times larger. Over the same period, 70% reported using AI for search more than they did the year before, and only 3% reported using it less.

Two details in that dataset complicate the standard demographic assumptions. Baby Boomers now find AI search more helpful (63%) than Gen Z does (47%), meaning the cohort with the most exposure is the most disillusioned. And the share of consumers saying heavy brand use of AI would reduce their trust in that brand roughly doubled from 20% to around 40%.

Local discovery has shifted faster than general search. BrightLocal’s Local Consumer Review Survey 2026, a representative panel of 1,002 United States adults published in February, found consumers using AI tools to find local business recommendations rising from 6% to 45% in a single year. That makes AI the third most-used local discovery channel behind Google and Facebook, ahead of Yelp and Tripadvisor. Google’s own share of local business discovery fell from 83% to 71% across the same period. BrightLocal’s supplemental AI trust report breaks the figure down further: ChatGPT was used by 31% of consumers for business recommendations, Google AI Mode by 23%, and 64% of consumers aged 30 to 44 have asked AI for a business recommendation against 24% of those aged 60 and over.

What people do after the answer appears

Yext surveyed 3,848 consumers globally in 2026 about local search behaviour, with an AI-user subgroup of roughly 1,630 respondents. The headline number reshapes how visibility should be understood: only 5% of consumers move directly from an AI answer to a purchase.

Everyone else verifies. After receiving an AI recommendation, 53% run a Google or Bing search to confirm or learn more, 49% visit the business website directly, 42% click through to the sources the AI cited, 28% look for reviews on Google or similar platforms, and 20% check social profiles. Those rates hold almost constant regardless of stated trust levels. In the same survey, 75% of AI users rate their trust in AI local recommendations at four or five out of five, and 59% say that trust has grown over the past year, and they verify anyway.

The Yext data also establishes a spending threshold that determines where this matters commercially. The median amount consumers globally are comfortable letting an AI spend without review is 25 United States dollars. Every commercially meaningful local decision, from a hotel booking to a dental appointment to a contractor quote, sits above that line.

What tips the decision once a consumer starts verifying is review signal. Yext found star ratings on third-party platforms ranking first among purchase influencers after an AI recommendation, review recency second, word of mouth from a trusted contact third, and review content and sentiment fourth. BrightLocal puts hard floors under those signals: 47% of consumers will not use a business with fewer than 20 reviews, and 31% will only use one rated 4.5 stars or higher, up from 17% a year earlier.

One counterweight belongs in the record, because the verification story is often told too neatly. GatherUp research reported by Search Engine Land found that 67% of consumers do not rigorously fact-check AI sources before choosing a local business. The two findings coexist because verification is uneven rather than absent. Consumers check something, and what they check is rarely the full source trail. A wrong answer that survives a casual glance at a business website still costs the booking.

The schema correction

The standard remedy sold against this problem is structured data. Add JSON-LD markup, define the business as a machine-readable entity, and AI systems read the facts directly instead of guessing. The 2026 evidence does not support that claim, and the gap between the claim and the evidence has become the most expensive misunderstanding in local marketing.

Ahrefs published a controlled study in May 2026, authored by Louise Linehan and Xibeijia Guan. The team began with an analysis of 6 million URLs which found that pages cited by AI were almost three times more likely to carry JSON-LD. Rather than publish that correlation as a finding, they tested it. They identified 1,885 pages that added JSON-LD for the first time between August 2025 and March 2026, matched each against control pages from other domains with similar pre-existing citation levels, and measured citations across Google AI Overviews, Google AI Mode and ChatGPT for 30 days either side of the change.

Google AI Mode rose 2.4% and ChatGPT rose 2.2%, both statistically indistinguishable from zero. Google AI Overviews fell 4.6%, a decline the authors report as statistically real while declining to attribute it to the markup itself. Four separate tests pointed the same way, including an event study checking for pre-existing divergence between the two groups.

The explanation for the original correlation is straightforward. Structured data lives on well-maintained, technically competent sites, and those sites also publish better content, earn more references and rank better in conventional search. The markup was riding the signal rather than creating it.

Ahrefs cross-referenced a separate experiment by searchVIU, which tested five major systems, ChatGPT, Claude, Perplexity, Gemini and Google AI Mode, for real-time schema use when fetching a page. None of them read it. During direct retrieval, every system extracted only visible HTML and ignored JSON-LD, Microdata and RDFa entirely.

Google states the same thing in its own documentation, which says plainly that no new machine-readable files or markup are required to appear in AI features. John Mueller has said structured data will not make a site rank better and that most sites see no visible change from adding it.

Three honest caveats belong alongside this, and any consultant presenting the study without them is misrepresenting it. Every page in the Ahrefs treated group already carried 100 or more AI citations before the change, so the study answers what schema does for pages already visible rather than what it does for a business with no footprint at all. The study pooled all schema types together, so type-specific effects remain unisolated. And Microsoft’s Fabrice Canel confirmed on record in March 2025 that Bing’s systems do use structured data, making Bing the one platform with a first-party confirmation.

Schema remains worthwhile as indexing infrastructure, entity resolution and rich-result eligibility. It is hygiene rather than a growth lever. Any agency in this city selling markup as an AI visibility product in 2026 is selling against the published evidence.

Where the citations actually come from

Omniscient Digital analysed more than 23,000 AI citations on branded queries. A brand’s own website accounted for 23% of them. Earned sources accounted for 48%, and commercial content published by other companies accounted for 30%, with rounding taking the total marginally over 100. Reviews and other forms of social proof accounted for 57% of citations on branded queries, and the researchers found that educational content wins during early discovery while social proof wins when buyers approach a decision.

Roughly 77% of the citation surface sits outside the domain a business controls. The same directional finding appears in independent work: a Search Engine Land analysis with Evertune covering roughly 25,000 of the most-cited URLs found that across nearly 400 million citations, 63% pointed to ranked listicles rather than brand homepages.

This is the Searchable finding approached from the opposite end. Chris Donnelly, Searchable’s co-founder, attributed the SME error rate directly to footprint, noting that platforms train on publicly available web data which skews toward larger brands. His recommended starting point was building third-party validation through directories, reviews, articles and online communities. Searchable also cites Similarweb data indicating that an AI recommendation drives roughly twice the traffic to a brand’s website within seven days compared to a competitor which was not recommended.

The two datasets converge on one conclusion. The asset determining how an assistant describes a business is not the business website, and it is not the markup on the business website. It is the density and consistency of independent references to that business across surfaces the model treats as corroboration.

Small channel, high intent

The commercial case for this work rests on quality of traffic rather than volume, and the honest version of that case includes both halves.

Multiple independent measurements agree that AI referral traffic converts far above conventional organic search. Semrush established the widely cited cross-industry baseline at 4.4 times the conversion rate of standard organic. Ahrefs found internally that AI-referred visitors made up 0.5% of sessions and drove 12.1% of signups, a 23-times differential representing the upper bound. Rocket Agency’s 18-month cross-industry study found ChatGPT visits converting at 5.1 times organic. The mechanism is intent: the model has already defined the problem, synthesised the options and produced a shortlist before the visitor arrives.

The counterweight is volume. Conductor’s 2026 benchmarks put AI referral traffic at approximately 1.08% of all website visits, despite year-on-year growth of 527%. A conversion premium applied to one per cent of traffic produces a fraction of the revenue impact of ordinary conversion applied to forty per cent of traffic.

The pressure on the older channel is what changes the calculation. BrightEdge recorded AI Overviews triggering on around 48% of tracked queries by February 2026, up from roughly 30% a year earlier, while Seer Interactive measured organic click-through rates falling 61% on queries where an AI Overview appears. Seer also found that brands cited in AI Overviews earn 35% more organic clicks and 91% more paid clicks than uncited brands on the same queries. The AI channel remains small in absolute terms. The channel it is eroding is not.

A note on the figures in circulation

Several numbers travel widely in this field without a traceable source, and an article about machine-generated misinformation has an obligation to hold its own figures to the standard it asks of the machines.

Claims that schema makes content 36% more likely to appear in AI summaries, that it produces 2.5 times higher citation rates, or that businesses lose 60% of visibility without it, do not resolve to published methodology. As of March 2026 there were no peer-reviewed studies on the effect of schema on AI search visibility. The 2024 Princeton and Georgia Tech GEO paper cited constantly in schema discussions tested in-content strategies such as adding citations and statistics, which worked, and did not test schema at all.

Two further figures in wide circulation deserve flagging. The claim that a specific high percentage of local SMEs lack a GEO plan has no identifiable primary source. The documented version comes from Centerfield research shared with Search Engine Land, which found 63% of marketers saying their company is not investing time or budget in GEO, only 9% with resources allocated to it, and 33% claiming a good or expert understanding of GEO against 72% for conventional SEO.

Where a figure below carries vendor research rather than independent study, the vendor is named. Readers can weigh the incentive accordingly.

The audit any owner can run this week

The practical sequence follows from the evidence rather than from the sales pitch, and the first step costs nothing but an hour.

Interrogate the assistants directly. Ask ChatGPT and Gemini the questions a customer asks. Opening hours. Address. Phone number. Services offered. Price range. Ask in Thai and in English, because the answers diverge. Record every response and mark each error.

Trace each error to its source. Wrong facts almost always come from a stale listing, an aggregator nobody claimed, or a page that states the information inside an image, a PDF or a booking widget. Retrieval systems read rendered text and ignore everything else.

State the facts in visible text. Opening hours, address, phone number, services and pricing belong in readable sentences on the website, in both languages where both markets matter. Dated first-party statements give retrieval engines something specific and recent to cite.

Reconcile every external listing against that source. Conflicting addresses across secondary platforms give a model competing candidates and no basis for choosing between them, which is precisely the condition under which fabrication occurs.

Check that AI crawlers are not blocked. Many sites block OAI-SearchBot, PerplexityBot, ClaudeBot and Google-Extended through default security settings without the owner knowing.

Build presence on the surfaces carrying the citation weight. Directories, review platforms, local publications and community references supply the corroboration owned channels cannot generate alone.

Re-run the audit quarterly. Corrections take weeks to months to propagate, and new stale listings appear continuously.

Where Golden Pages fits

Golden Pages was built on the assumption that a verified directory listing does two jobs at once, and the 2026 data confirms both.

The first job is corroboration. A verified, human-checked listing adds an independent reference to a business at exactly the layer where roughly 77% of AI citations originate, and where thin-footprint SMEs currently lose to larger competitors by a factor of five on brand confusion alone. Searchable’s own recommendation to businesses appearing incorrectly in AI answers was to build third-party validation through directories.

The second job is verification. Of the consumers who receive an AI recommendation, 42% click through to the cited sources and 53% run a confirming search. A listing agreeing with the business, the reviews and the assistant closes that loop. A listing contradicting any of them breaks it, and the consumer moves to a competitor whose record is consistent.

Membership of the Chiang Mai Business Network adds the layer no crawler manufactures, which is a record of businesses that other businesses in this city have actually dealt with. Peer endorsement was the strongest signal in local commerce before assistants arrived. A technology that fabricates facts about small companies at 5% has raised its value rather than reduced it.

The next era of local discovery belongs to businesses whose facts agree with themselves everywhere a machine reads and everywhere a customer checks.

Leave a Reply

Your email address will not be published. Required fields are marked *