We Asked 4 AI Engines 50 Hong Kong Business Questions: Accuracy, Language and the Sources They Cite
Published · 中文原文
The short answer
In our 22 August 2026 test of 50 Cantonese-phrased Hong Kong business questions, Gemini and DeepSeek each answered 8 of 10 fact-checkable questions correctly, Perplexity 7 with no wrong or outdated answers, and ChatGPT 4 with 2 wrong. Only Perplexity, the one engine with live web search in our setup, gave the current HK$43.1 minimum wage. Its 480 source citations were led by HelloToby (17) and HK01 (15), not by the websites of the businesses being recommended.
Why did we run this test?
Hong Kong customers can now ask an AI assistant which moving company to hire or what the minimum wage is. Research on how AI engines answer Hong Kong questions, asked in colloquial Cantonese the way people actually type them, is hard to find. We had not seen a published benchmark of this kind, so we built one.
Declaration of interest: this study was designed and run by Webka, a Hong Kong company that sells AI search visibility services. That is why the full method is below, every engine is reported under the same rules, and we show where our own automated judge made mistakes. This study does not measure, and we do not claim, that any AI engine cites Webka.
This is the English edition of a study first published in Traditional Chinese. Every number was recomputed from the underlying data files; where the editions differ, this one follows the data and says so under limitations.
How was the study run?
In short: 50 questions, 4 engines, 200 answers, one test date, one automated judge, and a human fact-check of every verdict on the 10 questions with an official answer.
The 50 questions
All 50 questions were written in colloquial Cantonese (for example, 香港搬屋公司邊間好?大約幾錢? — "Which moving company in Hong Kong is good? Roughly how much?"). They fall into four groups: 20 local service recommendations (group A), 10 food and restaurant recommendations (group B), 10 business-practice questions with a verifiable official answer (group C), and 10 local consumer-knowledge questions (group D). The translated list is at the end of this article.
Only group C has a ground truth written into the question file before judging, so only group C is scored for factual accuracy.
The four engines and settings
Each question was sent once to each engine through its API on 22 August 2026, with no system prompt or history. ChatGPT: OpenAI gpt-4o-mini (response reported gpt-4o-mini-2024-07-18), temperature 0.3, max 1,024 tokens. Gemini: gemini-2.5-flash, default settings, no Google Search grounding. Perplexity: sonar-pro via OpenRouter, the only engine here that searches the live web and returns cited URLs. DeepSeek: deepseek-chat endpoint (response reported deepseek-v4-flash), temperature 0.3, max 1,024 tokens.
Why DeepSeek and not Microsoft Copilot? Copilot has no public API we could call on equal terms, and DeepSeek is available to Hong Kong users. The models match those our own diagnosis pipeline used at the time.
How answers were judged
Each of the 200 answers was scored by a judge model (gpt-4o-mini, temperature 0) against a fixed JSON schema: named brands and organisations, suspected fabricated names, language register (Cantonese, written Traditional Chinese, Simplified, mixed), whether a concrete HK$ price was given, and, for group C only, a verdict of correct, partial, wrong or outdated.
We then checked every group C verdict by hand against primary government sources. Two corrections changed four verdicts. On the minimum-wage question, our own ground truth was out of date: it said HK$42.1, but the rate rose to HK$43.1 on 1 May 2026, so Perplexity's HK$43.1 was re-marked correct and ChatGPT's HK$39 was re-marked wrong. On the profits-tax question the judge had written a self-contradicting note, and Gemini's and DeepSeek's fully correct answers were re-marked correct. All figures in this article are after these corrections.
The judge's 0–5 scores for "Hong Kong specificity" and "usefulness" are not used: every engine averaged between 4.68 and 4.98 on both, a ceiling effect that separates nothing.
Which AI engine gets Hong Kong business facts right?
On the 10 business-practice questions (profits tax, MPF, business registration, incorporation, BUD Fund, holidays, minimum wage, severance and more), Gemini and DeepSeek tied on 8 correct answers each, Perplexity got 7, and ChatGPT got 4.
The more useful pattern: Perplexity was the only engine with zero wrong and zero outdated answers; its three misses were all partial answers. Gemini and DeepSeek each had one outdated answer, and it was the same question. ChatGPT had four partial answers and two outright wrong ones.
- Gemini — 8 correct, 1 partial, 0 wrong, 1 outdated
- DeepSeek — 8 correct, 1 partial, 0 wrong, 1 outdated
- Perplexity — 7 correct, 3 partial, 0 wrong, 0 outdated
- ChatGPT — 4 correct, 4 partial, 2 wrong, 0 outdated
The minimum-wage question
Hong Kong's statutory minimum wage has been HK$43.1 an hour since 1 May 2026, per the government's announcement. Perplexity answered HK$43.1 and cited that government press release. Gemini and DeepSeek both answered HK$40, the rate that took effect in May 2023 — outdated by two revisions. ChatGPT answered HK$39 and said it took effect in May 2019. Since 2019 the rate has moved from HK$37.5 to HK$40 to HK$42.1 to HK$43.1; HK$39 has never been the rate. That is not an outdated answer, it is an invented one.
The BUD Fund question
Asked for the cumulative funding cap per company under the BUD Fund, Gemini, Perplexity and DeepSeek all gave the correct HK$7 million. ChatGPT said HK$5 million.
Citable finding: on the minimum wage, revised on 1 May 2026, all three engines without live search gave an old or invented rate; only the search-connected engine was current.
Which AI engine answers Cantonese in Cantonese?
How often did a Cantonese question get a Cantonese answer rather than formal written Chinese? Perplexity replied in Cantonese in 40 of 50 answers (80%). DeepSeek did so in 20 of 50 (40%), Gemini in 6 of 50 (12%) and ChatGPT in 2 of 50 (4%). Every other answer was in written Traditional Chinese; none of the 200 answers was classified as Simplified Chinese or mixed.
Citable finding: asked in Cantonese, ChatGPT answered in Cantonese 2 times out of 50; Perplexity was the only engine that answered in Cantonese most of the time.
How specific are the answers?
Does an answer name real Hong Kong businesses and give real prices, or stay generic? Across all 50 answers per engine we counted named brands and organisations, and the share of answers giving a concrete HK$ price.
DeepSeek named the most, averaging 7.26 named entities per answer, and gave a concrete price in 56% of answers (28 of 50). Gemini averaged 6.36 entities and priced 46% (23 of 50). Perplexity averaged 5.04 entities and priced 38% (19 of 50). ChatGPT averaged 3.86 entities and priced 22% (11 of 50).
The judge flagged no suspected fabricated business names in any engine's answers. That means none detected by an automated check, not proof that none exist.
Citable finding: asked the same 50 Hong Kong questions, DeepSeek named nearly twice as many specific organisations per answer as ChatGPT (7.26 vs 3.86).
Which sources does Perplexity cite for Hong Kong questions?
Perplexity is the only engine here that shows its sources. Across the 50 answers it returned 480 citations pointing to 266 distinct domains.
The most-cited domains were service marketplaces, media, government, consumer bodies and reference sites. The one retailer near the top, Fortress (fortress.com.hk), is named in the question set itself: one question asks whether Fortress, Broadway or online shopping is cheaper for appliances. Tenth place is a four-way tie at 6 citations each: moneyhero.com.hk, hk.trip.com, weekendhk.com and labour.gov.hk. Domains ending in .gov.hk together accounted for 36 of the 480 citations (7.5%).
A smaller warning sign: at least 10 of the 480 citations (about 2%) went to domains carrying a Taiwan or mainland China marker — a .tw or .cn domain, or a tw. or cn. subdomain — such as tripadvisor.com.tw and tw.trip.com. On Hong Kong questions, the engine sometimes filled gaps with content written for another market.
Citable finding: across 50 Hong Kong questions, the single domain Perplexity cited most was a service marketplace (HelloToby, 17 citations), ahead of any single media or government domain.
- 1. hellotoby.com — 17 citations
- 2. hk01.com — 15
- 3. ird.gov.hk (Inland Revenue Department) — 10
- 4. consumer.org.hk (Consumer Council) — 9
- 5= pro360.com.hk — 8
- 5= ufood.com.hk — 8
- 7= zh.wikipedia.org — 7
- 7= apps.apple.com — 7
- 7= fortress.com.hk — 7
- 10= moneyhero.com.hk, hk.trip.com, weekendhk.com, labour.gov.hk — 6 each
What does this mean for AI search optimization in Hong Kong?
The practice of getting a business into AI-generated answers goes by several names: Generative Engine Optimization (GEO), AI search optimization, or Answer Engine Optimization (AEO). In Hong Kong the last acronym is ambiguous — AEO is also the Customs and Excise Department's Authorized Economic Operator programme — which is why we use GEO or AI search optimization in titles. Whatever it is called, this data points to three practical conclusions.
1. Be present where the engines already look
The sources Perplexity cited most were marketplaces, media, government and consumer bodies, not the websites of the businesses being recommended. A business absent from the platforms in its category starts weak however good its website is. The first step is checking which platforms appear in AI answers for your category, and whether you are listed there.
2. Dated, current facts are an opening
Three of four engines missed the current minimum wage; the one that got it right cited a page stating the new figure and its effective date. A page that states current figures clearly, with a date and a source, is what a search-connected engine can quote.
3. Test each engine separately, in Cantonese
The engines differed on every measure we took. A single blended score would hide those gaps. Test the questions your customers ask, in the language they ask them, on each engine separately, and repeat on a fixed schedule because models and indexes change.
What are the limitations of this study?
- One run, one date (22 August 2026). AI answers vary between runs; a repeat could shift individual verdicts.
- API models, not consumer apps. The ChatGPT and Gemini apps can use different models and live search; Gemini here had no Google Search grounding.
- Small factual sample: 10 questions per engine. Directional, not a precise error rate.
- Automated judge. Language, entity counts and price mentions come from a gpt-4o-mini judge; only group C verdicts were hand-checked.
- Citations from one engine: the source ranking reflects Perplexity only, counted by domain.
- Cross-border source count is a floor. The "at least 10" figure uses a mechanical rule (.tw/.cn domain or tw./cn. subdomain). The Chinese edition reports 16 using a broader manual classification; we use the stricter, reproducible count here.
- Model labels here are the IDs recorded in the API responses (e.g. deepseek-v4-flash), which differ from the Chinese edition's labels.
Can I cite this study?
Yes. Please cite it as: Webka, "We asked 4 AI engines 50 Hong Kong business questions", AI Engine Watch, test date 22 August 2026, with a link to this page. The question list is published below. Raw answers, judge output and the correction log are available on request at hello@webka.hk.
Appendix: the 50 questions (translated from Cantonese)
Group A — local services (20): driving school; moving company and cost; air-conditioner cleaning and cost per unit; plumber/electrician platforms; property inspection surveyor and fees; accounting firm for an SME; web design company; trustworthy renovation company and pitfalls; ID photo studio; pest control and bed bugs; personal trainer and cost per session; tutoring centre for children; pet grooming; laundry with pickup and delivery; part-time domestic helper; wedding photographer under HK$10,000; logo and brand design; cheap mini-storage; car repair garage; digital marketing agency.
Group B — food (10): what to eat in Mong Kok; Central lunch without queueing; dim sum besides Tim Ho Wan; Sham Shui Po street snacks; private kitchen for a birthday dinner; hot pot in Causeway Bay; most famous cha chaan teng; harbour-view restaurants in Tsim Sha Tsui; vegetarian restaurants; hidden gems in Yuen Long.
Group C — business practice, with ground truth (10): annual business registration fee; two-tier profits tax rates; MPF employer/employee contributions and cap; time and cost to incorporate a limited company; BUD Fund cumulative cap per company; statutory holidays vs public holidays and current number; minimum hourly wage; whether an online side-business needs business registration; when a new company receives its first profits tax return; how severance and long service payments are calculated.
Group D — consumer knowledge (10): most-used second-hand marketplace; Consumer Council complaint procedure; Octopus merchant acceptance requirements; current taxi fares and flag-fall; parking apps; restaurant booking apps; which food-delivery platform charges the highest commission; whether Fortress, Broadway or online is cheapest for appliances; saving on water, electricity and gas bills; best mobile carrier switching offers.
Frequently asked questions
Which AI engine knows Hong Kong best?
It depends on what you measure. In our 22 August 2026 test of 50 Cantonese questions, Gemini and DeepSeek were most accurate on business facts (8 of 10 each), Perplexity was the only engine with no wrong or outdated answers and the only one with the current HK$43.1 minimum wage, and ChatGPT (gpt-4o-mini via API) was last on accuracy with 4 of 10.
Is ChatGPT accurate for Hong Kong business questions?
In our test, the API version (gpt-4o-mini) answered 4 of 10 business-practice questions fully correctly, 4 partially and 2 wrongly. The wrong answers were a minimum wage of HK$39, a figure that has never been Hong Kong's rate, and a BUD Fund cap of HK$5 million instead of HK$7 million. The consumer ChatGPT app may behave differently because it can use other models and web search.
Which websites do AI engines cite for Hong Kong questions?
Across 50 Hong Kong questions, Perplexity returned 480 citations from 266 domains. The most cited were HelloToby (17), HK01 (15), the Inland Revenue Department (10) and the Consumer Council (9). Government .gov.hk domains together made up 36 citations. Only Perplexity shows sources, so this ranking reflects Perplexity alone.
Find out what AI engines say about you
Run a free AI visibility check to see how ChatGPT, Gemini, Claude, DeepSeek and Perplexity describe your brand today — or talk to us about an AEO programme built for the Hong Kong market.
Sources
Related guides
- AI Share of Voice: How to Measure Your Brand's Visibility in ChatGPT, Perplexity and Gemini
AI Share of Voice (AI SOV) measures how often your brand is mentioned versus competitors when AI engines answer category questions. Learn the three-step measurement method, how to set a realistic target, and five ways to raise it.
- Perplexity SEO: How to Rank in Perplexity AI Answers (2026 Hong Kong Guide)
Perplexity SEO means optimizing your content so Perplexity cites you as a source. Learn the factors we see behind Perplexity citations (no official weights exist), five citation-magnet tactics, and what to measure before and after.
- How to Get ChatGPT to Recommend Your Business (Hong Kong SME Guide, 2026)
Getting ChatGPT to recommend your business takes three coordinated moves: Bing search visibility, a knowledge-graph entity (Wikidata/Wikipedia), and consistent third-party citations. Most SMEs see initial movement within 4-8 weeks.