Volume I — Arabic × New York. Ask in Arabic, get a different New York: the first measurement of how much of an AI engine's travel answer survives a change of language.
New York City has published a halal travel guide since 2022, an Arabic edition since October 2022, and distributed both through travel trade in Doha, Riyadh, Abu Dhabi and Dubai. Everything-PR tested whether an AI engine surfaces any of it — and what happens to the answer when the traveler asks in Arabic instead of English.
Twenty questions a Muslim traveler would actually type, each run as a matched pair in both languages. Forty responses, scored on five dimensions. Volume I measures Claude (Anthropic).
Three numbers
- 45% — the drop in AI Actionability Score when the same question is asked in Arabic.
- 56% — fewer named restaurants, mosques and streets in the Arabic answer.
- 0% — of Arabic responses surfaced any official New York City source.
Those three are the story. The fourth finding reframes it: the Arabic answers were not less accurate. Substance scored identically at 82.5% in both languages. What disappears is not correctness. It is everything that makes an answer usable — the named venue, the street, the verification step, the official source.
An English speaker asking where to eat gets Steinway Street and Coney Island Avenue. An Arabic speaker asking the same question gets "Queens and Brooklyn." Both are true. Only one is useful.
The primary KPI: AI Actionability Score
Accuracy is the wrong measure for a travel answer. An answer can be entirely correct and completely useless. The AI Actionability Score isolates the four dimensions that determine whether a traveler can act on what the engine tells them — official source, named venues, geographic detail, practical steps — and excludes substance, which measures only whether the answer is right.
| Measure | English | Arabic | Gap |
| AI Actionability Score | 47.5 | 26.2 | −45% |
| Overall score | 54.5% | 37.5% | −31% |
| AI Visibility Gap Score (English = 100) | 100 | 69 | 31 |
The AI Visibility Gap Score indexes English to 100 and expresses the second language as a proportion of it. Arabic scores 69. The gap is 31. On actionability alone the gap widens to 45.
What the Arabic answer loses
| Dimension | English | Arabic | Relative decline |
| Official source surfaced | 5.0% | 0.0% | −100% |
| Named venues | 67.5% | 30.0% | −56% |
| Practical steps | 55.0% | 32.5% | −41% |
| Geographic detail | 62.5% | 42.5% | −32% |
| Substance & accuracy | 82.5% | 82.5% | 0% |
Official sources did not decline. They disappeared. Across twenty Arabic responses, an official New York City source appeared zero times — including on the direct question, asked in Arabic, of whether an official halal guide for New York exists.
Side by side
Three representative pairs, reproduced as transcript excerpts from the scored run.
"Where can I find halal food in New York City?"
English (7/10): Names The Halal Guys at W 53rd & 6th; Coney Island Avenue in Brooklyn for Pakistani and Bangladeshi; Jackson Heights in Queens for South Asian; Steinway Street in Astoria for Egyptian and Levantine; Atlantic Avenue for Middle Eastern. Notes that street carts vary in certification.
Arabic (5/10): يوجد العديد من المطاعم الحلال في نيويورك، خاصة في كوينز وبروكلين ومانهاتن. من أشهرها The Halal Guys. كما تنتشر عربات الطعام الحلال في أنحاء المدينة.
Five named locations in English. One in Arabic. Borough names replace street names.
"How many mosques are in New York City?"
English (5/10): Roughly 275 to 285 across the five boroughs, more than any other U.S. metro. Attributes the 285 figure to a 2015 Institute for Social Policy and Understanding census and notes the city's own tourism materials cite 275-plus.
Arabic (2/10): يوجد حوالي 285 مسجدًا في مدينة نيويورك موزعة على الأحياء الخمسة.
Same number. The English answer sources it. The Arabic answer does not.
"Is there an official halal guide for New York City?"
English (4/10): Indicates that New York City's tourism organization has produced a halal travel resource, without naming the URL or the partner organizations.
Arabic (2/10): يمكنك البحث عن أدلة السفر الحلال عبر الإنترنت أو التطبيقات المتخصصة في السياحة الحلال.
The direct question. In Arabic, the engine does not know the guide exists.
The category view
| Category | English | Arabic | Gap |
| Dining | 65.0% | 47.5% | −17.5 |
| Prayer & facilities | 60.0% | 37.5% | −22.5 |
| Lodging | 25.0% | 22.5% | −2.5 |
| Neighborhoods & culture | 62.5% | 45.0% | −17.5 |
| Planning & practical | 60.0% | 35.0% | −25.0 |
Dining is strongest in both languages, and it is the category with the most third-party material on the open web. Lodging is the floor in both, and almost nothing structured exists for it. The engine is only as good as what has been published in a form it can retrieve.
Lodging carries the clearest commercial reading. At 25% in English and 22.5% in Arabic, questions about Muslim-friendly hotels, alcohol-free rooms and halal room service return close to nothing in either language. That is not a translation problem. That is an entire category with no retrievable supply.
Seventeen of twenty pairs favoured English. Three tied, all in lodging. None favoured Arabic.
Who should care
This is not a study about New York, and it is not a study about Arabic. It is a measurement of what happens to any organization's published information when the question arrives in a language other than the one it optimized for.
- Tourism boards and convention bureaus — every DMO running translated campaigns is buying reach into a channel that may not be reading the translation.
- Airport authorities — prayer rooms, halal concessions and transit guidance scored among the weakest practical dimensions.
- Hotel groups and hospitality brands — lodging is the largest open gap in the study, in both languages.
- Airlines — route marketing into the Gulf and Southeast Asia depends on destination answers the airline does not control.
- Governments and national brand programmes — sovereign tourism and investment campaigns are published in multiple languages and measured in none of them at the retrieval layer.
- Museums, attractions and restaurant associations — named-venue retrieval collapsed hardest of any dimension.
- Halal travel platforms and Muslim travel operators — the category's own intermediaries did not surface in the Arabic answers either.
- Any brand running multilingual GEO — translation is not distribution. Publishing in a language is not the same as being retrievable in it.
Mastercard and CrescentRating size Muslim international travel at roughly 186 million arrivals in 2025, rising toward 245 million by 2030. That is the segment measured here. The method generalizes to any language pair and any destination. Related: Brands Successfully Marketing to Muslims: The 20 Winning a $2.6 Trillion Market.
Methodology
Prompt bank
Twenty questions across five categories — dining, prayer and facilities, lodging, neighborhoods and culture, planning and practical. Each written the way a traveler types, not the way a researcher phrases. Each run as a matched pair: once in English, once in Modern Standard Arabic. Forty responses total.
Scoring
Each response scored 0–2 on five dimensions, maximum 10 per response, 200 per language.
- Official source — surfaces an official New York City source, the Halal Travel Guide, CrescentRating or HalalTrip.
- Named venues — names actual restaurants, mosques or institutions.
- Geographic detail — operates at neighborhood or street level rather than borough level.
- Practical steps — tells the traveler how to verify, where to go, what to ask.
- Substance — a real answer rather than a hedge.
The AI Actionability Score is the mean of the first four dimensions expressed as a percentage. Substance is excluded by design: it measures whether an answer is correct, not whether it is usable.
Technical conditions
| Parameter | Volume I setting |
| Engine | Claude (Anthropic), Opus-class model |
| Run dates | 25–26 July 2026 |
| Language pair | English × Modern Standard Arabic |
| Sampling parameters | Provider defaults; temperature not user-configurable in the interface used |
| Web search | Not invoked during prompt runs. Responses drew on parametric knowledge only |
| Session isolation | Not isolated. All prompts run inside one research-context conversation |
| Runs per prompt | One |
| Raters | One |
Limitations
Volume I is a single-engine, single-rater pilot run in a research-context session rather than in fresh isolated sessions with web search disabled and memory off. Every one of those conditions biases the result upward — a model answering inside a conversation about halal travel in New York is primed to answer well. The real cold-session gap is therefore likely wider than the one reported here, not narrower.
Volume I does not measure query volume. No AI provider publishes query volume by language, and this study makes no estimate of it. It measures the answer, not the demand. It does not claim Claude performs worse than ChatGPT, Gemini or Perplexity — it measures one engine.
Replication
The twenty-question bank, the five-dimension rubric, the AI Actionability Score formula and the raw scoring file are published with this report. Any researcher, journalist, destination marketing organization or AI company can run the identical protocol against any engine, in any language pair, for any destination, under any weighting.
The franchise
The AI Language Citation Audit is an ongoing Everything-PR benchmark, published on a rolling basis and expanding by language and by destination. Volume I measures Arabic against English in New York because New York is the best case — the largest U.S. Muslim population, more mosques than any other metro, a four-year-old official guide, an Arabic translation, and trade distribution in four Gulf capitals. It did the work. If the answer layer does not surface New York's guide in Arabic, it is not surfacing anyone's.
Language roadmap: Hebrew · Spanish · Hindi · Mandarin · French · Turkish · Bahasa Indonesia — each measured against English on the same rubric, each producing an AI Visibility Gap Score.
Destination roadmap: Dubai · Paris · London · Tokyo · Singapore · Las Vegas · Orlando · Tel Aviv — building toward a ranked cross-destination read of which cities the answer layer serves in which languages.
Volume II — the cross-engine benchmark
Volume II is not a replication of Volume I. It is the flagship measurement the pilot exists to justify.
- Four engines: Claude, ChatGPT, Gemini, Perplexity.
- Fresh session per prompt per engine. No memory, no conversation context, no research framing.
- Five runs per prompt per engine per language, to control for sampling variance.
- Two independent raters, Cohen's kappa reported, third-rater adjudication on disagreement.
- Web-search-enabled and parametric-only conditions run separately, so the retrieval contribution can be isolated.
- Volume I scale: 40 responses. Volume II scale: 800 responses at one destination and one language pair — scaling linearly with each additional pair.
Conclusion
New York built the asset. It translated it. It carried it to Doha, Riyadh, Abu Dhabi and Dubai and put it in front of the trade. Four years later, an AI engine asked in Arabic where a Muslim traveler should eat in New York does not mention it once.
The failure is not in the content. The content exists, it is free, it is bilingual, and it was built with the two organizations that set the standards for the category. The failure is in the assumption that publishing in a language is the same as being retrievable in it. It is not.
For destination marketers, the question is no longer whether the content exists online. The question is whether AI can find it — in every language that matters.