Falcon-Emirati: When an LLM Learns the Dialect, the Culture, and the Nuance

2026年10月6日 06:44
站內 AI 整理稿

Back to Articles Falcon-Emirati: When an LLM Learns the Dialect, the Culture, and the Nuance Team Article Published October 6, 2026 Upvote 7 +1 Shaikha Alsuwaidi Shaikha710 Follow tiiuae Omar saif alkaabi Omar-Alkaabi Follow tiiuae Maitha Alhammadi MaithaAlhammadi Follow tiiuae Ahmed Alzubaidi amztheory Follow tiiuae Mohammed Alyafeai Alyafeai Follow tiiuae Leen AlQadi LeenAlQadi Follow tiiuae Basma Boussaha basma-b Follow tiiuae Hakim Hacid HakimHacid Follow tiiuae Arabic is really a family of languages living under one name.

Modern Standard Arabic is what you read in the news or a textbook, but it's rarely how people actually talk to each other.

In the UAE, day-to-day conversation, humor, negotiation, and storytelling happen in Emirati Arabic, a Gulf dialect with its own vocabulary, its own rhythm, and a culture wrapped tightly around it.

Emirati poetry, especially nabati poetry, along with proverbs and short anecdotes, carries meaning that doesn't survive a literal, word-for-word reading.A model that only knows MSA can translate every word of an Emirati sentence and still miss what it actually means.

That's the gap Falcon-Emirati-7B is built to close.It's a dialect-specialized model on top of Falcon-H1-Arabic, aimed at understanding and generating Emirati Arabic the way a native speaker would: the vocabulary, the tone, and the cultural context behind it.

Built on Falcon-H1-Arabic We didn't start from scratch.Falcon-Emirati-7B is built on Falcon-H1-Arabic, our Arabic model family that already set new benchmarks for the language earlier this year.

Falcon-H1-Arabic uses the Falcon-H1 hybrid architecture: State Space Models (Mamba) and Transformer attention running in parallel inside every block, with their outputs fused before each block's projection.

That combination gives the linear-time efficiency of Mamba on long sequences while keeping the precision of attention for long-range dependencies, which matters for a morphologically rich language like Arabic.

The family spans three scales (3B, 7B, and 34B parameters) with context windows up to 128K and 256K tokens, and it was already trained on a broad mix of MSA and dialectal Arabic (Gulf, Levantine, Egyptian, Maghrebi) alongside English and multilingual data.

That gave us a strong starting point: a model that already understood Arabic broadly, handled long context well, and had some dialectal exposure baked in.

Falcon-Emirati-7B takes that foundation and pushes it specifically toward the Emirati dialect, the vocabulary, the grammar, and the cultural knowledge that a general Arabic model, however capable, doesn't pick up on its own.We built Falcon-Emirati-7B on the 7B variant specifically.

It's the sweet spot in the family: large enough to hold onto the nuance that dialect adaptation needs, but small enough that both training and inference stay practical.

The 34B model would likely push quality a bit further, but at a training and serving cost that doesn't make sense for a dialect-specialized chat model, and the 3B model doesn't leave enough headroom for the depth of cultural and linguistic understanding we were after.

7B gave us the best balance of quality against training and inference cost.Why Dialect Adaptation Is Hard Turning a general Arabic model into an Emirati-dialect specialist sounds like a smaller job than building the base model in the first place.It isn't.

A few things make it genuinely difficult: Emirati is mostly a spoken dialect.It shows up far less in writing online than MSA, or even other Gulf and Levantine dialects, so there just isn't as much raw text to learn from.Meaning is often non-literal.

Idioms, proverbs, and poetic references lean on shared cultural context, not surface vocabulary.There's no established playbook.

There isn't a well-documented recipe for how much dialectal data is enough, how to mix it with MSA and general Arabic, or which training stage (continued pre-training, SFT, or preference optimization) matters most for picking up a dialect.That last point shaped how we worked.

A lot of building Falcon-Emirati-7B came down to trial and error: testing different data mixes, training stages, and supervision strategies, and using both human judgment and benchmark scores to figure out what actually moved the needle.

Our Approach to Data We built a dedicated Emirati data pipeline on top of Falcon-H1-Arabic's pretraining, drawing on three complementary sources.1.

Authentic Emirati-Dialect Web Data We crawled and curated content from Emirati websites and forums written natively in the dialect, not translated or transliterated from MSA.

This is where we got our ground truth: how Emiratis actually write and speak online, the everyday phrasing, the colloquial expressions, and the natural back-and-forth between Emirati and MSA that shows up in real usage.2.

MSA Data About Emirati Culture and Identity Alongside the dialectal text, we pulled in MSA-language material specifically about Emirati culture, heritage, and language: articles and references on local customs, values, history, and social norms, including how Emiratis are perceived and stereotyped.

This doesn't teach the model to write in dialect, but it teaches the model what it's talking about when Emirati topics come up, things like heritage, etiquette, and the context a native speaker just knows.3.

Synthetic Data, Guided by Glossaries and Style Rules Authentic dialectal text alone wasn't enough to cover the range of topics a chat model actually needs to handle day to day.So we generated a large amount of synthetic Emirati-dialect data to fill the gaps.

We didn't just let a generator model improvise in "Gulf-ish" Arabic.We constrained it with strict rules and glossaries and dictionaries built specifically for Emirati vocabulary and grammar.

Those guardrails made the difference between synthetic output that reads as authentically Emirati and output that's grammatically fine but sounds off to anyone who actually speaks the dialect.

Finding the Right Adaptation Recipe Since there's no standard recipe for MSA-to-dialect adaptation, we treated the training strategy itself as something to figure out experimentally.

We ran ablations on how much dialectal data to inject and at which stage of training, how to balance authentic crawled data against synthetic data without the model overfitting to synthetic patterns, and how much MSA cultural context was actually needed to keep it culturally grounded rather than just fluent on the surface.

At each step we leaned on a mix of automatic scoring and native-speaker review, since automatic metrics alone don't capture naturalness, tone, or cultural fit well enough to trust on their own.

Evaluation Methodology We tracked progress throughout training with two complementary approaches: Manual Evaluation by Native Speakers Emirati native speakers reviewed model outputs directly, judging not just whether an answer was correct but whether it sounded right: naturalness, tone, cultural appropriateness.

These are the things a benchmark score won't tell you but a native ear catches immediately.Automatic Evaluation on Alyah For quantitative tracking, we used Alyah (الياه, "North Star"), a benchmark we and the community released specifically to evaluate Emirati-dialect capability in Arabic LLMs.

Alyah is a fully native multiple-choice benchmark of 1,173 samples, collected manually from native Emirati speakers and spanning categories from everyday greetings and etiquette to figurative language, heritage knowledge, and Emirati poetry: the categories where dialect and culture matter most and where generic Arabic models tend to struggle.

Full details on Alyah's construction and category breakdown are available in our benchmark blog post, and background on the base model family is available in the Falcon-H1-Arabic announcement.Results Falcon-Emirati-7B scores 84.

83% on Alyah, ahead of every other Arabic and multilingual model we compared it against, including several models many times its size.The chart below shows where it lands next to a representative set of leading instruction-tuned models on the Alyah leaderboard.

Alyah accuracy (%), instruction-tuned models.Falcon-H1-Arabic family models are excluded from this comparison since Falcon-Emirati-7B is built on top of them.What the Results Tell Us A couple of things jump out from this comparison.Size alone doesn't buy you dialect competence.

Some of the largest multilingual models here score well below smaller, more dialect-aware ones, which tells you Emirati proficiency has to be trained for on purpose, not picked up as a side effect of scale.

The models that do best also tend to be Arabic-native or Arabic-focused to begin with, which lines up with what we saw during our own ablations: general Arabic and dialect coverage is a necessary starting point, but it still takes targeted, dialect-specific work to close the rest of the gap, particularly on the hardest parts of Alyah, like poetry, heritage knowledge, and the language-and-dialect category itself.

This also matches what came out of the Alyah benchmark release more broadly: even strong models show real degradation once you move into genuinely dialectal, culturally embedded content.That gap doesn't close on its own with bigger models.

It takes data and evaluation built specifically for the dialect.Beyond Multiple Choice: LLM-as-Judge Evaluation Multiple-choice accuracy tells you whether a model can recognize the right answer among four options.

It doesn't tell you whether the model will actually produce Emirati Arabic on its own when someone just talks to it.So alongside Alyah, we ran a second evaluation: open-ended generation on the same 1,173 Alyah questions, scored by an LLM judge (Gemini 3.

7 Flash) against five models, Falcon-Emirati-7B, ALLaM-7B-Instruct-preview, gemma-3-27b-it, Jais-2-8B-Chat, and Fanar-2-27B-Instruct, chosen as the strongest competing models from the Alyah leaderboard.

The judge scored each answer on two separate dimensions: whether the content was correct, and, independently, whether the answer actually came back in Emirati dialect rather than MSA.

We report both a partial-credit score (the judge's graded assessment) and a stricter pass/fail version, plus how often each model abstained instead of answering.LLM-judged correctness on the 1,173 Alyah questions, open-ended generation, Gemini 3.7 as judge.

LLM-judged dialect fidelity on the same questions: does the answer actually come back in Emirati, or does the model default to MSA?Falcon-Emirati-7B leads on correctness, but the real gap is in the second chart.On dialect fidelity, Falcon-Emirati-7B scores 0.52 (partial credit) against 0.

05 for ALLaM, 0.03 for gemma-3-27b-it, 0.02 for Jais-2-8B-Chat, and effectively 0.00 for Fanar-2-27B-Instruct.That's not a small edge, it's close to two orders of magnitude at the low end.

In practice, this means the other models often know the right answer but say it in Modern Standard Arabic by default, even when asked directly in Emirati.Falcon-Emirati-7B is the only one of the five that reliably answers back in the dialect it was asked in.

Fanar-2-27B-Instruct stands out for a second reason too: it abstains far more than any other model, declining to answer 26.2% of the time, versus under 5% for every other model in the comparison.Combined with its correctness score of 0.

27 (partial credit), the lowest of the five, it suggests a model that is both less willing and less able to engage with Emirati-specific content, rather than one that's just answering in the wrong register.Dialect fidelity by Alyah category, partial credit.

Falcon-Emirati-7B is the only model that consistently switches into Emirati; the others stay in MSA across nearly every category.Breaking dialect fidelity down by category makes the pattern even clearer.

It holds across every single category in Alyah, from everyday greetings to poetry, which suggests this isn't a narrow trick learned for a handful of question typ

Related

相關文章

MarkTechPost AI生成式AI

Reka 發表 Rho-1:一款 19B 參數的全能推理模型,能理解、生成影片並輸出機器人動作

Reka 發表了 Rho-1 的研究預覽版,這是一款從零訓練的 19B 參數全能推理模型。單一神經網路能理解與生成文字、圖片和影片,對其進行推理,並輸出機器人動作。Reka 將其定位為可直接取代透過不同模態模型傳遞工作的代理型流水線。Rho-1 的改變之處:當今多數多模態系統採用流水線架構,由中央模型規劃後,將任務分配給圖片、影片或偵測等專用模型。每一次交接都會增加延遲,且每個專用模型只能處理狹義的請求。Rho-1 消除了這些交接環節,將文字、視覺與機器人動作都轉化為單一上下文視窗中的 token。根據 Reka 的研究,一次未經編輯的會話展示了完整流程:模型繪製燈塔、標出邊框、讓它動起來、編輯影片,並最終輸出……」

2 小時前
鈦媒體生成式AI

快手的視頻Agent,會不會來晚了?

AI價值官2026.10.05 17:36 · 來自浙江全文4219字00:00 / 11:40視頻模型廠商,正集體從"生成一段畫面"走向"交付一部成片"。文 | AI價值官,作者丨納瓦,編輯丨星野9 月 28 日,快手在模型與 Agent 兩層同時出手:白天,Agent 創作工具 likli 開啟內測;晚間,新一代模型 Kling 4.0 官宣將於 10 月上線。這並非快手的獨門動作,字節、MiniMax、阿里千問都已推出各自的創作 Agent。當模型能力被逐漸拉平,競爭正從模型轉向工作流與商業化。

15 小時前
鈦媒體生成式AI

AI耳機蓄勢,芯片廠商待發

半導體產業縱橫2026.10.05 15:32 · 來自內蒙古全文5079字00:00 / 15:12AI耳機還沒有成為一個邊界清晰的品類,芯片廠商卻已經開始為它準備下一代平臺。文 | 半導體產業縱橫傳統耳機,卒?AI耳機正在無線耳機市場中成為主流。據Research and Markets的測算,2025年全球AI耳機市場規模約為59.9億美元,預計到2026年將攀升至74.2億美元,並在2030年達到173.4億美元。消費電子行業裡最常被重複的一句判斷是:所有硬件,都值得用AI重新做一遍。

17 小時前
鈦媒體生成式AI

硅谷AI,正在開源

影子備忘錄2026.10.05 15:32 · 來自廣東全文4084字00:00 / 10:59大模型從哪家強到哪家合適的轉變。文 | 影子備忘錄過去兩年,硅谷AI圈最常被問到的問題只有一個:哪家的模型最強?但2026年,這個問題正在被另一個更務實的問題取代:完成同一項任務,到底有沒有必要用最貴的模型?今年6月,打車巨頭Uber透露了一個讓整個行業沉默的數字:僅2026年前四個月,公司就耗盡了全年的AI預算,原因僅僅是員工大量使用AI編碼工具。

17 小時前