推出 Falcon ASR 語音辨識模型

2026年10月7日 13:21
站內 AI 整理稿

Back to Articles Introducing Falcon ASR Team Article Published October 7, 2026 Upvote - Abdul Muneer abdulmuneertii Follow tiiuae Lepauloux Ludovick Follow tiiuae Rishabh Saraf rishabh-saraf Follow tiiuae Shamsa Hamad shamsaH Follow tiiuae English · العربية Arabic WER: 20.92% · Parameters: 1.

6B · Emirati WER (TII evaluation): 22.73% We’re introducing Falcon-ASR, our 1.6 billion parameter speech recognition model for Arabic, with a particular focus on the Emirati dialect.

Developed at the Technology Innovation Institute (TII) in Abu Dhabi, it also supports English, French, Spanish and Portuguese.In our evaluation, Falcon-ASR achieved an average word error rate of 20.92% across six Arabic test sets, compared with the best published result of 23.

17% in the leaderboard snapshot we used.On our internal Emirati evaluation, it recorded the lowest word and character error rates among the systems we compared.We also support word-level timestamps for transcriptions, linking each transcribed word to its position in the audio.

You can try Falcon-ASR in our Hugging Face Demo.Recognising spoken Arabic Arabic speech varies by region, speaker and setting.A model that handles a formal news broadcast may still struggle with a conversation in Emirati or with speech recorded over a phone line.

Dialectal Arabic also has fewer transcribed resources than Modern Standard Arabic (MSA), which makes training and evaluation harder.We trained Falcon-ASR on Emirati, MSA, other Gulf and Arabic dialects, and English.

Our aim is to transcribe the words people use in everyday speech, including dialectal forms and changes between languages.Arabic benchmark results The Open Universal Arabic ASR Leaderboard, maintained by the ELM Research Center, ranks systems by the equal-weight average WER across six test sets.

It also reports character error rate (CER).Lower values are better for both metrics.Our Falcon-ASR evaluation follows this protocol.Model Parameters Avg WER (%) Avg CER (%) Falcon-ASR 1.6B 20.92 8.79 Audar-ASR-V1-Turbo 2.35B 23.17 9.23 Cohere Transcribe Arabic (07-2026) 2.0B 25.87 11.

80 omniASR LLM 7B 7.0B 28.32 12.52 WER = Word Error Rate; CER = Character Error Rate.A lower value indicates better performance.We evaluated Falcon-ASR on the same six benchmarks using the leaderboard’s pinned manifests.

Competitor figures are the published leaderboard averages checked on 30 September 2026.Falcon-ASR’s average WER is 2.25 percentage points better than the best published result in that snapshot.Evaluating Emirati speech Public evaluation data already includes Emirati: Casablanca has a UAE subset.

We complement that coverage with an internal evaluation of additional Emirati and Gulf speech, using held-out recordings and human-validated transcripts to assess transcription accuracy beyond the public UAE subset.In our internal Emirati evaluation, Falcon-ASR achieved 22.73% WER and 10.

19% CER: Model Parameters WER (%) CER (%) Falcon-ASR 1.6B 22.73 10.19 Qwen3-Omni-30B-A3B-Instruct 30.0B (3.0B active) 26.80 12.72 Audar-ASR-V1-Turbo 2.35B 27.89 13.75 Cohere Transcribe Arabic (07-2026) 2.0B 31.05 18.07 Qwen3-ASR-1.7B-hf 2.0B 31.52 13.35 Audar-ASR-V1-Flash 0.78B 32.87 15.

36 Falcon-ASR has the lowest WER and CER among the systems compared here.Its WER is 4.07 percentage points below Qwen3-Omni, the next best result.The results show improved transcription accuracy at both the word and character level on this evaluation.

Training for different recording conditions We included background noise, overlapping speech, music, room reverberation and telephony effects, as well as variations in speed and pitch.

We applied the same treatment to Emirati recordings, exposing the model to a range of conditions it may encounter in meetings, calls and other everyday recordings.English and other languages Falcon-ASR also transcribes English with the same model weights.

In our evaluation on the seven public English test sets used by the Hugging Face Open ASR Leaderboard, it achieved a mean WER of 5.74%.Test set WER (%) LibriSpeech clean 1.75 LibriSpeech other 4.21 SPGISpeech 2.02 VoxPopuli 3.87 GigaSpeech 8.15 AMI 8.33 Earnings-22 11.

86 The model also supports French, Spanish and Portuguese.All five languages use the same weights, without requiring a language flag.The output is a transcript in the language spoken.Model foundation Falcon-ASR builds on our Falcon3-Audio work.

The architecture and training approach for Falcon3-Audio are described in Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data.Try Falcon ASR Our Hugging Face Demo Space lets you try Falcon-ASR and explore its transcription capabilities.

API access and native applications are planned.We invite you to try the Demo with your own recordings.نقدّم Falcon ASR نقدّم Falcon-ASR، نموذجنا للتعرف على الكلام العربي بحجم 1.6 مليار معلمة، مع اهتمام خاص باللهجة الإماراتية.

طوّرنا النموذج في معهد الابتكار التكنولوجي (TII) في أبوظبي، وهو يدعم أيضًا اللغات الإنجليزية والفرنسية والإسبانية والبرتغالية.في تقييمنا، حقق Falcon-ASR متوسط معدل خطأ في الكلمات بلغ 20.92% عبر ست مجموعات اختبار باللغة العربية، مقارنةً بأفضل نتيجة منشورة بلغت 23.

17% في نسخة لوحة المتصدرين التي استخدمناها.وفي تقييمنا الداخلي للهجة الإماراتية، سجّل النموذج أدنى معدلات خطأ في الكلمات والأحرف بين الأنظمة التي قارناها.ندعم أيضاً الطوابع الزمنية على مستوى الكلمات في عمليات التفريغ النصي، بحيث ترتبط كل كلمة بموضعها في التسجيل الصوتي.

يمكنكم تجربة Falcon-ASR عبر عرضنا التجريبي على Hugging Face.التعرف على العربية المنطوقة يختلف الكلام العربي باختلاف المنطقة والمتحدث وظروف التسجيل.فقد ينجح نموذج في تفريغ نشرة إخبارية رسمية، ثم يجد صعوبة في تفريغ محادثة باللهجة الإماراتية أو تسجيل عبر الهاتف.

كما أن الموارد الصوتية المفرّغة نصيًا للهجات العربية أقل من تلك المتاحة للعربية الفصحى، مما يزيد صعوبة التدريب والتقييم.درّبنا Falcon-ASR على اللهجة الإماراتية والعربية الفصحى ولهجات خليجية وعربية أخرى، إلى جانب الإنجليزية.

وهدفنا هو تفريغ الكلمات التي يستخدمها الناس في حديثهم اليومي، بما في ذلك الصيغ اللهجية والانتقال بين اللغات.

نتائج الاختبارات العربية ترتّب لوحة المتصدرين المفتوحة الشاملة للتعرف على الكلام العربي، التي يديرها مركز ELM للأبحاث، الأنظمة وفق متوسط معدل خطأ الكلمات عبر ست مجموعات اختبار، بوزن متساوٍ لكل مجموعة.وتعرض أيضًا معدل خطأ الأحرف (CER).وكلما انخفضت قيمة أي من المقياسين، كان الأداء أفضل.

ويتبع تقييمنا للنموذج هذا البروتوكول.Model Parameters Avg WER (%) Avg CER (%) Falcon-ASR 1.6B 20.92 8.79 Audar-ASR-V1-Turbo 2.35B 23.17 9.23 Cohere Transcribe Arabic (07-2026) 2.0B 25.87 11.80 omniASR LLM 7B 7.0B 28.32 12.52 WER هو معدل خطأ الكلمات؛ وCER هو معدل خطأ الأحرف.

تشير القيمة الأقل إلى أداء أفضل.قيّمنا Falcon-ASR على مجموعات الاختبار الست نفسها، باستخدام قوائم العينات المثبّتة في لوحة المتصدرين.وأرقام النماذج المنافسة هي المتوسطات المنشورة في اللوحة، والتي جرى التحقق منها في 30 سبتمبر 2026.وكان متوسط خطأ الكلمات للنموذج أفضل بمقدار 2.

25 نقطة مئوية من أفضل نتيجة منشورة في تلك النسخة.تقييم اللهجة الإماراتية تتوافر بالفعل بيانات عامة لتقييم اللهجة الإماراتية، إذ تضم مجموعة Casablanca قسمًا خاصًا بالإمارات.

ونكمّل هذه التغطية بتقييم داخلي لتسجيلات إضافية من الكلام الإماراتي والخليجي، باستخدام تسجيلات مخصّصة للاختبار ونصوص مرجعية خضعت لمراجعة بشرية، لتقييم دقة التفريغ على مواد إضافية إلى جانب البيانات الإماراتية العامة.حقق Falcon-ASR في تقييمنا الإماراتي الداخلي معدل خطأ كلمات قدره 22.

73٪ ومعدل خطأ أحرف قدره 10.19٪: Model Parameters WER (%) CER (%) Falcon-ASR 1.6B 22.73 10.19 Qwen3-Omni-30B-A3B-Instruct 30.0B (3.0B active) 26.80 12.72 Audar-ASR-V1-Turbo 2.35B 27.89 13.75 Cohere Transcribe Arabic (07-2026) 2.0B 31.05 18.07 Qwen3-ASR-1.7B-hf 2.0B 31.52 13.35 Audar-ASR-V1-Flash 0.

78B 32.87 15.36 سجّل Falcon-ASR أقل معدل لخطأ الكلمات والأحرف بين الأنظمة المقارَنة هنا.وكان معدل خطأ الكلمات أقل بمقدار 4.07 نقطة مئوية من Qwen3-Omni، صاحب النتيجة التالية.وتُظهر النتائج تحسنًا في دقة التفريغ على مستوى الكلمات والأحرف في هذا التقييم.

التدريب على ظروف تسجيل مختلفة ضمّنا بيانات التدريب ضوضاء خلفية وكلامًا متداخلًا وموسيقى وصدى الصوت وتأثيرات الاتصالات الهاتفية، إلى جانب تغيّرات في سرعة الكلام وحدّة الصوت.

وطبّقنا المعالجة نفسها على التسجيلات الإماراتية، لتهيئة النموذج للتعامل مع ظروف مختلفة قد يواجهها في الاجتماعات والمكالمات والتسجيلات اليومية الأخرى.الإنجليزية واللغات الأخرى يفرّغ Falcon-ASR الكلام الإنجليزي باستخدام أوزان النموذج نفسها.

وفي تقييمنا على مجموعات الاختبار الإنجليزية العامة السبع المستخدمة في لوحة Hugging Face المفتوحة للتعرف على الكلام، بلغ متوسط معدل خطأ الكلمات 5.74٪.Test set WER (%) LibriSpeech clean 1.75 LibriSpeech other 4.21 SPGISpeech 2.02 VoxPopuli 3.87 GigaSpeech 8.15 AMI 8.33 Earnings-22 11.

86 يدعم النموذج أيضًا الفرنسية والإسبانية والبرتغالية.وتستخدم اللغات الخمس أوزانًا واحدة، دون الحاجة إلى تحديد اللغة مسبقًا.ويكون الناتج تفريغًا نصيًا باللغة المنطوقة.أساس النموذج يستند Falcon-ASR إلى أعمالنا في Falcon3-Audio.

وتعرض ورقة Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data بنية Falcon3-Audio ونهج تدريبه.تجربة Falcon ASR يتيح عرضنا التجريبي على Hugging Face تجربة Falcon-ASR واستكشاف قدراته في تفريغ الكلام.أما الوصول عبر واجهة API والتطبيقات الأصلية فهو مخطّط له.

ندعوكم إلى تجربة النموذج باستخدام تسجيلاتكم.

More from this author Falcon-Emirati: When an LLM Learns the Dialect, the Culture, and the Nuance 20 October 6, 2026 Falcon OCR Arabic: 270M Parameters State-of-the-Art Arabic OCR 13 October 6, 2026 Community EditPreview Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images Comment · Sign up or log in to comment Upvote -

Related

相關文章

智東西模型更新

800億!曝DeepSeek新融資,即將IPO

編譯 | 畢偉豪 編輯|心緣 10月6日消息,據外媒彭博社今日報道,DeepSeek即將敲定至少800億元人民幣的新一輪融資,騰訊和寧德時代承諾的出資金額均位居該輪融資最高的一批。據彭博社報道,知情人士稱,根據已簽署的條款,該輪融資的最終募資總額可能接近1000億元,為計劃在2027年初進行的IPO鋪路。

1 天前
MarkTechPost AI模型更新

超越特定領域的世界模型:JEPA-Anything 以單一方法涵蓋七大領域

來自 PhAI Labs、香港中文大學、復旦大學、史丹佛大學、牛津大學與普林斯頓大學的研究團隊釋出 JEPA-Anything,這是一個與領域無關的世界模型建構框架。它不為每個領域設計專屬預測模型,而是用一套共享學習方案應對截然不同的系統。該框架採用正交預測分解(OPF)技術擴充了聯合嵌入預測架構(JEPA)。研究團隊在七大領域進行測試:視覺、生物學、臨床軌跡、控制、分子動力學、物理場域與天氣。JEPA-Anything 解決了什麼問題?標準 JEPA(如 I-JEPA 或 V-JEPA 2)使用上下文編碼器、EMA 目標編碼器與一個預測器,而預測器只輸出單一的整體目標嵌入,研究團隊稱此為容量分配問題。

1 天前
量子位模型更新

最火AI崗位FDE:月薪5萬,都幹這些…

FDE(前線部署工程師)是近期最受關注的AI職位之一,海外年薪中位數約20萬美元,國內大廠也開出月薪三到五萬元。這份工作強調駐場梳理客戶的業務本體(Ontology),溝通時間佔七成以上,開發僅約三成。從業者認為,FDE與傳統外包不同,關鍵在於能否將經驗沉澱回自家產品並複用。

3 天前
MarkTechPost AI模型更新

DeepSeek Harness v0.2 為其開源代理框架帶來官方桌面應用程式

DeepSeek 已為 DeepSeek Harness (dsh) 推出官方桌面應用程式,dsh 是其開源代理框架。該應用程式隨 v0.2 預覽版一同發布,安裝檔支援 macOS(Apple 晶片)與 Windows(64 位元)。目前可作為預覽版部署使用,使用者可從 deepseek.com/harness 下載,或執行 npx @deepseek-ai/dsh web。DeepSeek 提醒未來可能會有破壞相容性的變更。 v0.2 新增內容:框架是將模型轉化為代理的執行環境,能讀取檔案、執行指令並維持計畫。v0.2 預覽版針對日常工作與程式開發進行優化。內建功能包括:預載常用辦公室與開發工具;新增插件管理頁面,可安裝、設定、啟用與停用插件;以及右側邊欄提供檔案與差異審查預覽。

3 天前