BanglaDialecto: An End-to-End AI-Powered Regional Speech Standardization
Md Nazmus Sadat Samin*, Jawad Ibn Ahad*, Tanjila Ahmed Medha, Fuad Rahman, Mohammad Ruhul Amin, Nabeel Mohammed, Shafin Rahman
✅ Accepted at IEEE International Conference on Big Data 2024 · (* Equal Contribution)
📄 Paper: arXiv:2411.10879
💻 Code & Dataset: github.com/EncryptedBinary/BanglaDialecto
Abstract
This project addresses recognizing Bangladeshi dialects and converting diverse Bengali regional accents into standardized formal Bengali speech. We develop a large Noakhali Dialect Dataset (NDD) — 10 hours of regional speech with parallel standard Bangla transcriptions — and design a full end-to-end pipeline with three stages: (1) dialect speech recognition (ASR), (2) dialect-to-standard Bangla machine translation (MT), and (3) standard speech synthesis (TTS). Fine-tuned Whisper Large V2 achieves CER 0.8%, WER 1.5% on dialect ASR; BanglaT5 achieves BLEU 41.6% for dialect→standard translation. AlignTTS completes the pipeline with high-quality Bangla speech synthesis.
Introduction
Bangladesh has diverse regional dialects (Noakhali, Sylheti, Chittagonian, Mymensingh, etc.) that differ significantly from standard formal Bangla in phonology, vocabulary, and grammar. Speakers face barriers in voice-assisted services, transcription tools, and communication systems built only for standard Bangla. Prior work on Bangla dialect standardization is limited by small datasets and pipeline fragmentation.
We build the first end-to-end AI pipeline for Noakhali dialect standardization, chaining state-of-the-art ASR → MT → TTS models, all fine-tuned on purpose-built dialectal data.
Task Formulation: For n dialect speech signals {s^i}, each with dialect text t_d^i and standard text t_s^i:
- Model ℱ₁: transcribe
s^i → t_d^i(Dialect ASR) - Model ℱ₂: translate
t_d^i → t_s^i(Dialect MT) - AlignTTS: generate standard speech from
t_s^i(TTS)
🎯 Key Contributions
- Noakhali Dialect Dataset (NDD) — 10 hours of annotated dialect speech (7,200 samples after segmentation) with parallel standard Bangla translations; collected from YouTube, Facebook, and direct interviews with 24 native speakers.
- End-to-End Dialect Standardization — ASR + MT + TTS chained into a seamless pipeline.
- Whisper Fine-tuning — CER 0.8%, WER 1.5% on Noakhali dialect — significantly outperforming all prior Bangla ASR baselines.
- BanglaT5 Fine-tuning — BLEU 41.6% for dialect→standard translation, surpassing mBART50, IndicBART, mT5.
Methodology
Figure: NDD dataset statistics. 24 participants (75% male, 25% female), age 18–48, with audio collected at 16 kHz mono, 16 kHz sample rate.
III-A. NDD Dataset Construction
Collection: 10 hours of Noakhali regional speech from YouTube, Facebook, and interviews. Participants read standard paragraphs in their native Noakhali accent across diverse sub-regions.
Dataset Properties:
- Unique Characters: 72
- Total Unique Words: 11,246
- Audio: 10 hours, 16 kHz, mono, 16-bit
Annotation: Qualified human evaluators from Noakhali annotated dialect transcriptions and standard Bangla translations in a 3-column format: (s^i, t_d^i, t_s^i).
III-B. Preprocessing
- Denoising: Background noise removed using noisereduce library
- Segmentation: Each signal split into fixed 5-second segments
N^i = ⌊T^i/5⌋segments per recording - Text Alignment: Dialect and standard text chunked to match speech segments
Post-segmentation split:
| Split | Samples |
|---|---|
| Training | 6,270 |
| Validation | 810 |
| Test | 120 |
III-C. Model Pipeline
Stage 1 — Dialect ASR (Whisper fine-tuning)
- Input: dialect audio segments
s^i_k - Output: dialect text
t_d^i_k - Training: 10 epochs, batch size 16
- Models: Whisper-base (74M) to Whisper-Large V2 (1550M)
Stage 2 — Dialect MT (BanglaT5 fine-tuning)
- Input: dialect text
t_d^i_k - Output: standard Bangla
t_s^i_k - Training: 25 epochs, batch size 6
- Models: XLM-ProphetNet, mBART50, IndicBART, mT5, BanglaT5
Stage 3 — TTS (AlignTTS)
- Input: standard Bangla text
- Output: synthesized speech
- AlignTTS replaces attention with alignment loss + dynamic programming for stable Bangla synthesis
📊 Table I — NDD Dataset Details
| Property | Value |
|---|---|
| Unique characters | 72 |
| Total unique words | 11,246 |
| Total audio duration | 10 hours |
| Sample rate | 16 kHz |
| Bit depth | 16 |
| Channels | Mono |
| Participants | 24 (18M, 6F) |
| Age range | 18–48 |
📊 Table II — Comparison with Prior Dialect Standardization Work
ASR (Dialect Speech → Dialect Text):
| Method | Language | CER (%) ↓ | WER (%) ↓ |
|---|---|---|---|
| Messaoudi et al. | Tunisian dialect | 18.7 | 24.4 |
| Ying et al. | Sichuan dialect | 25.24 | 57.02 |
| Alam et al. | Standard Bangla | 13.856 | 39.29 |
| Nasr et al. | Arabian dialect | 20.0 | 30.0 |
| Safieh et al. | Jordanian dialect | 26.4 | 51.50 |
| Phung et al. | Vietnamese | 4.79 | 2.79 |
| Ours | Noakhali Dialect | 0.8 | 1.5 |
MT (Dialect Text → Standard Text):
| Method | Language | CER (%) ↓ | WER (%) ↓ | BLEU (%) ↑ |
|---|---|---|---|---|
| Faria et al. | Mymensingh dialect | 8.23 | 15.48 | 69.06 |
| Faria et al. | Noakhali dialect | 20.3 | 38.7 | 47.43 |
| Kchaou et al. | Tunisian dialect | — | — | 60.00 |
| Kabir et al. | Sylhet Dialect | — | 14.89 | — |
| Hamed et al. | Arabic | — | — | 18.00 |
| Baruah et al. | Assamese | — | — | 50.19 |
| Prama et al. | Sylhet Dialect | — | — | 57.40 |
| Ours | Noakhali dialect | 20.2 | 38.2 | 41.6 |
📊 Table III — Pre-trained vs Fine-tuned Model Performance
Dialect ASR (Whisper Models):
| Model | Parameters | Pre-trained CER (%) | Pre-trained WER (%) | Fine-tuned CER (%) | Fine-tuned WER (%) |
|---|---|---|---|---|---|
| Sequential Transformer | 24M | 87.0 | 179.8 | — | — |
| Whisper-base | 74M | 230.8 | 232.2 | 20.6 | 47.1 |
| Whisper-small | 244M | 389.8 | 411.9 | 2.0 | 4.0 |
| Whisper-medium | 769M | 259.2 | 169.0 | 1.5 | 2.0 |
| Whisper-Large V2 | 1550M | 135.2 | 167.5 | 0.8 | 1.5 |
Dialect MT (Translation Models):
| Model | Parameters | Pre-train CER | Pre-train WER | Pre-train BLEU | Fine-tuned CER | Fine-tuned WER | Fine-tuned BLEU |
|---|---|---|---|---|---|---|---|
| mBART50 | 611M | 248.5 | 612.8 | 0.45% | 79.8 | 416.8 | 3.0% |
| IndicBART | 244M | 166.5 | 409.8 | 11.9% | 22.8 | 118.6 | 27.7% |
| mT5-base | 582M | 275.4 | 405.4 | 5.6% | 34.4 | 59.2 | 18.6% |
| BanglaT5 | 247M | 178.9 | 281.6 | 22.7% | 21.3 | 38.2 | 41.6% |
📚 Citation
@inproceedings{samin2024bangladialecto,
title={BanglaDialecto: An End-to-End AI-Powered Regional Speech Standardization},
author={Samin, Md Nazmus Sadat and Ahad, Jawad Ibn and Medha, Tanjila Ahmed and Rahman, Fuad and Amin, Mohammad Ruhul and Mohammed, Nabeel and Rahman, Shafin},
booktitle={IEEE International Conference on Big Data (BigData)},
year={2024}
}
