Skip to content

Bangla ASR → SRT

Drop in a video, get Bangla subtitles — powered by a rescoring pipeline that beats Whisper's own best guess.

100:00:01,200 --> 00:00:03,900200:00:04,100 --> 00:00:06,800.SRT ↓

The problem

Whisper's single best transcript is often not its best transcript. Can we recover accuracy for Bangla without fine-tuning — no weight updates at all?

What I built

  • Pipeline: VAD chunking → group-aware split → Whisper N-best generation on Modal GPUs
  • Probability-weighted Minimum Bayes Risk (MBR) rescoring over the N-best lists
  • A video-to-SRT transcription web tool on top
  • Supervised by Dr. Nabeel Mohammed (NSU), two-semester capstone aimed at publication

Outcome

  • 1-best WER 34.85% → 32.29% on 2,271 utterances
  • ~42% of the oracle WER headroom recovered, zero fine-tuning

Stack

  • Python
  • Whisper
  • Modal
  • MBR