Bangla ASR → SRT
Drop in a video, get Bangla subtitles — powered by a rescoring pipeline that beats Whisper's own best guess.
The problem
Whisper's single best transcript is often not its best transcript. Can we recover accuracy for Bangla without fine-tuning — no weight updates at all?
What I built
- Pipeline: VAD chunking → group-aware split → Whisper N-best generation on Modal GPUs
- Probability-weighted Minimum Bayes Risk (MBR) rescoring over the N-best lists
- A video-to-SRT transcription web tool on top
- Supervised by Dr. Nabeel Mohammed (NSU), two-semester capstone aimed at publication
Outcome
- 1-best WER 34.85% → 32.29% on 2,271 utterances
- ~42% of the oracle WER headroom recovered, zero fine-tuning
Stack
- Python
- Whisper
- Modal
- MBR