Parsing Subtitles at Scale: Stripping SRT/WebVTT Timestamps and Building Clean Transcript Engines
A deep technical breakdown of parsing SubRip (SRT) and WebVTT caption formats, regular expression timestamp stripping, cue text normalization, and building zero-latency browser transcript tools.
Video and audio content creators, journalists, and accessibility developers frequently need to extract readable prose transcripts from video subtitle files (.srt and .vtt). However, raw subtitle files are filled with timestamp counters, line index numbers, positioning coordinates, and HTML styling tags. Converting raw caption files into clean, readable blog posts or summary notes requires robust regular expression text normalization.
1. Understanding Subtitle Formats (SRT vs WebVTT)
The two dominant subtitle formats on the web are SubRip (.srt) and WebVTT (.vtt):
SubRip (SRT) Syntax:
1
00:00:01,500 --> 00:00:04,200
Welcome to our web architecture tutorial.
2
00:00:04,500 --> 00:00:07,800
Today we are building high-performance static tools. WebVTT (VTT) Syntax:
WEBVTT - Video Transcript
00:00:01.500 --> 00:00:04.200 align:middle line:80%
Welcome to our web architecture tutorial.
00:00:04.500 --> 00:00:07.800
Today we are building high-performance static tools. 2. Regex Patterns for Timecode & Index Stripping
To extract pure text paragraphs, a multi-stage regular expression engine strips file headers, sequence numbers, and timestamp lines:
- WebVTT Header:
/^WEBVTT.*$/gm - Timecodes:
/^\d2:\d2:\d2[\.,]\d3\s*-->\s*\d2:\d2:\d2[\.,]\d3.*$/gm - Sequence Numbers:
/^\d+$/gm - Cue Styling Tags:
/<[^>]*>/g
3. Normalizing Paragraph Breaks and Caption Tags
Subtitle tracks break sentences mid-thought to fit on-screen line lengths. After stripping timecodes, consecutive caption lines must be merged into natural sentences while preserving true paragraph breaks.
4. Pure JavaScript Client-Side Transcript Engine
Here is a production-ready JavaScript function for converting raw subtitle strings into continuous text:
function cleanSubtitleText(rawText) {
return rawText
// Remove WebVTT header lines
.replace(/^WEBVTT.*$/gm, '')
// Remove timecode lines (both SRT comma and VTT dot separators)
.replace(/^d{2}:d{2}:d{2}[.,]d{3}s*-->s*d{2}:d{2}:d{2}[.,]d{3}.*$/gm, '')
// Remove standalone numeric line counters
.replace(/^d+$/gm, '')
// Remove HTML cue styling tags (e.g. <v Speaker>, <b>, <i>)
.replace(/<[^>]*>/g, '')
// Collapse multiple blank lines into single line breaks
.replace(/
s*
/g, '
')
// Trim remaining leading/trailing whitespace
.trim();
} 5. Practical Use Cases for Subtitle Transcripts
- ✓ SEO Video Transcripts: Generating blog posts and indexable text transcripts from YouTube SRT downloads.
- ✓ AI Model Ingestion: Feeding clean podcast transcripts into LLM summarizers without timestamp noise.
- ✓ Content Editing: Reviewing spoken video scripts for copy editing and article syndication.