October 2, 2026 · 11 min read
TTML Master: Caption File Formats for Video Producers
Use a TTML/DFXP master as your canonical captions. Export and test SRT/WebVTT per platform, use SCC/EBU for broadcast, or QR live captions for events.

Use SRT for maximum compatibility, WebVTT for browser-native styling and positioning, TTML or DFXP for production and interchange work, and SCC for CEA-608 broadcast workflows. Each format solves a different problem. Picking the wrong one causes dropped styling, broken timing, or a file your platform simply won’t read.
TL;DR:
- WebVTT files must be served as UTF-8 with the correct MIME type (
text/vtt) to ensure styling and positioning are preserved in browsers.- Captions should be exported from a master TTML or DFXP file to retain styling, positioning, and metadata, rather than creating each format separately.
- For broadcast workflows, SCC and EBU-STL formats are non-interchangeable and must match platform-specific delivery specifications.
- Testing caption files on the actual playback platform and verifying encoding, MIME type, and timing is crucial before publication.
- Live captioning requires real-time delivery methods, such as QR-code-based apps, since sidecar files can’t support immediate streaming scenarios.
Table of Contents
- Common caption and subtitle formats at a glance
- What each format actually supports under the hood
- How to go from raw transcript to tested caption file
- How to test captions before they go live
- Why a master file beats patching exports forever
- The real lesson in choosing a caption format
- Live captioning for events that can’t wait for a file
- FAQ
- Sources
Common caption and subtitle formats at a glance
Before you pick a format, know what each one is actually built for. Here’s the quick version:
- SRT (.srt): Plain text, widest compatibility, the safe default for most platforms including YouTube.
- WebVTT (.vtt): Browser-native, supports styling and positioning, required for HTML5 video.
- TTML/DFXP (.ttml/.dfxp): XML-based, built for production and interchange, preserves styling and positioning across systems.
- SCC (.scc): Represents CEA-608 data directly, the standard for broadcast captioning pipelines.
- EBU-STL (.stl): A European broadcast binary format, not interchangeable with SCC despite similar use cases.
- SBV (.sbv) and SAMI (.smi): Older or platform-specific formats, mostly legacy at this point.
Encoding matters as much as the extension. WebVTT files must be UTF-8 and served as text/vtt, according to MDN’s Web Video Text Tracks documentation. YouTube has its own encoding rule too: it supports basic SRT, but only when the file is plain UTF-8, per YouTube’s help documentation. Styling and positioning you built into a rich master file often get dropped the moment you export to a simpler delivery format, so check the result, not just the export.
What each format actually supports under the hood
SRT is the plainest of the bunch: numbered cues, timestamps in HH:MM:SS,mmm format, and a line or two of text. No header, no styling blocks, no positioning data. The Library of Congress format description treats it as a widely adopted sidecar format, but one you should think of as a compatibility deliverable rather than a format to build your workflow around. It’s simple because it has to be. That simplicity is also its ceiling.
WebVTT starts with a literal WEBVTT line at the top of the file, and from there it can carry a lot more than SRT ever will. STYLE and REGION blocks let you control appearance. Cue settings handle position and alignment. According to MDN’s WebVTT reference, the format also supports multiple track kinds, meaning the same structure can serve captions, subtitles, chapters, or time-aligned metadata depending on how you tag it. Browsers expose these as separate tracks a viewer can toggle, which is part of why WebVTT is the format HTML5 video actually expects. None of this works, though, if the MIME type and encoding are wrong. A WebVTT file served without text/vtt or saved outside UTF-8 will often get silently ignored by the player, even when the cue syntax inside it is flawless.
TTML and DFXP take a different approach entirely: full XML structure, with named profiles like SMPTE-TT or iTT that define exactly how styling and timing should behave. The Library of Congress entry on TTML describes it as a format built for production and interchange, the kind of file that moves between editing systems, caption vendors, and broadcast pipelines without losing its structure. This is why a TTML master makes sense as your canonical file: it holds the detail that gets stripped out of simpler formats. Convert TTML down to SRT and you’ll typically lose speaker labels, positioning, and styling in the process. If a speaker label matters for the SRT version, write it into the text payload manually, because the conversion won’t carry it over for you.
SCC and CEA-608, along with EBU-STL, live in broadcast workflows where the format has to match the pipeline exactly. SCC represents CEA-608 data directly, and platform guidance points to it specifically for that kind of workflow, per YouTube’s supported formats page. EBU-STL is a separate binary format used in European broadcast systems, and the two are not interchangeable. Swapping one for the other without a delivery spec in hand is a common and avoidable mistake.

The recurring pitfalls across all of these: files saved with a byte-order mark that breaks parsing, mismatched language tags that confuse multi-track setups, and styling that gets quietly dropped the moment a file lands on a platform that doesn’t honor it.
How to go from raw transcript to tested caption file
A caption file is only as good as the process behind it; understanding why captions matter in video production can help justify this workflow. Here’s the pipeline that holds up across most projects:
- Transcribe the audio, capturing speaker changes and meaningful non-speech sound like music cues or applause.
- Divide the transcript into cues, keeping line length and reading speed comfortable. Shorter cues that match natural speech pauses read better than long blocks of text.
- Build or update your master file, ideally TTML or DFXP, since it preserves styling, positioning, and metadata that simpler formats drop.
- Export your delivery formats (SRT, VTT) from that master rather than authoring each one separately by hand.
- Check the export for lost metadata. Conversions strip things, so don’t assume the output matches your intent.
- Test on the actual destination player, not just in a text editor, since syntax that looks correct can still fail to render.
A sidecar file is different from burned-in captions. Sidecar files stay toggleable and separate from the video; burned-in captions are permanently part of the picture and can’t be turned off, a distinction University of Texas’s captioning guidance lays out clearly. Keep that in mind when a client asks for “captions” without specifying which kind.
Before you deliver anything, run through a short checklist: correct filename and extension, UTF-8 encoding, correct MIME type for web delivery, accurate language tags, and any platform-specific rule your destination enforces. YouTube’s support page spells out several of its own constraints worth checking against before upload.
Pro Tip: Keep your TTML master under version control, so every SRT or VTT export traces back to a single source instead of five slightly different hand-edited files.
For a closer look at cue timing specifically, see this guide on timing captions for live and prerecorded video.
How to test captions before they go live
Testing is not optional, and it’s where most caption problems actually surface. For HTML5 video, set up the <track> element with the right kind, srclang, and label attributes, and watch for CORS issues if the caption file lives on a different domain than the video, a requirement MDN’s track element guide covers in detail.
Common failure patterns and what they usually mean:
- Missing cues: often a malformed timestamp or a stray character breaking the parser.
- Garbled characters: almost always an encoding mismatch or a leftover byte-order mark.
- Timing drift: frame rate mismatches between the source video and the caption file.
- Track not loading at all: wrong MIME type, missing CORS headers, or an incorrect file path.
Open the raw file in a plain text editor first. If the structure looks wrong there, it will look wrong everywhere else too.
Why a master file beats patching exports forever
Sidecar files make sense for prerecorded content you control end to end: upload once, test once, done. Live events don’t work that way. There’s no file to proofread in advance, which means the captioning has to happen in real time, over a stream or a direct connection to the room.
The practical answer is to keep both approaches in your toolkit. Maintain a rich master timed-text file for anything prerecorded, and export tested derivatives per platform. For live events, you need a different delivery method entirely. Live Caption AI provides QR-code-based live caption delivery to attendees’ own phones, with domain-specific vocabulary models and real-time translation support, a different problem than sidecar files solve.
The real lesson in choosing a caption format
Most guides treat format choice like a single decision: pick one file type and move on. That’s backwards. The format that matters most is the one you never deliver: your master file. Everything downstream, SRT for YouTube, VTT for your website player, gets exported from it, not authored by hand each time.

The industry’s obsession with SRT as a default is understandable but overstated. SRT is a fine delivery format and a poor place to do any real work, since it can’t hold styling, positioning, or speaker metadata in the first place. Treat it as the output of a process, not the process itself.
The other thing creators underestimate: testing on the actual destination player matters more than getting the syntax theoretically correct. A WebVTT file can be flawless by spec and still fail silently because of a MIME type mismatch or an encoding slip. Build the habit of testing before you ship, every time, not just when something looks off.
If you’re producing anything with recurring captioning needs, a TTML master and a short export checklist will save more time than any converter tool.
— Ryan
Live captioning for events that can’t wait for a file
Sidecar files work great until the event is happening right now. For live services, conferences, or multilingual sessions, Live Caption AI turns any phone into a caption receiver through a QR code: no hardware, no stenographer. Medical, Worship, and Finance vocabulary models plus your own Custom Terms, translation into one of 29 languages per session, and session audio is never stored.

Plans start with a free tier, and the Professional plan is available for a monthly subscription for teams ready to add cloud-based features. For settings with specialized terminology, the professional captioning page walks through domain vocabulary options in more depth.
FAQ
What can open an SRT file?
Any plain text editor can open an SRT file, since it’s just structured text with no binary encoding. Most video editing software, media players, and caption tools also support direct SRT import for editing or playback.
Does YouTube use SRT or VTT?
YouTube accepts both, but its own documentation recommends basic SRT for most uploaders, with the requirement that the file be plain UTF-8, according to YouTube’s supported caption formats page. WebVTT is also supported, though advanced styling features may not render as expected on the platform.
Which is better, SRT or PGS?
They solve different problems: SRT is a text-based sidecar format meant for editable, toggleable captions, while PGS is an image-based subtitle format typically used on Blu-ray discs. For most web and production workflows, SRT or WebVTT is the more practical and editable choice.
What is an SRT or VTT file?
Both are sidecar caption files that sit alongside a video rather than being baked into the picture. SRT uses simple numbered cues with timestamps, while WebVTT adds a formal header and supports styling, positioning, and multiple track types, as described in MDN’s WebVTT documentation.
Sources
- SubRip Subtitle format (SRT) — Library of Congress
- Web Video Text Tracks Format — MDN
- Supported subtitle and closed caption files — YouTube Help