Convert MP3 to VTT Captions
Upload an MP3 and get back a properly paced WebVTT file, ready to drop into an HTML5 <track> tag, a webinar replay, or a course platform that expects .vtt instead of .srt.
Upload audio or video
MP3, WAV, M4A, AAC, OGG, OPUS and more · up to 100 MB on the free plan
or drag and drop it here
Record in your browser
Meetings and calls up to 30 minutes free — longer on paid plans
Start recordingFree account · 60 trial minutes, then 30 minutes a month · no credit card required
Plenty of caption tools assume you're editing desktop video and only ever export SRT. But if you're publishing to a website, embedding a lecture recording in an LMS, or uploading a webinar replay to a platform that reads the WebVTT spec, SRT often gets rejected outright — you need a real .vtt file with the right header and cue formatting. Transkio transcribes your MP3 and exports it as WebVTT directly, so you're not stuck converting formats by hand afterward.
Upload the MP3 file itself — Transkio doesn't fetch audio from a URL, so grab the recording from your webinar host, recorder, or export tool first. From there, Transkio's speech-to-text engine transcribes the audio, times each line to the source recording, and paces the resulting captions into cues sized for on-screen reading before you export the finished .vtt file.
WebVTT shows up specifically wherever web standards matter more than desktop editing conventions: a video embedded directly in a web page using the HTML5 <video> and <track> elements, a webinar replay hosted on a browser-based player, a lecture recording uploaded to a learning management system, or a course platform that validates uploaded captions against the WebVTT spec and rejects anything else. If your destination is a browser rather than a video editor, there's a good chance it wants .vtt specifically, not .srt.
The two formats look deceptively similar — both are plain text, both use numbered or sequential timed blocks of caption text — which is exactly why hand-converting one into the other is so error-prone. Small syntax differences (a comma instead of a period in the timestamp, a missing header line) are enough to make a file that looks right fail silently in a browser. Transkio sidesteps all of that by generating real WebVTT directly from the transcript, rather than relabeling an SRT file and hoping a browser doesn't notice.
SRT Doesn't Work Everywhere WebVTT Does
WebVTT is the caption format the HTML5 <track> element actually expects, and it's what most web-embedded video players, some webinar hosting platforms, and a number of LMS or online-course tools require for uploaded captions. SRT is common in desktop editors like Premiere or Final Cut, but it isn't part of any web standard — plenty of browser-based players won't parse it, or will silently fail to display anything.
That mismatch shows up most with recorded audio: a webinar session, a course lecture, a podcast episode meant for a web page. You've got an MP3, but the destination wants captions in a format the file doesn't come in, and manually retyping timestamps into VTT syntax for a 45-minute recording isn't a reasonable use of anyone's afternoon.
It gets worse when someone tries the obvious shortcut: renaming an .srt file to .vtt, or find-and-replacing commas with periods in the timestamps. That fixes the timestamp format but skips the required "WEBVTT" header line the spec demands at the top of the file, along with other small structural differences — the result often looks correct in a text editor and still fails to load in a browser, with no error message explaining why.
And because a webinar or lecture recording is usually saved as MP3 or another compressed audio format, there's no video track to extract captions from in the first place — the captions have to be generated from the audio itself, timed to it, and then formatted specifically as WebVTT rather than assumed to work as a generic subtitle file.
Upload the MP3, Export WebVTT Directly
Transkio accepts MP3 (along with WAV, M4A, AAC, OGG, OPUS, and FLAC) and transcribes it with word-level timing. When you export, choose VTT and you get a standards-compliant .vtt file with cues split at natural word boundaries and kept within readable length and duration limits — no manual reflowing needed before it goes live on a page.
Every transcript is editable before export too: click any line to jump the player to that moment, fix a misheard word or name, or rename a speaker, and the changes carry through to your VTT export. Transkio's free plan includes 60 trial minutes and then 30 minutes a month, with VTT export available at every plan tier — no card required to start.
The exported file includes the required WEBVTT header, correctly formatted period-based timestamps, and cues structured the way the specification expects — nothing to hand-fix afterward before dropping it into a <track> element or a course platform's caption upload field.
The same transcript exports as SRT and plain TXT too, so a single upload covers a web-ready VTT file, an SRT for a desktop editor, and a readable transcript document without transcribing the recording more than once.
What WebVTT Actually Is
WebVTT — Web Video Text Tracks — is a caption format defined by the W3C specifically for the web, meant to work with the HTML5 <track> element inside a <video> tag. Every valid .vtt file starts with the literal text "WEBVTT" as its first line, followed by a blank line, and then a sequence of cues: a timestamp range using periods for milliseconds (00:00:03.400 --> 00:00:06.800), followed by one or more lines of caption text, separated from the next cue by a blank line.
Beyond the basic timed-text structure, the WebVTT spec also defines optional features like cue identifiers, positioning and alignment settings, and voice spans that mark which speaker is talking — capabilities that go beyond what SRT was ever designed to support, since WebVTT was built for the more capable environment of a web browser rather than a 1990s desktop DVD-ripping tool.
None of those optional features are required for a caption file to work correctly — a valid WebVTT file can be as simple as the header plus a handful of timed cues, which is exactly what a straightforward MP3-to-VTT conversion produces.
WebVTT vs. SRT: The Real Syntax Differences
The two formats are close cousins, which is part of why confusing them is so easy. WebVTT requires the "WEBVTT" header line at the top of the file; SRT has no header at all. WebVTT timestamps use a period before the milliseconds (00:00:03.400); SRT uses a comma (00:00:03,400). SRT numbers every cue with a sequence number before the timestamp line; WebVTT cue identifiers are optional and rarely needed for a basic caption track.
Beyond syntax, WebVTT is a genuine web standard maintained by the W3C, with native browser support built into every modern browser's video element. SRT has no formal governing standard and became universal purely through decades of adoption in desktop editing software — which is exactly why it works everywhere on a desktop editing timeline and inconsistently everywhere on the web.
For a straightforward captioning need — get the words on screen, timed correctly — the practical difference for a viewer is invisible; both formats display as timed text. The difference only matters at the point of parsing: a strict web player checking for the WEBVTT header and period-based timestamps will reject a file that doesn't have them, no matter how correct the actual caption text is.
Where WebVTT Is Specifically Required
The clearest case is the HTML5 <track> element itself: <video><track kind="captions" src="captions.vtt" srclang="en" label="English"></video> only accepts WebVTT — there's no way to point that element at an .srt file and have a browser render it. Any custom video player built directly on the HTML5 video element inherits this same requirement.
Beyond the raw HTML element, a number of JavaScript video-player libraries used across the web, and some webinar hosting and course platforms, standardize on WebVTT as the caption format they accept for uploads, precisely because it's the format native to the browser environment they're built on. Where a platform accepts either format, WebVTT tends to be the safer default for anything web-embedded; where it explicitly requires WebVTT, SRT simply won't work.
This is different from a desktop editing context, where the reverse is often true — SRT is the more universally recognized format across Premiere, Final Cut, DaVinci Resolve, and similar tools, even though several also accept WebVTT. Which format actually matters comes down entirely to where the caption file needs to work, not which one is objectively "better."
Adding a VTT Caption Track to a Web Page
A basic implementation looks like this: <video controls src="lecture.mp4"><track kind="captions" src="lecture.vtt" srclang="en" label="English" default></track></video>. The kind="captions" attribute tells the browser this track is a caption track rather than, say, chapter markers or descriptions, srclang sets the language code, and default makes it load automatically rather than requiring a viewer to enable it manually.
If your content is audio-only rather than a video file — a webinar recording or podcast episode saved as MP3 — the same <track> element still works once you've paired the MP3 with a minimal video wrapper (even just a static image), which is a common pattern for publishing audio-only content anywhere a caption track is expected.
The exported VTT file from Transkio drops straight into the src attribute with no further editing required, assuming the file name and path match wherever you've hosted it alongside the video.
Speaker Labels in WebVTT: Voice Spans
WebVTT supports a feature SRT has no real equivalent for: voice spans, which mark which speaker is talking within a cue using a <v Name> tag. On the Elite plan and above, once speaker detection has identified each voice in a multi-speaker recording, exporting with speaker labels on wraps each caption's text in a voice span using that speaker's label or name — support for actually displaying that distinction visually depends on the specific player rendering the file, but the information is present in the exported file either way.
For a single-narrator webinar or lecture, this doesn't come into play — voice spans only add value once more than one person is speaking in the recording, and a solo-narrator VTT export stays as plain timed cues without them.
From Webinar Recording to a Captioned Web Replay
A typical version of this workflow: record or export the webinar as an MP3 (or pull the audio from a recorded video call), upload it to Transkio, review and correct the transcript, then export VTT. Pair that VTT file with the webinar's recorded video (or a simple static-image video wrapper if only audio was captured) and publish both to whatever page or platform hosts the replay.
The same pattern applies to course content: a recorded lecture saved as MP3, transcribed and exported as VTT, then uploaded alongside the lecture video to an LMS that expects a WebVTT caption file for accessibility and search purposes within the course platform.
VTT File Encoding and Common Gotchas
The single most common WebVTT bug is a missing or malformed header — the file has to start with exactly "WEBVTT" as the first line (optionally followed by a title or notes on the same line), with a blank line before the first cue. A file missing this, or with anything else on the first line, can fail to load in a strict player with no visible error.
The second most common issue, especially when converting by hand from SRT, is leaving commas in the timestamps instead of periods — a WebVTT parser expects 00:00:03.400, not 00:00:03,400, and will typically fail the entire cue or the whole file rather than gracefully handling the wrong punctuation.
A less obvious one is character encoding and byte-order marks: some text editors save files with a UTF-8 byte-order mark (BOM) at the very start, which can sit before the WEBVTT header and cause certain strict parsers to miss it entirely. Transkio's exports avoid all of these issues by generating the file programmatically from the transcript rather than by hand-editing or converting an existing file.
Choosing Between VTT, SRT, and Plain Text
It helps to think of the three export formats as answering three different questions rather than competing for the same job. Plain text answers "what was said" — a readable document for reference, search, or repurposing into blog posts and show notes, with no timing information at all. SRT and VTT both answer "what was said, and exactly when" — but for different playback environments.
A practical rule of thumb: if the destination is a web page, an HTML5 video embed, or a course platform's caption upload field, reach for VTT first. If the destination is a desktop editing timeline — Premiere, Final Cut, DaVinci Resolve — reach for SRT. If you're not sure which the destination expects, exporting both takes seconds since they come from the same reviewed transcript, and you can simply try the one that loads.
Some platforms genuinely accept either format interchangeably, in which case the choice comes down to whichever your existing workflow already uses. There's no accuracy or quality difference between the two — they carry the identical caption text and timing, just wrapped in different syntax.
Common Mistakes When Publishing VTT Captions
Beyond header and timestamp formatting, the most frequent real-world mistake is a file-extension or MIME-type mismatch — some web servers need to be configured to serve .vtt files with the correct text/vtt content type, and a misconfigured server can cause a browser to refuse the file even when its contents are perfectly valid WebVTT.
Another common one is a srclang attribute that doesn't match the actual language of the captions — this doesn't break playback, but it does mean accessibility tools and browser language-selection UI report the wrong language for the track, which matters more once a page serves the same video with multiple caption tracks in different languages.
And, similar to SRT, re-timing or trimming the source audio after generating a VTT file breaks the sync between the captions and the media — any edit to the audio after transcription means re-uploading the edited version and exporting a fresh VTT rather than trying to manually shift existing timestamps.
MP3 Quality and Caption Accuracy
The same rules that apply to any transcription apply here: caption accuracy tracks how clearly the speech was recorded, not the MP3's bitrate or the caption format it ends up exported as. A webinar recorded through a decent headset microphone in a quiet room will produce cleaner captions than the same session recorded through a laptop's built-in mic across a noisy room, regardless of which format the captions get exported to afterward.
For webinars and lectures specifically, cross-talk during a Q&A segment and audience members speaking from a distant or muted microphone tend to be the biggest accuracy factors — worth a slightly more careful review pass on those sections before publishing captions that will sit on a public or course-restricted page indefinitely.
WebVTT Styling and Layout Options
Beyond plain timed text, WebVTT supports optional cue settings that control where and how a caption appears on screen — position, alignment, and line placement can all be set per cue in the file. It also supports a small set of inline styling tags borrowed loosely from HTML: bold, italic, and underline markup can be applied to specific words or phrases within a cue if a player supports rendering them.
None of this is required for a working caption file. A straightforward MP3-to-VTT conversion doesn't need positioning hints or inline styling — the browser's default caption rendering handles placement and appearance perfectly well for the vast majority of webinar and lecture content. These options exist for cases like on-screen captions that need to avoid covering a speaker's face in a specific corner of the frame, which is a video-production concern rather than something an audio-only transcription workflow needs to solve.
Transkio's VTT exports keep to the simple, universally supported subset of the spec — clean cues with accurate timing — rather than adding styling that most destinations won't render anyway and that adds no value for an audio-derived caption track.
VTT Files and Search Engines
One underappreciated reason to publish a WebVTT track alongside embedded video is that some search engines and site crawlers can index caption text associated with a video, effectively making spoken content in a webinar or lecture recording searchable the same way visible page text is. This isn't guaranteed or something Transkio can control — it depends entirely on how a given search engine or platform handles caption files — but it's a commonly cited reason to add real captions to a video page beyond accessibility alone.
A plain-text transcript published on the same page accomplishes something similar and is generally the more reliable of the two for search visibility, since it doesn't depend on a search engine choosing to parse an associated .vtt file. Publishing both — the VTT track for the video player and a text transcript underneath for readers and crawlers — covers both bases without extra transcription work, since they come from the same upload.
Multi-Language Caption Tracks on One Page
A single video or embedded webinar can carry more than one <track> element, each pointing at a different .vtt file for a different language, letting a viewer choose their preferred caption language from the player's built-in menu. This is a common pattern for webinars or courses with an international audience — one source-language VTT plus one or more translated VTT files.
Transkio can produce the translated text starting on the Pro plan, but translation currently outputs a translated transcript document, not a re-timed translated VTT file — pairing translated text back to the original audio's timing to build a second caption track is a manual step, not an automatic export option today.
How It Works, Step by Step
- 1
Upload the MP3
Drag in your webinar, lecture, or podcast recording — up to 100 MB free.
- 2
Transkio Transcribes and Times It
Speech-to-text runs automatically, with cues paced at natural word boundaries for on-screen reading.
- 3
Review the Transcript
Click any line to jump the audio to that point, fix a misheard word, or rename a speaker before exporting.
- 4
Export as WebVTT
Download a standards-compliant .vtt file with the correct header, ready for a <track> element or a course platform upload.
- 5
Publish Alongside Your Video
Upload the .vtt file to your hosting platform or reference it in a <track> element, then confirm captions display correctly before you publish.
What You Get
See it work
See the Difference
A real example of how the same recording reads before and after Transkio.
Webinar clip, transcribed and captioned
ok so today were gonna go over uh the three main pricing tiers and uh what each one includes so lets jump right in
WEBVTT 00:00:00.000 --> 00:00:03.400 Okay, so today we're going to go over 00:00:03.400 --> 00:00:06.800 the three main pricing tiers and what each one includes. 00:00:06.800 --> 00:00:08.500 So let's jump right in.
Frequently asked questions
What's the actual difference between VTT and SRT?
WebVTT (.vtt) is the caption format defined by the W3C for use on the web — it's what the HTML5 <track> element reads natively in browsers. SRT (SubRip) predates the web standard and is more commonly used in desktop video editors. They look similar (timestamps plus text) but use different syntax, and many web players and some webinar or LMS platforms only accept one or the other.
Is VTT export available on the free plan?
Yes. VTT is included at every plan tier, including free — along with TXT and SRT. The free plan gives you 60 trial minutes, then 30 minutes per month, no card required.
Can I paste a link to my webinar recording instead of uploading the file?
No — Transkio works from an uploaded file or an in-browser recording, not a pasted URL. Export or download the audio from your webinar or recording platform first, then upload the MP3 directly.
Will the VTT captions be in a different language than the audio?
Transkio transcribes in the source language by default. Translated text transcripts are available starting on the Pro plan, but translation currently produces a translated text transcript only — not a translated VTT or SRT caption file.
How big of a file or how long of a recording can I upload?
Free accounts can upload files up to 100 MB. Paid plans raise those limits considerably — Ultra supports files up to roughly 8 GB and recording sessions as long as 8 hours, which comfortably covers a full-day webinar or lecture series.
Can I just rename an SRT file to .vtt instead of exporting a real one?
It won't reliably work. Renaming the extension doesn't add the required WEBVTT header line or fix the timestamp format (SRT uses a comma before milliseconds, WebVTT uses a period), so a renamed SRT file often fails to load in strict web players even though the extension looks correct.
Does WebVTT support showing which speaker is talking?
Yes, using a feature called voice spans, which mark each cue with the speaker's label. On the Elite plan and above, exporting with speaker labels on includes this in the VTT file — whether a given player visually displays it depends on that player's own support for voice spans.
Why won't my browser display the VTT file even though it looks correct?
Common causes: a missing or malformed WEBVTT header on the first line, a web server serving the file with the wrong content type instead of text/vtt, or commas left in the timestamps from a manual SRT conversion. Transkio's exports are generated correctly by default, which avoids all three.
Can I use the same VTT file in a video editor as well as on my website?
Many desktop editors do accept WebVTT alongside SRT, but support varies by tool. If you need the captions in both places and your editor doesn't accept .vtt, export SRT for the editor and VTT for the web from the same reviewed transcript — no need to transcribe twice.
Does the srclang attribute need to match the actual audio language?
It should. The srclang attribute on the HTML5 <track> element doesn't affect whether captions display, but it tells the browser and accessibility tools what language the track is in, which matters especially once a page has multiple caption tracks in different languages.
What happens if I edit the MP3 after generating a VTT file?
Any trim, cut, or speed change to the audio shifts the timing of everything after that point, breaking sync with the existing VTT file. Re-upload the edited audio and export a fresh VTT rather than trying to manually adjust the old timestamps.
Is WebVTT required for accessibility compliance on a webpage?
WebVTT is the native captioning format for HTML5 video and a practical, widely supported choice for making web video accessible, but Transkio doesn't carry any specific accessibility certification. Whether your page meets a particular legal or regulatory accessibility standard is something to verify against your own requirements.
Does a course platform or LMS need anything special beyond the .vtt file itself?
That depends entirely on the specific platform — some accept a direct file upload for captions, others expect the VTT referenced via an HTML5 <track> element in embedded video. Check your platform's own documentation for exactly how it expects captions to be attached; the VTT file itself is standard either way.
Can I get both a transcript document and a VTT file from the same MP3 upload?
Yes. Upload once, review the transcript, and export TXT for a readable document and VTT for web captions separately — both come from the same transcription, so there's no need to process the audio twice.
Does the MP3's audio quality affect how the VTT captions look?
Caption accuracy — the words themselves — depends on recording clarity, not the MP3's bitrate. A quiet room and a consistent, close microphone matter far more to the final caption quality than any audio-quality setting in your recording software.
Can I add positioning or styling to the VTT captions?
WebVTT supports optional cue positioning and basic inline styling, but Transkio's exports stick to clean, simply timed cues rather than adding styling most players won't render anyway. For an audio-derived caption track, default caption placement covers the vast majority of use cases without any manual adjustment.
Can one video have caption tracks in more than one language?
Yes — the HTML5 <track> element supports multiple caption tracks, each pointing at its own .vtt file, letting viewers pick a language from the player's menu. Transkio can produce translated transcript text for a second language starting on the Pro plan, but building a second, re-timed VTT file from that translation is a manual step today, not an automatic export.
Will publishing a VTT file help my webinar page show up in search results?
It can, since some search engines index caption text associated with embedded video, but this isn't guaranteed and depends on the platform. Publishing a plain-text transcript alongside the video is generally the more reliable way to make the spoken content searchable, and both come from the same Transkio upload.
Do I need a video file to use a VTT caption track, or does audio-only work?
The HTML5 <track> element technically attaches to a <video> element, so audio-only content usually gets paired with a minimal video wrapper (even a static image) to host the caption track. The VTT file itself is generated the same way regardless — from the MP3's transcript and timing.
Should I export VTT, SRT, or just a plain text transcript?
It depends on where the captions need to live. Choose VTT for web embeds, HTML5 video, and most course or webinar platforms; choose SRT for desktop editors like Premiere or Final Cut; choose plain text for a readable document, show notes, or repurposing content elsewhere. All three export from the same reviewed transcript, so there's no cost to generating more than one.
Can I edit the VTT file's text after exporting it?
You can open and edit the plain-text .vtt file in any text editor afterward, but it's much easier to make corrections in Transkio's transcript editor before exporting — fixes there apply cleanly to every export format at once, rather than needing to hand-edit timestamps and text separately in a raw VTT file.
Does Transkio support other subtitle or caption formats besides SRT and VTT?
Transkio's export formats are TXT, SRT, and VTT on every plan, with JSON and DOCX added starting on paid plans for structured data and document use. That covers the formats actually required by web players, desktop editors, and course platforms in practice.
Is there a limit on how many MP3 files I can convert to VTT?
There's no cap on the number of files — the limit is total minutes transcribed per month. Free accounts get 60 trial minutes then 30 minutes monthly; paid plans raise that considerably, and Ultra removes the monthly minute cap entirely.
Do I need any technical or coding knowledge to use the exported VTT file?
No coding is required to generate the file itself — Transkio produces a complete, ready-to-use .vtt file. Adding it to a web page does require pointing an HTML5 <track> element at the file, which is a small piece of markup most webinar or course platforms handle for you through their own upload interface rather than requiring you to write it by hand.
Can I preview the captions before publishing them?
Yes — review the transcript inside Transkio's editor before exporting, where clicking any line jumps the audio player to that exact moment so you can confirm the text is accurate. Once published, most video platforms also offer a preview of the caption track as it will appear to viewers, worth checking once before a webinar replay or course lecture goes fully live.
Do speaker names carry over automatically into the VTT voice spans?
Once you rename a speaker in Transkio's transcript editor — replacing a generic "Speaker 1" label with an actual name — that name is used consistently throughout the transcript and carries into the voice spans of the VTT export, so the caption file reflects the same names you assigned during review.
Does Transkio keep a copy of the original MP3 after I export the VTT file?
Yes — your uploaded audio and its transcript stay in your account after processing, so you can come back later to re-export in a different format or make further edits without re-uploading. You can delete a recording and its transcript from your account at any time.
Related Tools and Guides
Turn Your MP3 Into Web-Ready VTT Captions
Upload your webinar or lecture recording and export a properly paced WebVTT file in minutes. Free plan includes 60 trial minutes, no card required.
Transcribe for free- 30 free minutes, no card required
- Transcripts in minutes, not hours
- 50+ languages