Skip to content
Transkio

Convert MP3 to SRT Subtitles

Upload an MP3 voiceover, podcast clip, or any audio file and get back a timed SRT file ready to drop into your video editor. SRT export is on the free plan.

Upload audio or video

MP3, WAV, M4A, AAC, OGG, OPUS and more · up to 100 MB on the free plan

or drag and drop it here

Record in your browser

Meetings and calls up to 30 minutes free — longer on paid plans

Start recording

Free account · 60 trial minutes, then 30 minutes a month · no credit card required

A lot of short-form video creators record their voiceover separately as an MP3 before they ever open an editor — a script read into a phone or a mic, exported as audio, then synced to footage later in Premiere, CapCut, or Descript. The problem is that MP3 file has no captions attached to it, and typing out timestamps by hand for every line is slow and easy to get wrong.

Transkio turns that MP3 into a subtitle file automatically. Upload the audio, let the transcription engine listen to it, and export the result as an SRT with cue numbers and timecodes already in place — no manual timing, no separate captioning tool, and it works across dozens of languages.

It's not just short-form video creators who end up needing an SRT from audio, either. A podcaster turning an episode into a video upload for YouTube — often just a static cover image or waveform animation behind the audio — needs captions to make that upload watchable with the sound off and more discoverable in search. Someone repurposing a long interview into a handful of short, captioned clips for social media needs the same thing. And anyone making an audio-only recording more accessible, whether that's a lecture, a sermon, or an internal training recording, often finds that a caption file is the more usable end product, even when there's no video involved at all.

None of these situations start with a subtitle file — they start with an MP3. The gap between the two is exactly what this page exists to close: upload the audio, and get a properly timed .srt file back, without opening a caption editor or typing a single timecode by hand.

An MP3 Voiceover Doesn't Come with Captions Built In

If you record narration or a voiceover script as a standalone MP3 — separate from the video edit itself — you're left with audio and nothing else. Most video editors can burn in or overlay captions, but only if you feed them a subtitle file with real timecodes, not just a block of text.

Writing that timing by hand means scrubbing through the clip second by second, guessing where each line starts and ends, and redoing it every time you trim the audio. For a 60-second Reel that might be tolerable once; for a weekly upload schedule it isn't.

A plain transcript doesn't solve this either. Even a perfectly accurate, well-punctuated text transcript is just a block of words with no timing information attached — useful for reading, useless for placing text on top of video frame by frame. Captioning software and video editors need a structured file that says exactly which words appear on screen between which two timestamps, and that structure doesn't exist until something builds it.

There's also a quieter cost to skipping captions altogether: a large share of video is watched muted, in feeds, in offices, on transit. Audio-only content that gets turned into video without captions loses most of its audience the moment autoplay starts with the sound off, regardless of how good the narration actually is.

Upload the MP3, Download a Ready-to-Import SRT

Upload your MP3 (WAV, M4A, AAC, OGG, OPUS, and FLAC are accepted too, along with video files if you ever record straight to video). Transkio transcribes the audio and paces it into properly formatted subtitle cues — split at natural word boundaries so no line lingers too long or crams too much text on screen — then lets you export as an SRT with a single click.

Before exporting, the transcript is fully editable: click any line to jump the audio to that moment, fix a misheard word, or rename a speaker. What you download is the corrected version, not the raw first pass.

The same underlying transcript also exports as plain TXT or as WebVTT, so a single upload covers a reading transcript, an SRT for a desktop editor, and a VTT for a web player without transcribing the same audio three separate times. Whichever format you need this time, the others stay available from the same transcript later.

Because timing comes directly from the speech-recognition model rather than a guess, cue boundaries land close to where the words are actually spoken — no manual scrubbing to find where a sentence starts, and no drift over the length of a long recording the way hand-timed captions can accumulate.

What an SRT File Actually Is

SRT stands for SubRip Text, named after the SubRip software that originally defined the format for ripping subtitles off DVDs. Structurally, it's about as simple as a file format gets: plain text, split into numbered blocks called cues, each one containing a sequence number, a start and end timecode in the format hours:minutes:seconds,milliseconds, and one or more lines of caption text, with a blank line separating each cue from the next.

There's no video, no styling information, and no metadata beyond the timing and the text itself — which is exactly why it's so widely supported. Almost any software that displays captions, from professional editors to a phone's built-in video player, can parse that structure without needing a specialized library or a complex spec to implement.

That simplicity is also its limit. SRT has no built-in way to mark who's speaking, no native support for styling like bold or color, and technically no formal standard body behind it the way WebVTT has the W3C — it became universal through sheer adoption rather than a formal specification. In practice, none of that stops it from being the single most widely accepted caption format across desktop editors and upload platforms today.

Why Someone Needs Captions for Audio-Only Content

The most common path is turning a podcast into a video upload. Podcast hosting has increasingly moved toward video-first platforms, and even a simple video — a static cover image, a waveform animation, or a talking-head recording — needs captions to hold attention with sound off and to be indexed properly by a platform's search and recommendation systems. An MP3-only podcast has no captions to draw on, so they have to be generated from the audio directly.

Repurposing is another common driver: pulling the strongest three minutes out of an hour-long interview and turning it into a captioned short-form clip for social platforms. The source is audio (or video with audio that matters more than the picture), and the destination format all but requires burned-in or overlay captions, since short-form platforms are overwhelmingly watched muted by default.

Accessibility is a third, less flashy but genuinely important reason. A lecture, sermon, or training recording that only exists as audio excludes anyone who's Deaf or hard of hearing, or anyone who simply can't have sound on where they are. Generating a caption file from the recording — even one that's never turned into a formal video — gives that audio a text-timed equivalent that a captioning-aware player or an accessibility tool can actually use.

SRT vs. a Plain-Text Transcript: Different Jobs

A transcript and an SRT file often come from the exact same transcription, but they solve different problems and it's worth being clear about which one you actually need before exporting. A plain-text transcript is meant to be read start to finish, or searched, or pasted into a document — a script, an article, a set of notes. It has no concept of on-screen timing because it isn't meant to appear on screen at all.

An SRT file is meant to be consumed a few seconds at a time, synced to a video frame, by someone reading while also watching. That constraint shapes everything about how it's formatted: lines have to be short enough to read in the time they're displayed, cues have to break at natural pauses rather than arbitrary word counts, and text that would read perfectly well in a document can feel rushed or cramped once it's paced against video.

The practical rule: if a human is going to read the content as a document, export TXT. If the words need to appear synced to video or audio playback in a captioning-aware player, export SRT (or VTT, covered on its own page). Both come from the same underlying transcript, so there's no need to choose upfront — export whichever one the destination actually needs, and switch later if requirements change.

Where an SRT File Actually Gets Used

Desktop video editors are the most common destination — Premiere Pro, Final Cut Pro, DaVinci Resolve, CapCut, and Descript all import standard .srt files directly, letting you place, style, and burn in captions as part of the edit rather than typing them from scratch inside the editor.

Upload platforms are the other major use: YouTube accepts an .srt file as a caption track alongside a video upload, and so do most other major video hosts. Uploading a caption file this way is generally more accurate than relying on a platform's own auto-captions, since the timing and wording already reflect a transcript you've reviewed and corrected yourself.

Beyond editors and upload platforms, SRT files also work in standalone media players like VLC, which can load a .srt file alongside a video and display synced subtitles without any burning-in step at all — useful for previewing captions before committing to a final export, or for sharing a video and its captions as two separate files.

How Caption Pacing Actually Works

Good captions aren't just correct text placed at roughly the right time — they're paced so an average reader can actually finish each line before it disappears. That means limiting how many characters appear per line, keeping each cue on screen long enough to be read but not so long that it lags behind the audio, and breaking at natural phrase boundaries instead of mid-sentence wherever possible.

Transkio's transcription pipeline builds cues with these constraints already applied, splitting at natural word boundaries and keeping each cue within sensible character and duration limits rather than dumping the whole transcript into oversized blocks. That's the difference between a technically valid SRT file and one that's actually pleasant to read while watching.

None of that pacing is locked in place, though — because the underlying transcript is editable before export, you can adjust wording, fix a name, or tighten up a line that reads awkwardly once it's paced against the video, and the exported SRT reflects whatever the transcript says at the moment you download it.

From Podcast Episode to Captioned Video: A Common Workflow

A typical version of this workflow looks like: record or export the podcast episode as an MP3, upload it to Transkio, review and correct the transcript, then export SRT. Separately, pair that MP3 with a static image or a simple waveform animation in a video editor of choice, and drop the SRT in as the caption track. The result is a captioned video ready for YouTube or any platform that expects a video file rather than a raw audio upload — all without retyping a single line of the episode.

The same pattern applies to a livestream or webinar that was only recorded as audio: export the recording as MP3, transcribe it, and generate SRT for whatever video wrapper the recording eventually gets published in. The transcription step is identical regardless of what happens to the video side afterward.

Multiple Speakers in a Caption File

A caption track with more than one speaker on screen can still tell them apart. On the Elite plan and above, once speaker detection has identified who's talking in each segment, exporting with speaker labels turned on prefixes each caption line with the speaker's name or label — useful for an interview clip or a two-host podcast segment where it matters to a viewer which voice is talking.

For a solo voiceover or single-narrator recording, none of this applies — speaker labels only add value once there's more than one voice in the recording to distinguish, and a single-speaker SRT export stays clean with just the caption text and timing.

Encoding and Compatibility Quirks Worth Knowing About

SRT files are plain text, which sounds simple until a caption with an accented character, a curly quote, or a non-English name gets garbled after upload to a platform that expected a different text encoding. Transkio exports SRT files in standard UTF-8 encoding, which is what virtually every modern editor and upload platform expects, so this is rarely something you need to think about — but if a caption file from another source ever displays oddly, mismatched encoding is usually the first thing to check.

Line-break handling is another quiet source of bugs when captions are hand-edited in a plain text editor rather than a purpose-built tool — inconsistent line endings between operating systems can occasionally confuse a strict parser. Editing captions inside Transkio's transcript editor sidesteps this entirely, since the export is generated fresh from the corrected transcript rather than a manually patched text file.

Reviewing Captions Before They Go Live

Captions get scrutinized differently than a private transcript does, because they're visible to every viewer, permanently, on the video itself. A misheard word in a document you're reading privately is a minor annoyance; the same mistake burned into a public video's captions is a mistake an audience actually sees. It's worth treating the review pass for an SRT export a little more carefully than you might for a transcript meant only for internal use.

The editor makes that review fast rather than tedious: click any caption line to jump the audio to that exact moment, listen, and correct it in place. Doing that pass once before exporting is almost always faster than fixing captions after they're already burned into an exported video.

SRT for Recurring, Weekly-Upload Workflows

For a creator publishing on a regular schedule — a weekly podcast, a daily short-form drop — the real value of automating MP3-to-SRT conversion isn't any single episode, it's what it removes from a recurring workload. Hand-timing captions for one video is tedious; hand-timing them every single week, on top of everything else involved in publishing on a schedule, is where manual captioning workflows tend to quietly get skipped altogether, and the content ships without captions as a result.

Because the transcription and export steps take a few minutes regardless of how many times you've done it before, folding this into a publishing routine doesn't compound in effort the way manual timing does. The same upload, review, export sequence works identically on episode one and episode two hundred.

Naming and Pairing an SRT File With Your Video

Most video editors and upload platforms match a subtitle file to a video by file name and by importing it as an explicit caption track, not by anything embedded in the SRT itself — so it's worth naming the exported .srt file to clearly match its video (matching base filenames, e.g. episode-42.mp4 and episode-42.srt, is the convention several editors auto-detect) before you drop it into a project or upload it alongside a video file.

If you're exporting an MP3-derived SRT to pair with a video that doesn't exist yet — say, you'll build the video around a static image after the fact — it's fine to export the SRT first and hold onto it until the video is ready. The timecodes are anchored to the audio itself, so as long as the final video uses that same audio track without re-timing it, the captions will still line up correctly.

SRT for Accessibility, Without the Certification Claims

Captions are one of the most concrete things a creator can do to make audio content more accessible — a viewer who's Deaf or hard of hearing, watching in a noisy or sound-sensitive environment, or simply reading along while multitasking all benefit from an accurate caption track the same way. Generating an SRT from an MP3 recording is a genuinely useful accessibility step for content that would otherwise only exist as sound.

That said, it's worth being precise about what this does and doesn't guarantee: Transkio produces an accurate, reviewable caption file, but it doesn't carry any specific accessibility certification, and meeting a particular legal or regulatory accessibility standard for your content is your responsibility to verify against whatever requirements apply to you. Review the exported captions the same way you would any transcript before treating them as a compliance deliverable.

MP3 Quality and What It Means for Caption Accuracy

Caption accuracy tracks the same variables as transcript accuracy generally: how close the microphone was, how much background noise crept in, and how many people were talking at once, rather than the MP3's bitrate or file size. A modest, well-recorded voice memo compressed at a standard bitrate will produce cleaner captions than a poorly recorded file saved at a needlessly high bitrate.

For scripted voiceover specifically — the most common source of MP3-to-SRT conversions — recording in a quiet room with a consistent mic distance is the single highest-leverage thing you can do before uploading. It matters more than any setting inside a recording app, and it's the difference between a transcript that needs a light proofread and one that needs a heavier editing pass before the captions are ready to publish.

Common Mistakes When Turning Audio Into Captioned Video

The most frequent one is trimming or re-timing the audio after generating the SRT file, then wondering why the captions no longer line up — any cut, speed change, or added silence shifts every timestamp after that point. If the audio changes, regenerate the transcript and export a fresh SRT rather than trying to manually shift an old one.

A second common mistake is skipping the review pass because the audio was recorded from a clean, scripted read and assumed to transcribe perfectly. Even well-recorded, scripted narration produces occasional misheard words, especially with brand names, numbers, or uncommon terms — a quick pass through the transcript before export catches these before they're visible on screen to an audience.

A third is exporting captions before deciding on final video length — if a video will ultimately be trimmed down from a longer recording, it's usually simpler to trim the audio itself, re-export the SRT from the trimmed version, and skip trying to manually delete and renumber cues in a text editor afterward.

How It Works, Step by Step

  1. 1

    Upload the MP3

    Drag in the audio file — up to 100 MB free — or any other supported audio or video format.

  2. 2

    Transkio Transcribes and Times It

    Speech-to-text runs automatically, with cues already paced at natural word boundaries.

  3. 3

    Review the Transcript

    Click any line to jump the audio to that point, fix a misheard word, or rename a speaker before exporting.

  4. 4

    Export as SRT

    Download a standards-compliant .srt file ready to import into your video editor or upload alongside your video.

What You Get

SRT Export on the Free Plan

No paywall on the core feature. Free accounts get 60 trial minutes, then 30 minutes a month, exporting in TXT, SRT, VTT — no card required to start.

Cues Paced for Readability

Captions are split at natural word boundaries with sensible character and duration limits per cue, so lines don't blow past what a viewer can read in the time it's on screen.

Built for the Editor You Already Use

The exported .srt file imports directly into Premiere, CapCut, Descript, DaVinci Resolve, or any tool that reads standard SubRip subtitles — no reformatting needed.

Fix Mistakes Before You Export

Click any line in the transcript to jump the audio to that exact spot, correct a word, or rename a speaker — the SRT you download reflects your edits, not the raw transcription.

Handles More Than MP3

Accepts WAV, M4A, AAC, OGG, OPUS, and FLAC audio as well, plus common video formats, up to 100 MB per file on the free plan.

Speaker Labels When You Need Them

Recording an interview or multi-voice clip as MP3? Speaker detection separates who said what, available from the Elite plan up.

One Transcript, Three Caption-Friendly Exports

The same reviewed transcript exports as TXT, SRT, or VTT, so switching formats later doesn't mean transcribing the audio again.

Standard UTF-8 Encoding

Exports use standard text encoding that upload platforms and editors expect, so accented characters and non-English names display correctly.

See it work

See the Difference

A real example of how the same recording reads before and after Transkio.

Short-form voiceover clip, transcribed and timed

Before — raw auto-captions

so today im gonna show you three things that changed my morning routine and honestly the third one surprised me the most so stick around

After — Transkio transcript

1 00:00:00,000 --> 00:00:03,200 So today I'm going to show you three things 2 00:00:03,200 --> 00:00:06,400 that changed my morning routine, and honestly 3 00:00:06,400 --> 00:00:09,000 the third one surprised me the most. Stick around!

Frequently asked questions

What's the difference between a transcript and an SRT file?

A plain transcript is just the spoken words as continuous text — useful for reading or reuse, but a video editor can't place it on screen. An SRT file breaks that same content into numbered cues, each with a start and end timecode, which is the format video editors expect for captions.

Is SRT export really available on the free plan?

Yes. SRT is one of the export formats on every plan, including free, alongside TXT and VTT. Free accounts get 60 trial minutes and then 30 minutes a month, with no card required to sign up.

Do I have to upload an MP3, or can other audio formats work?

MP3 is fully supported, along with WAV, M4A, AAC, OGG, OPUS, and FLAC audio, plus MP4, MOV, WebM, MKV, and MPEG video if you ever record straight to video instead of audio-only.

Can I edit the captions before I export the SRT?

Yes. The transcript is editable before you export — click a line to jump the audio to that point, correct any misheard words, or rename speakers. The SRT you download reflects those edits.

Which video editors can open the SRT file?

Any editor that supports standard SubRip subtitles, which covers Premiere Pro, CapCut, Descript, DaVinci Resolve, Final Cut Pro, and most others — you import the .srt file directly alongside your footage.

Why would I want SRT captions for something that's only audio, with no video?

Two common reasons: you're about to pair the audio with a simple video (a static image, a waveform animation) for a platform like YouTube that expects video, or you want an accessible, text-timed version of an audio-only recording for anyone who can't or doesn't want to listen with sound.

Do the captions include speaker names for an interview or multi-host podcast?

Yes, when speaker detection is on. On the Elite plan and above, each caption line can be prefixed with the speaker's label or name, which helps a viewer follow who's talking in a multi-voice recording.

Can I get a translated SRT file in a different language?

Not currently. On the Pro plan and up, you can generate a translated text transcript of the recording, but subtitle exports (SRT and VTT) stay in the audio's original language — translated caption files aren't available yet.

What's the actual file format inside an .srt file?

Plain UTF-8 text, structured as numbered cues: a sequence number, a start and end timecode in hours:minutes:seconds,milliseconds format, one or more lines of caption text, and a blank line before the next cue. No video, images, or styling are embedded in the file itself.

Will the SRT file work if I upload it directly to YouTube?

Yes. YouTube accepts standard .srt caption files uploaded alongside a video, and using a reviewed, corrected SRT this way is generally more accurate than relying on YouTube's own automatic captions.

How is this different from just letting YouTube auto-generate captions after I upload?

YouTube's auto-captions are generated for live viewing and can't be corrected before they're shown, have no punctuation worth relying on, and don't separate speakers. Generating and reviewing an SRT beforehand means the captions your audience sees are accurate and properly formatted from the moment the video goes live.

Can I burn the captions directly into the video instead of using a separate SRT file?

Transkio exports a standard .srt file rather than a burned-in video — burning captions permanently into the video frames happens in your video editor, using the SRT file as the caption source. Most editors support this as a straightforward export or render option.

What happens if my MP3 has long pauses or silence in it?

Silence and pauses don't generate captions, which is correct behavior — an SRT file only needs cues where there's actual speech to caption. Long gaps between spoken sections simply appear as gaps between cues in the timeline.

Is there a limit to how long an MP3 can be for SRT export?

File size limits scale by plan — 100 MB on Free up to 8 GB on Ultra — which comfortably covers everything from a short voiceover clip to a multi-hour podcast episode.

Can I convert the same MP3 to VTT instead of SRT?

Yes, from the same uploaded file and the same reviewed transcript — export SRT, VTT, or plain TXT without re-transcribing. The mp3-to-vtt page covers when WebVTT specifically is the format you need.

If I trim or edit the MP3 after generating an SRT, do the captions still line up?

No — any cut, trim, or speed change shifts the timing of everything after that point in the audio. If you edit the source audio, re-upload the edited version and export a fresh SRT rather than reusing the old one.

Do I need to name the SRT file a specific way for my editor to recognize it?

Most editors auto-detect a subtitle file when it shares the same base filename as the video (for example, episode.mp4 and episode.srt in the same folder), so matching those names before importing avoids having to manually attach the caption track.

Can I use the SRT file with a podcast host or platform that expects video captions?

If the destination expects a video with a caption track, pair your MP3 with a simple video (a static image or waveform works) in an editor and add the SRT as its caption track; if the destination accepts a standalone .srt file alongside a video upload, you can upload the two files separately without combining them yourself.

Does generating an SRT file from my MP3 satisfy legal accessibility requirements?

Transkio produces an accurate, reviewable caption file, which is a genuinely useful accessibility step, but it doesn't carry any specific accessibility certification. Whether your content meets a particular legal or regulatory accessibility standard is something to verify against your own requirements, not something any transcription tool can certify on your behalf.

Does the MP3's bitrate or audio quality affect how accurate the captions are?

Not directly. Caption accuracy tracks how clearly the speech was recorded — microphone distance, background noise, how many people are talking — not the file's bitrate or compression level. A modest-bitrate recording made in a quiet room typically produces cleaner captions than a high-bitrate file recorded in a noisy space.

Can I generate an SRT from a recording that isn't in English?

Yes — 50+ languages, auto-detected or selectable before you upload. The exported SRT file uses the same language as the transcript.

I publish on a weekly schedule — does this workflow scale to that?

Yes. The upload, review, and export sequence takes a few minutes regardless of how many times you repeat it, so it holds up fine as a recurring step in a weekly or even daily publishing routine, rather than something that gets more tedious the more episodes you produce.

Can I download the audio track back out after generating an SRT, if I only have the transcript now?

The original MP3 file stays stored with your transcript in your account for as long as you keep it there, so you can always go back to the source audio — you're not left with only the text after transcription completes.

Do I need to record my voiceover in one continuous take for this to work well?

No. Recording in multiple takes and stitching them together in an audio editor beforehand works fine — Transkio transcribes and times whatever the final MP3 actually contains, regardless of how many separate takes it was assembled from.

What happens to filler words like "um" and "uh" in the exported captions?

The transcript captures speech as spoken, including filler words, but since it's fully editable before export, it's easy to trim obvious filler out of the caption text during review if you'd rather the captions read more cleanly than the raw audio sounds.

Is there a way to preview how the captions will look before exporting?

The transcript editor shows the paced, cue-by-cue text before you export, so you can see how lines are split and how long each one stays on screen. For a full visual preview against actual video footage, loading the exported SRT into a media player like VLC alongside your video is the fastest check.

Can I re-export an SRT after making further edits later?

Yes. The transcript stays stored and editable in your account, so you can go back, adjust wording or speaker labels, and re-export a fresh SRT file at any time without re-uploading or re-transcribing the original audio.

Does Transkio work for voiceover recorded on a phone rather than professional gear?

Yes — a script read into a phone's voice memo app is one of the most common sources of MP3-to-SRT conversions on this page. Recording quality still matters (a quiet room and a consistent distance from the mic help), but no special equipment is required.

Turn Your MP3 Voiceover Into Captions

Upload the audio and export a properly timed SRT file in minutes — free plan included, no card required.

Transcribe for free
  • 30 free minutes, no card required
  • Transcripts in minutes, not hours
  • 50+ languages