Skip to content
Transkio

Video Subtitle Generator & Caption Generator

Upload a video, get back accurate, properly paced SRT and VTT subtitle files — free tier included, no card required.

Upload audio or video

MP3, WAV, M4A, AAC, OGG, OPUS and more · up to 100 MB on the free plan

or drag and drop it here

Record in your browser

Meetings and calls up to 30 minutes free — longer on paid plans

Start recording

Free account · 60 trial minutes, then 30 minutes a month · no credit card required

Most video today is watched with the sound off — scrolling a feed on a train, sitting in an open office, or auto-playing muted while someone decides whether it's worth tapping for sound. A subtitle or caption file is what keeps a video understandable in every one of those situations, and it's also what a deaf or hard-of-hearing viewer, a non-native speaker following along, or a search engine trying to index what the video is actually about all rely on. Transkio turns any video file into a downloadable subtitle track without you typing a single timecode by hand.

Upload your video and Transkio transcribes its audio track, then formats that transcript into properly paced subtitle cues — split at natural word boundaries, kept within readable character and duration limits per cue — and exports it as an SRT or VTT file ready to hand to your video editor, LMS, or upload directly to YouTube. You can review and correct the transcript first, so what ends up as captions is exactly right.

"Subtitles," "captions," and "burned-in captions" get used interchangeably in everyday conversation, but they're not quite the same thing, and the difference matters once you're deciding what to export. A subtitle or caption file is a separate, timestamped text file that a video player reads alongside the footage; burned-in captions are pixels drawn permanently onto the video itself. Transkio produces the first kind — SRT and VTT files — because a separate text file stays flexible: it can be re-styled, swapped, turned off, or handed to an editor to burn in later, while a caption that's already burned in can't be undone once it's rendered.

The people uploading video here cover a wide range: a marketing team publishing short clips to feeds where sound defaults to off, a course creator meeting an accessibility requirement for an online class, a corporate learning-and-development team captioning training for a hearing-impaired employee, a conference organizer captioning a recorded keynote, a YouTuber who wants a searchable transcript alongside the video. The underlying job is identical every time: turn spoken audio into a subtitle file that's accurate, properly timed, and ready to attach wherever the video lives.

Manually Timing Captions Eats Hours You Don't Have

Writing subtitles by hand means transcribing the video, then guessing at timestamps, then nudging each cue back and forth in an editor until the words line up with the audio. For a 10-minute video that's easily an hour of tedious work, and it's the first thing that gets skipped when a deadline is tight — which is how videos end up published with no captions at all, locking out viewers who are deaf, hard of hearing, or just watching muted on the train.

The stakes are higher than a missed afternoon. Schools, many workplaces, and public-facing organizations are commonly expected to provide captioned video under accessibility frameworks like the ADA and WCAG, and a video with no captions can simply be unreadable to someone who can't hear it at all — there's no muted-audio fallback for a viewer who is deaf or hard of hearing to begin with.

There's also a platform-visibility cost to skipping captions. Search engines and recommendation systems can't watch a video, but they can read a caption file, and a meaningful share of what determines whether a video gets surfaced for search or recommended in-feed comes down to whatever text is attached to it. Publish without a caption track and you're relying entirely on a title and thumbnail to do work that text could otherwise help with.

And hand-timed captions are fragile in a way people don't expect until it happens: re-cut the video by even a second or trim the intro, and every cue drifts out of sync, meaning the whole timing pass often has to be redone rather than simply shifted.

Upload Once, Export a Synced Subtitle File

Transkio listens to your video's audio track, generates a full transcript, and automatically paces it into subtitle cues — timed to the audio and split so no cue is a wall of text or a flash of one word. Fix any misheard names or terms in the editable transcript, then export as SRT or VTT and drop the file straight into your editing software, video host, or course platform.

Timing isn't eyeballed — each cue is generated from the actual word-level timestamps produced during transcription, so a cue's start and end time reflect when those words were genuinely spoken, not a rough guess nudged into place by hand afterward.

Because the subtitle file and the underlying transcript are the same data viewed two ways, correcting a misheard word in the transcript editor updates the exported caption file too — there's no separate subtitle-only editing pass to redo after you've already fixed the transcript once.

If your source is audio without a picture track, the same pacing and export logic is available through audio to text or, for a file already saved as MP3, MP3 to SRT — just without a video frame attached to it.

Subtitles, Captions, and Burned-In Captions: What's Actually Different

In everyday conversation, people use "subtitles" and "captions" as if they're the same word, and for most purposes they are — both describe a timed text track synced to a video's dialogue. Historically, "captions" (particularly closed captions) implied a same-language track that also describes relevant non-speech sound, like [phone rings] or [applause], written for viewers who can't hear the audio at all, while "subtitles" implied dialogue-only text, often for viewers who can hear the audio but don't understand the spoken language. In practice today, especially in English-language software and platforms, the two words are used interchangeably, and this page treats them the same way.

"Burned-in" or "open" captions are a different thing entirely: instead of a separate file, the text is rendered directly into the video's pixels, permanently, by the video-editing software. Once burned in, captions can't be turned off, resized, restyled, or swapped for another language — they're just part of the picture from that point forward, the same as any other visual element in the frame.

SRT and VTT files, by contrast, stay separate from the video. A media player, a website, or a course platform reads the subtitle file alongside the video and renders the text on top at playback time, which means the same video file can carry multiple caption tracks, a viewer can toggle captions on or off, and updating a caption is as simple as swapping the text file — no re-export of the video itself required.

SRT vs. VTT: Which Format Should You Actually Use

SRT (SubRip Text) is the older and more broadly compatible of the two formats: plain numbered blocks of text, each with a start and end timecode, readable by essentially every video editor, media player, and captioning tool in wide use. If you're not sure what a destination platform or editing tool expects, SRT is the safer default — it's supported almost everywhere.

WebVTT (.vtt) is the web-native standard, designed to work with the HTML5 <track> element that browsers use to display captions on embedded video. It supports a few extra features SRT doesn't — basic text positioning and simple styling cues — which matters if you're embedding a captioned video directly on your own website rather than uploading to a third-party platform.

Transkio exports both formats from the same transcript, so there's no need to pick one and commit — download SRT for your editor or a platform that expects it, and VTT for a web player, without transcribing twice or converting between formats yourself.

Subtitle Timing and Readability: The Rules That Actually Matter

A technically accurate caption can still be unreadable if it's paced badly. Captioning practice that's built up across broadcast and streaming over the years converges on a few widely used guidelines: keep a line to roughly 32–42 characters so it fits comfortably across the bottom of a frame without wrapping awkwardly, cap a cue at two lines so it doesn't cover too much of the picture, and pace the text at a reading speed a viewer can actually keep up with — often discussed as somewhere in the range of 15–17 characters per second, which works out to roughly two to three words per second for average sentence lengths.

Cue duration matters too: a cue that flashes for under a second reads as a blur even if the text is short, while a cue that lingers for eight or ten seconds encourages a viewer to re-read it and lose track of the video. A comfortable range for most spoken-word content sits between roughly one and seven seconds per cue, with the exact number depending on how much text that cue carries.

Where a cue breaks matters as much as how long it lasts. Splitting mid-word or mid-phrase — cutting "the launch is moving to" from "the fourteenth" — forces a viewer's eye to do extra work reconnecting the sentence across two cues. Breaking at a natural clause boundary, a comma, or the end of a sentence keeps each cue readable as a unit on its own.

Transkio applies these general norms automatically when it paces cues from the word-level transcript, so a generated subtitle file already reads at a reasonable speed with sensible line breaks — you're reviewing and correcting wording, not re-timing the whole file by hand.

Why Silent Autoplay Made Captions the Default, Not an Extra

Social and short-form video feeds default to muted autoplay, and a viewer decides within the first couple of seconds whether a clip is worth unmuting — if the video can't communicate anything without sound, most of that scrolling audience never engages with it at all. A captioned clip can land its point purely through text on screen, which matters enormously in an environment where sound is the exception, not the assumption.

The same holds true off social media: open-plan offices, shared living spaces, public transit, and waiting rooms are all places people routinely watch video without headphones and without sound. A training video, a product walkthrough, or a recorded webinar that only works with audio on is a video that a meaningful share of its intended audience simply can't consume in the moment they encounter it.

None of this requires guessing at exact numbers to take seriously — it's a description of how people actually watch video day to day, and it's the reason captioning has moved from an accessibility nice-to-have to something most video creators treat as a standard publishing step, the same way they'd never skip a thumbnail or a title.

Accessibility: What ADA and WCAG Actually Ask For

In the United States, the Americans with Disabilities Act (ADA) has been interpreted by courts and enforcement guidance to extend accessibility obligations to digital content, including video, for many businesses and public-facing organizations. Internationally, the Web Content Accessibility Guidelines (WCAG) set out specific success criteria for video, and captioning prerecorded audio content is one of the foundational requirements — it's treated as a baseline, not an advanced feature.

It's worth being precise about what a subtitle file does and doesn't do here: exporting an SRT or VTT file and attaching it to a video is the mechanical step that makes captioning possible, and it addresses the caption-specific requirement directly. Whether a given video, site, or organization is fully compliant with ADA or WCAG as a whole is a broader legal and technical question that depends on far more than one caption file — color contrast, keyboard navigation, alt text, and other criteria all factor in, and that determination is something your organization (ideally with legal or accessibility-specialist input) makes for itself. Transkio isn't a compliance certification service and doesn't audit or certify a site's accessibility status.

This matters most concretely for schools and universities publishing lecture or course video, government agencies and contractors, and any consumer-facing company whose video content reaches a general public audience — these are exactly the contexts where captioning is most often an explicit, documented expectation rather than a courtesy.

Captions for Non-Native-Speaking Audiences

A viewer who's fluent enough to follow spoken content but not a native speaker often benefits enormously from reading along in the same language — captions fill in the gaps that come from an unfamiliar accent, fast speech, background noise, or jargon that's easy to mishear but easy to read. This is a distinct use case from translation: the caption is in the same language the video was recorded in, just rendered as text alongside the audio.

If your audience genuinely doesn't read the language the video was recorded in, that's a translation need rather than a captioning one — Transkio's translation feature (available from the Pro plan) produces a translated text version of the transcript, though as covered in the FAQs below, that's a separate translated document rather than a translated subtitle file you can drop straight onto the video today.

Even without translation, publishing accurate, same-language captions widens who a video actually works for — a global product demo, an international training session, or a conference talk with attendees from many linguistic backgrounds all become more usable the moment reliable captions exist.

Subtitles and Video SEO

Search engines can crawl text; they can't watch a video and understand what's being said in it. A transcript or caption file attached to a video page gives a search engine something concrete to index — the actual words spoken, the topics covered, the terminology used — rather than relying entirely on a title, description, and whatever metadata was manually entered.

Platforms like YouTube read uploaded caption files directly, and that text can factor into search relevance, auto-generated chapters, and how accurately the platform can summarize or match a video against a search query. A video with no caption track is, from a search engine's point of view, mostly just a thumbnail and a title.

For a video embedded on your own website — a product page, a course landing page, a blog post — the same logic applies to the page itself: a visible or embedded transcript alongside the video gives that page substantially more indexable, keyword-relevant content than the video alone ever could.

The Full Workflow: From Upload to Exported Subtitle File

Start by uploading the video file itself — MP4, MOV, WebM, MKV, or MPEG. Transkio extracts the audio track internally and runs it through transcription, which typically finishes in a few minutes for anything under an hour of runtime, regardless of how the video was originally shot or edited.

The result opens in the transcript editor before any subtitle file is generated. This step matters: any word the transcription model misheard will otherwise carry straight through into the exported captions, so a quick read-through — fixing a misheard name, an acronym, or an unusual term — is the difference between a captioned video that's accurate and one that quietly embarrasses you the moment someone reads along.

Once the transcript reads the way you want, export SRT or VTT (or both). The file downloads ready to attach: drop it into your video editor's caption track, upload it alongside the video file on YouTube or your LMS, or reference it from an HTML5 <track> element if you're embedding video on your own site.

Where Your Exported File Actually Goes

YouTube accepts SRT and VTT files uploaded directly alongside a video through its own captions manager, and treats an uploaded caption track as more reliable than its own automatic captions for both accuracy and search indexing. Most learning-management systems (Panopto, Kaltura, and similar platforms commonly used for course video) accept the same file types for the same reason — accessibility and searchability inside the course library.

Editing software — Premiere Pro, Final Cut Pro, DaVinci Resolve, CapCut, and most others — imports SRT natively as a caption or subtitle track you can review on the timeline, restyle to match your brand, and burn in at export time if that's genuinely what the destination platform requires (some ad placements and certain social formats only accept burned-in text, with no separate caption-track option).

For a video embedded with a plain HTML5 <video> tag on your own site, a VTT file referenced through a <track> element gives visitors a native captions toggle in the player controls, without any third-party captioning script or plugin needed.

Common Captioning Mistakes to Avoid

The most common mistake is exporting straight from the raw transcription without reading it first — a misheard name or technical term in the transcript becomes a permanent error in the caption file, and once that file is attached to a published video, most viewers will assume the mistake is real rather than a transcription slip.

A second common mistake is over-packing cues: trying to fit an entire sentence onto one line to "save space" produces a wall of text a viewer can't actually read at normal speaking pace. Trust the pacing rather than fighting it — a shorter cue that changes more often is almost always more readable than a longer one that lingers.

A third is burning captions in permanently before you're sure you're done editing the video. Once text is baked into the pixels, fixing a typo means re-exporting the whole video, not just a text file — keep working from the separate SRT or VTT file for as long as possible, and only burn in captions as the very last, deliberate step if the destination truly requires it.

Short-Form Clips vs. Long-Form Video: Different Captioning Priorities

A 30-second social clip and a 90-minute recorded webinar both need captions, but the priorities shift with length. A short clip lives or dies on its first few seconds of muted autoplay, so every cue needs to be immediately legible — short lines, generous timing, nothing that requires a second read to parse before the viewer scrolls past.

A long-form recording — a lecture, a webinar, a multi-speaker panel — puts more weight on consistency over the full runtime and on the transcript underneath the captions being genuinely accurate, since a viewer is far more likely to search or skim a long video's transcript afterward than to read every caption cue in real time. Reviewing the transcript before export matters more, proportionally, the longer the recording runs, simply because there are more opportunities for a misheard name or term to slip through.

Either way, the underlying workflow doesn't change — upload, review the transcript, export — but it's worth spending review time in proportion to how much the video's length raises the stakes of an uncaught error.

Captioning Recorded Meetings, Webinars, and Screen Recordings

Recorded meetings and webinars are a slightly different captioning case from produced video: the audio is often less controlled — multiple people on different microphones, occasional cross-talk, a shared screen with no separate narration — and the captions matter less for a public audience and more for people who missed the live session or need to search back through it later.

Transkio's speaker detection (available from the Elite plan) is particularly useful here: labeling who said what in the transcript makes the review pass faster on a multi-person recording, even though the exported caption cues themselves follow the standard convention of dialogue-only text. Combined with searchable timestamps, a captioned meeting recording becomes something a teammate can skim for the two minutes that matter instead of watching from the start.

The same workflow applies whether the recording came from Zoom, Google Meet, a screen-capture tool, or a browser recording made directly in Transkio — upload the resulting video file and the rest of the process is identical.

How It Works, Step by Step

  1. 1

    Upload Your Video

    Drag in MP4, MOV, WebM, MKV, or MPEG — up to 100 MB free. Transkio pulls the audio track automatically.

  2. 2

    Review the Transcript

    Click any line to jump the video player to that moment, and fix any misheard names or terms before exporting.

  3. 3

    Export SRT or VTT

    Download a properly paced, timed subtitle file, split at natural word boundaries and readable character limits.

  4. 4

    Attach It Anywhere

    Upload alongside the video on YouTube or your LMS, import into your editor, or reference it from your site's video player.

What You Get

Free SRT and VTT Export

Caption export isn't locked behind a paid tier — every Free account can export SRT and VTT files, up to 60 trial minutes then 30 minutes a month.

Properly Paced Cues, Not a Wall of Text

Subtitles are split at natural word boundaries and kept within readable character and duration limits per cue, so viewers aren't stuck reading a paragraph in two seconds.

Edit the Transcript Before You Export

Click any line to jump the video player to that moment, fix a misspelled name or technical term, and rename speakers — your corrections carry straight through to the caption file.

Works From a Direct Upload, No URL Fetching

Upload the video file itself — MP4, MOV, WebM, MKV, or MPEG — and Transkio pulls the audio track for you. There's no link-paste shortcut, so keep your source file handy.

50+ languages

Transcribe lecture recordings, product demos, and interviews in dozens of source languages and get accurate captions back in that same language.

Same Transcript Powers Captions and a Readable Script

The transcript behind your subtitle file is also a clean, readable document on its own — useful for show notes, a blog recap, or an accessible text version of the video.

Both SRT and VTT, No Extra Conversion

Export whichever format your platform or editor expects — or both — from the same transcript, without running a separate file-conversion tool afterward.

Timecodes From Real Word-Level Timestamps

Cue timing comes from when words were actually spoken during transcription, not a rough manual estimate nudged into place after the fact.

Speaker Labels Carry Into the Transcript

Multi-speaker videos get each voice labeled in the underlying transcript on Elite and up, making it easier to keep track of who said what while reviewing before export.

See it work

See the Difference

A real example of how the same recording reads before and after Transkio.

Lecture captioning example

Before — raw auto-captions

so today were gonna talk about uh cellular respiration and uh specifically how mitochondria uh produce energy for the cell

After — Transkio transcript

1 00:00:00,000 --> 00:00:03,500 So today we're going to talk about cellular respiration — 2 00:00:03,500 --> 00:00:07,000 specifically how mitochondria produce energy for the cell.

Frequently asked questions

Is a caption generator the same thing as a subtitle generator?

Yes — on Transkio, captions and subtitles come from the same transcript and the same export step. Whichever word you use, you'll end up with a timed SRT or VTT file synced to your video's audio.

Does adding captions actually help with accessibility compliance?

Captions are the core requirement in most accessibility standards for video, including ADA and WCAG, since they give deaf and hard-of-hearing viewers access to spoken content. Transkio exports a standard SRT or VTT file you can attach wherever your video is hosted, but full legal compliance depends on more than captions alone — Transkio doesn't certify or audit accessibility compliance.

Is SRT and VTT export available on the free plan?

Yes. Every Free account can export TXT, SRT, VTT files, with 60 trial minutes to start and 30 minutes included every month after — no card required.

Can I get translated subtitles in another language?

Not yet. Transkio's translation feature (available on Pro and above) produces a translated text transcript, but SRT/VTT subtitle export is currently original-language only — the captions match the language spoken in your video.

What video formats can I upload?

MP4, MOV, WebM, MKV, and MPEG are all supported. Upload the file directly — Transkio pulls the audio track from it automatically, so there's no need to extract or convert audio yourself first.

What's the actual difference between SRT and VTT?

Both are timed text files that pair captions with timecodes, and most editors and platforms accept either. SRT is the older, more universally supported format; VTT (WebVTT) is the web standard built for HTML5 video players and supports a few extra styling and positioning options. When in doubt, SRT is the safer default for editors and third-party platforms, and VTT is the right choice for embedding captions on your own website.

Is there a recommended reading speed or pacing for subtitles?

Widely used captioning guidelines suggest keeping pace to roughly 15–17 characters per second — about two to three words per second for typical sentences — so a viewer can read a cue in the time it's on screen. Transkio's automatic pacing follows these general norms when it generates cues from your transcript's word-level timestamps.

How many characters should be on one subtitle line?

A common guideline is roughly 32–42 characters per line, with a cue capped at two lines, so the text fits comfortably across the frame without wrapping awkwardly or covering too much of the picture. Longer sentences get split across multiple cues rather than crammed onto one.

Can I burn captions permanently into my video instead of using a separate file?

Transkio exports separate SRT and VTT files rather than a burned-in video, which keeps your options open — you can restyle, translate, or turn off a separate file, none of which is possible once text is rendered into the video pixels. If a specific platform requires burned-in captions, most video editors can import the SRT file and burn it in at export time as a final step.

Do captions need to be word-for-word, or can they be condensed?

Transkio produces word-for-word captions based on what was actually said, since that's the safest default for accuracy, accessibility, and search indexing. You can edit the transcript before exporting if you want to lightly condense a cue for readability — corrections you make there carry straight through to the exported file.

Will an exported SRT or VTT file work on YouTube?

Yes. YouTube accepts SRT and VTT files uploaded directly through its captions manager alongside your video, and an uploaded caption track is generally more accurate and more useful for search than YouTube's own automatic captions.

Can I import the SRT file into Premiere Pro, Final Cut, or CapCut?

Yes, all of these (and most other editing software) import SRT natively as a caption or subtitle track on the timeline, where you can restyle it, adjust individual cues, or burn it into the final export if needed.

Can I add speaker labels to a video's captions for a multi-person video?

Speaker detection labels each person's turns in the underlying transcript on Elite and up, which makes reviewing a multi-speaker video much easier before export. Exported SRT/VTT cues currently follow the standard captioning convention of dialogue text without an inline speaker tag on each line.

Does background music or sound effects get captioned too?

Transkio transcribes spoken words, not ambient music or non-speech sound effects, so an exported caption file covers dialogue and narration rather than acting as a full sound-description track. For content that needs sound-effect descriptions as part of accessibility captioning, those would need to be added manually to the exported file.

Is there a length limit on videos I can generate subtitles for?

Video length is bounded by your plan's file size and monthly minutes rather than a fixed cap — free accounts can upload files up to 100 MB, with higher limits on paid plans, which in practice covers everything from a short clip to a multi-hour recording.

Should a short social clip and a long webinar be captioned differently?

The export workflow is identical either way, but review priorities shift: a short clip depends on every cue being instantly readable during muted autoplay, while a long recording benefits more from a careful transcript review, since there's more audio for a misheard name or term to slip through undetected.

Can I caption a recorded Zoom meeting or webinar the same way?

Yes. Upload the recorded meeting or webinar video the same way you would any other file, and Transkio transcribes and paces it into subtitle cues. Speaker detection (from the Elite plan) labels who's talking in the transcript, which makes reviewing a multi-person recording noticeably faster before you export.

Do I need special software to view an SRT or VTT file before uploading it anywhere?

No — both are plain text files you can open in any text editor to check them. You don't need dedicated subtitle software to review, edit, or verify the file; a text editor and the video itself are enough to confirm the captions look right.

Why does Transkio export separate SRT/VTT files instead of burning captions into the video directly?

A separate file keeps your options open in ways a burned-in caption can't: you can restyle it, swap it for another language later, hand it to an editor to re-time against a re-cut version, or simply turn captions off for a viewer who doesn't want them. Burning text permanently into the video pixels is a one-way step best left until you're certain no further changes are coming.

Can I re-export captions if I trim or re-cut my video afterward?

Yes — since the subtitle file is generated from the transcript's word-level timestamps rather than hand-placed, the safest approach after a re-cut is to re-upload the updated video and export fresh cues from it, rather than trying to manually shift an old SRT or VTT file to match new edit points.

Get Captions on Your Next Video in Minutes

Upload a video, review the transcript, and export a properly paced SRT or VTT file — free tier included, no card required.

Transcribe for free
  • 30 free minutes, no card required
  • Transcripts in minutes, not hours
  • 50+ languages