English
English subtitle generator with word-level timing.
Automatic English subtitles, timed to every individual word rather than to the line — which is what makes karaoke highlights and word-by-word reveals land on the beat.
An English subtitle generator transcribes spoken English and writes timed subtitles from it. LumaCaption times every word individually rather than every line, which is what allows per-word emphasis and karaoke reveals, and exports either burned-in video or an SRT, VTT or TXT file.
Most subtitle generators time a line: they decide that a phrase occupies the stretch between 4.2 and 6.1 seconds and put it on screen for that window. That is enough for accessibility subtitles and is all a broadcast workflow ever needed.
It is not enough for anything that moves. A karaoke highlight, a word-by-word reveal, an accent that travels through a phrase — all of them need to know when each individual word was said, and a line-level timestamp simply does not contain that information. Tools that time lines and then divide the window evenly between the words produce captions that drift out of sync with the voice within a couple of seconds, and the drift is most visible exactly where the delivery matters most.
This page is about the English path specifically. If you speak Hindi or move between Hindi and English mid-sentence, the Hindi and Hinglish pages cover behaviour that differs enough to be worth reading separately.
Every word timed, not every line
Per-word timestamps come out of transcription rather than being interpolated from a line window. That is the difference between a karaoke highlight that tracks your voice and one that drifts.
Two engines, one for accuracy and one for volume
A premium engine for the clips that matter and a lighter one for everyday work, so you spend the accurate minutes where the audio is hard rather than on every test render.
Burned-in or sidecar, both free
Export a styled MP4 with the captions rendered into the picture, or download SRT, VTT or TXT. Subtitle files are unlimited on every plan including free.
The full catalogue applies
All 294 presets across ten lanes work with English, with per-word colour, placement and emphasis on top. Nothing is gated behind a plan.
What actually breaks English transcription
Accuracy claims in this category are close to meaningless because every tool performs well on the audio they are demonstrated with — one speaker, close mic, quiet room. The differences show up on the recordings people actually make, and they are worth knowing about because most are fixable at the source for less effort than switching tools.
Overlapping speech is the hardest case and no tool solves it well: when two people talk over each other, transcription has to choose, and it will drop one of them. Music under speech is the second, particularly music with vocals, since the model has no principled way to separate a sung word from a spoken one. Room reverb is third and the most underrated — a phone held at arm’s length in a room with hard walls and no soft furnishing smears the consonants that distinguish similar words.
Proper nouns are a category of their own. Names, brands and places are exactly the words a transcript cannot guess from context, and they are also the words that look worst when wrong. Fixing them in the transcript before styling takes seconds; noticing them after export does not.
- Get the mic closer
- The single highest-return change. Halving the distance to the microphone does more for accuracy than any tool choice in this category.
- Duck the music
- If you control the mix, drop the bed under speech. If you do not, expect to correct more words — this is not a tool-selection problem.
- Fix names first
- Scan the transcript for names and brands before you pick a style. They are the errors viewers notice and the ones no automatic pass will catch.
- Check numbers
- Spoken numbers are transcribed inconsistently across every engine — "fifteen hundred" and "1500" are both defensible. Pick one and make it consistent, especially if a number is the emphasised word.
Burned-in subtitles or a subtitle file?
These are different products for different destinations and the choice is not a matter of taste. Burned-in captions are pixels in the video: they appear on every platform, cannot be switched off, and are the only option that works on Reels, Shorts and TikTok, where a sidecar file is either unsupported or ignored.
A subtitle file is a separate track the player renders. It can be toggled off, translated, and read by search engines — which is why it is the right answer for YouTube long-form, for a site embed, and for anything with an accessibility requirement. It is also the wrong answer for short-form, because most viewers will never turn it on.
The pragmatic split most people land on: burn in for short-form, sidecar for long-form, and both for anything that goes to more than one place. Both exports are free here on every plan, so there is nothing to weigh commercially.
How it works
Add your video
Drop in the clip. The video file itself is never uploaded — audio is extracted from it for transcription.
Transcribe and fix names
Every word is timed individually. Scan for proper nouns and numbers, which are the errors worth catching before styling.
Style and export
Pick from 294 presets, then export a burned-in MP4 or download SRT, VTT or TXT. Exports are unlimited and free.
Common questions
How accurate are automatic English subtitles?
On clean single-speaker audio, good enough that most people change only proper nouns. Accuracy falls off sharply with overlapping speakers, music under speech and room reverb — and it falls off for every tool in this category, not just one, so testing your own worst recording tells you more than any published accuracy figure.
What is the difference between word-level and line-level timing?
Line-level timing records when a phrase starts and ends. Word-level timing records when each individual word was spoken. Only the second can drive karaoke highlights, word-by-word reveals or a travelling emphasis, because those need to know where the voice is right now — not just which phrase is on screen.
Can I download an SRT file?
Yes. SRT, VTT and TXT downloads are unlimited on every plan including free. There is no SRT import, though — the tool produces subtitle files, it does not take one in. If you already have an SRT and want it burned into a video, ffmpeg, HandBrake or Subtitle Edit will do that.
Do I need to upload my video?
The video file is never uploaded. Audio extracted from it is sent for transcription, and rendering happens on your own device — which is also why exports are unlimited and free rather than metered.
Will subtitles help my video get more views?
Indirectly. Most short-form video is watched without sound, so a clip with no on-screen text loses viewers in the first seconds. Subtitles do not change how a platform ranks a video by themselves; what they change is how long people watch, and that is what the ranking responds to.
