GrowMyTube
🎬 Video tool

Auto captions for YouTube videos

Turn speech into captions with Whisper AI in your browser. Download an SRT or VTT file for YouTube, or pick an animated word-by-word style and burn it into your video.

Add a video or audio file and get timed captions in minutes. The Whisper speech model runs on your own device, so nothing is uploaded. Download an SRT or VTT to upload in YouTube Studio, or burn one of 24 animated styles, including one or two words at a time in the centre, into the video itself.

How to use

  1. Add your video or audio. Drop an MP4, MOV, WebM, MP3 or WAV file into the tool. Files up to about 10 minutes work best.
  2. Pick a language and a speed. Leave the language on auto-detect or choose yours. The page shows an estimated wait for each model size, measured on your device after the first run.
  3. Fix any words. Read through the generated lines and correct names, brands and jargon. If you only change spelling, the word timing stays exact.
  4. Export or burn in. Download an SRT or VTT to upload in YouTube Studio, or choose a caption style and create a new MP4 with the captions drawn onto it.

Features

Nothing is uploaded

The speech model runs on your device. Your video, which may be unpublished, never reaches a server.

SRT, VTT and TXT export

Upload the SRT or VTT in YouTube Studio as a subtitle track, or keep the plain text as a transcript.

24 animated caption styles

Word-by-word captions with highlights, boxes, glow and gradients, with your own colours, fonts and animation.

Honest wait times

The page measures your device once and shows how long each model size will take for your clip, with a faster option when a wait would be long.

Clean MP4 for the burned-in version

The captioned video is rebuilt as a standard H.264 and AAC MP4 with constant frame rate, so it plays and uploads without glitches.

What people use it for

Captions on YouTube: the two routes

There are two different things people mean by captions. A subtitle file such as an SRT is a separate text track that viewers can turn on or off, and YouTube indexes its words. Burned-in captions are drawn onto the picture, so they always show and cannot be switched off, which is the look most Shorts use.

This tool does both from the same transcript. For a long video, upload the SRT. For a Short, burn in an animated style. You can do both for the same video and keep the file for later.

Getting accurate captions

Clear speech with little background noise gives the best result. Music under the voice, several people talking at once and strong accents lower accuracy, so always read the lines before you publish. If the page offers a larger model for your clip, it is slower but usually more accurate.

Music and sound effects can make speech models invent sound descriptions like [Music]. The tool removes those by default and collapses runaway repeats, and you can switch that off if you want sound descriptions kept.

Adding the SRT in YouTube Studio

Open YouTube Studio, choose Subtitles, pick the video and a language, then choose Add and upload the file with timing. YouTube accepts SRT and VTT among other formats. After uploading, preview a few lines against the audio before publishing.

How the speech model runs in your browser

The tool uses an open speech model called Whisper, compiled so it can run inside a web page. On a computer with WebGPU it uses the graphics chip, otherwise it falls back to the processor. The audio is read from your file, cut into pieces of about a minute at quiet moments for longer clips, and turned into words with timestamps. None of that leaves your device.

The first time you use a model size, your browser downloads it and keeps it. After that it starts quickly. The page measures how fast your device is on the first clip of a minute or longer, and from then on it shows an estimated wait for each model size, so you can choose between a faster, less accurate run and a slower, more accurate one.

Choosing a model size

Tiny is the quickest and smallest. It is fine for clear speech in a quiet room, and it makes more mistakes with accents, jargon and background noise. Base is a good default. Small is the most accurate of the three and the slowest, and it needs a reasonably capable device and a large one-time download.

If a model produces nothing but punctuation on your device, which can happen with some graphics chips, the tool detects it, tries another setting and tells you. You do not need to do anything.

Making captions that read well

For subtitle files, shorter lines are easier to read, especially on a phone. The line length option breaks captions at about 28 or 42 characters. For burned-in captions on Shorts, one or two words at a time in the centre of the frame is a popular style, and the position control lets you move captions away from the area covered by YouTube's own buttons.

Always read the transcript once before you publish. Names, numbers, brand names and technical terms are the words speech models most often get wrong, and they are the words viewers notice most.

Troubleshooting

The captions are mostly wrong

Check that the language is set correctly instead of auto-detect, try a larger model size, and make sure the speech is not buried under music. A short, clear clip is a good test.

It stops or the tab crashes on a long video

Browsers limit memory per tab. Keep clips to about ten minutes, close other tabs, and use the Tiny or Base model. You can split a long video and caption the parts.

Words appear slightly early or late in the animated captions

Word timing comes from the model and can drift on fast speech or music. The page tells you when it had to re-estimate timing. Fix any line that looks wrong in the editor, and check the preview before burning in.

The burned-in video has no audio or stutters

Keep the tab open and visible while it records, because browsers slow down hidden tabs. The result is rebuilt as a standard MP4 afterwards, which fixes most playback glitches.

Limits and good to know

FAQ

Is the caption generator free?

Yes. There is no sign-up, no watermark and no daily limit.

Is my video uploaded?

No. The speech model runs in your browser on your own device.

Which file does YouTube want?

Download the SRT or VTT and upload it under Subtitles in YouTube Studio. Both are accepted.

Can I get captions that show one or two words at a time?

Yes. Choose an animated style and set words at a time to one, two or three, then burn the captions into the video.

How accurate are the captions?

Good on clear speech and weaker with music or crosstalk. Always read and fix the lines before publishing.

Why does the first run take longer?

The speech model downloads once. After that it starts quickly, and the page shows how long the next run should take.

Do captions help a video get views?

Captions make a video usable for people who watch without sound or who are deaf or hard of hearing, and they give YouTube the text of what is said. They are not a guaranteed ranking boost, but they do help viewers stay.

Can I translate the captions?

This tool transcribes in the language spoken. YouTube Studio can add subtitle tracks in other languages once you upload a file, and you can give the SRT to a translator.

What is the difference between an SRT and burned-in captions?

An SRT is a separate track viewers can turn on or off, and YouTube reads its text. Burned-in captions are part of the picture, always visible, and cannot be switched off.

More auto captions pages