Skip to content

Captions and voice-over

Studio can turn speech in your video into captions, and turn text into a voice-over.

Transcribe listens to a sequence and turns speech into timed captions on the timeline. The result is ordinary caption text that you can edit.

Transcription can run:

  • On your own hardware, using speech models that run locally.
  • On Cloud GPU, after you confirm an estimate. Only the sound is sent, never the picture. It handles up to 4 hours per file, and English, German, Spanish, French, Italian, Dutch, Portuguese and Polish. Captions come back as VTT or SRT, with at most two lines and 42 characters per line by default.

Speaker labels aren’t available yet.

Depending on the format you export, captions can be burned into the picture, saved as a separate file next to the video, or included as a track viewers can turn on and off.

Type a script and Studio reads it aloud, placing the result on the voice track.

  • On your own hardware, several voices are included.
  • On Cloud GPU, there are eight English voices, a speed from 0.5× to 2×, and scripts up to 100,000 characters. The script is deleted with the job.

You can also record your own voice-over with a microphone. See Audio and music.