Put spoken words on screen.
Automatic captions give your audience a written version of what they hear. Tap Captions's caption workflow starts with speech in a local video, processes it on your iPhone and gives you text to review.
Start with the speech in your video
Use a clip with a clear voice and as little competing audio as possible. Recorded speech is transcribed into words. This is speech recognition, which is different from the speech synthesis used to make a voiceover.
Captions are useful for viewers watching with the sound off and for people who need a text alternative to speech. Automatic output is a first pass. Names, accents, specialist terms and overlapping speakers can need correction.
Review before you share
Read the captions while listening to the original audio. Correct names, numbers and punctuation. Check the start and end of each phrase, then play the video again without sound to see whether it still makes sense.
When a sound matters to the meaning, include it in the text. For example, an off-camera doorbell may explain why a speaker turns away. Speech transcription alone does not guarantee a complete accessibility caption track.
Keep the text readable
Use short phrases and a size that remains legible on a phone. Keep important words away from faces, demonstrations and the controls a social app places over the video. Read more about caption styling.
The export places captions in the video image so they remain visible when the video is shared. These are open captions. They are different from a separate caption track that a viewer can turn on or off.
Continue reading
Use the iPhone subtitle guide or learn the difference between captions and subtitles.