About audio descriptions and captions
People who haven’t worked in accessibility before often use “captions” and “audio description” as if they’re two names for the same thing — subtitles, more or less. They aren’t. They solve two different problems for two different audiences, and a video with only one of them is accessible to only one of those audiences.
What captions do
Section titled “What captions do”Captions are a text track: dialogue, and the significant sounds that carry meaning, timed to appear as they’re heard. Their primary audience is Deaf and hard-of-hearing viewers, for whom captions are the only way to access what’s being said and what’s happening sonically — a door slamming, tense music building, a phone ringing off-screen. But captions serve a much wider group in practice than that primary audience alone: anyone watching with the sound off in a public space or open office, anyone whose second language they read more confidently than they follow by ear, anyone in a noisy environment. That breadth is part of why captions have become an expected default on video content generally, not a niche accommodation.
What audio description does
Section titled “What audio description does”Audio description is a narration track, inserted into the natural pauses in a video’s existing dialogue, that describes what’s visually happening — actions, settings, on-screen text, who’s on screen and what they’re doing, visual jokes or reveals that a script wouldn’t otherwise mention out loud. Its audience is blind and low-vision viewers, who already have full access to the dialogue and sound design through the normal audio track, but have no way to perceive what’s only conveyed visually. Where captions solve “I can’t hear this,” audio description solves “I can’t see this.”
Because it has to fit into gaps between existing dialogue, audio description is a genuinely different writing discipline from captioning. A caption can be as long as the line of dialogue it transcribes. A description has to say something useful about the screen in whatever silence is actually available, which means being selective about what matters most — the detail a sighted viewer would register instantly (who just walked in, what changed in the room) rather than an exhaustive account of the frame.
Why a fully accessible video needs both
Section titled “Why a fully accessible video needs both”Because captions and audio description serve different senses, providing only one leaves out the audience the other was for. A video with excellent captions and no audio description is well-served for Deaf and hard-of-hearing viewers but tells a blind viewer nothing about what’s on screen beyond the dialogue they can already hear. A video with audio description and no captions does the reverse. Treating one as a substitute for the other is the most common mistake in this space — usually not out of neglect, but because both get filed under the same mental bucket of “accessibility for video” even though they answer to entirely different needs. Genuinely accessible video treats them as two separate, necessary layers, not alternate versions of the same fix.
Where this connects
Section titled “Where this connects”Content Studio generates both tracks for a video project, independently of each other, so you can produce captions, audio description, or both depending on what the project needs. For the hands-on steps, see Generate and edit audio descriptions and Turn on and refine captions, or walk through the full first pass in Describe your first video. Once you’re ready to hand a video off, Export a video covers getting both tracks out in a usable format — see Export formats for what each format actually contains. For the assistive technology on the receiving end of these tracks, see About assistive technology.