Readability
Caption best practices: how to make captions people can actually read
The working rules for captions that are easy to read on a phone, with the figures professional style guides actually use.
Updated 9 min readBy the AutoCap team
A caption only works if someone can read it in the time it’s on screen and still see what’s happening in the picture. Broadcasters and streaming services have turned that into precise rules. Social video changes some of the conditions — the screen is tall and small, the styling is louder, the viewer is scrolling — but those rules are still the right starting point, and most of them carry over directly.
The guidelines worth knowing
Four references come up again and again. The figures in this guide come from their current versions. Where they disagree, it’s usually because they were written for different screens and audiences.
- The BBC Subtitle Guidelines (version 1.2.5, March 2026) cover broadcast and online video and include specific guidance for vertical 9:16 video.
- Netflix’s English (USA) Timed Text Style Guide sets the rules for subtitles and SDH delivered to Netflix.
- The DCMP Captioning Key, from the US Described and Captioned Media Program, is written with educational media in mind.
- WCAG 2.2 requires captions for prerecorded video (success criterion 1.2.2) but doesn’t specify line length, speed or style.
Line length
Long lines make the eye travel further and cover more of the picture. The BBC limits broadcast subtitles to 37 characters per line, a constraint inherited from Teletext. For online video it sets the limit by width instead: 68% of the width of a landscape 16:9 video, or 90% of a vertical 9:16 video. Netflix’s English guide allows up to 42 characters per line.
Vertical video is narrow, so far fewer characters fit at a readable size. The BBC gives a useful equivalence: 37 characters in a 75%-wide region of a 16:9 video corresponds to 25 characters across 90% of a 9:16 frame. For TikTok, Reels and Shorts, treat about 25 characters per line as a ceiling, and go shorter if your caption style uses large, heavy type.
Lines per caption
Netflix allows a maximum of two lines, and the DCMP Captioning Key prefers no more than two. The BBC recommends two lines for landscape, 4:3 and square video, and up to three for vertical 9:16 video, where each line is shorter. On a busy vertical edit, one line — or two short ones — keeps more of the picture clear and gives the viewer less to take in at once.
Reading speed
Reading speed is measured in words per minute (wpm) or characters per second (cps). Here’s what each guideline allows:
| Guideline | Limit | Applies to |
|---|---|---|
| BBC Subtitle Guidelines | 160–180 wpm (0.33–0.375 seconds per word) | Recommended subtitle speed |
| Netflix English (USA) | Up to 20 cps | Adult programs |
| Netflix English (USA) | Up to 17 cps | Children’s programs |
| DCMP Captioning Key | No more than 130, 140 or 160 wpm | Lower-, middle- and upper-level educational media |
To check a caption, divide its character count by the time it’s on screen. A 38-character caption shown for 1.5 seconds runs at about 25 cps — faster than Netflix allows for adult programs. Show it longer if the speech leaves room, or split it into two captions.
Fast talkers are the hard case in social video. The BBC notes that viewers tend to prefer verbatim subtitles, and on social media the speaker’s exact phrasing is often the point. Rather than rewriting what someone said, break it into more, shorter captions timed to the speech. Condense only when the words truly can’t be read in time, and never in a way that changes the meaning.
Timing and duration
- Start with the speech. The BBC’s guidelines say to match a subtitle to the moment speech begins and to keep any lag behind the speech to a minimum.
- Minimum duration. Netflix sets a minimum of five-sixths of a second (20 frames at 24 frames per second). The DCMP Captioning Key sets 40 frames — 1 second and 10 frames. The BBC suggests aiming for around 0.3 seconds per word, so a four-word subtitle stays up for about 1.2 seconds.
- Maximum duration. Netflix caps a subtitle at 7 seconds and DCMP at 6. A caption left up long after the words stop no longer matches what the viewer hears.
- Shot changes. The BBC advises against subtitles that straddle a shot change: split the sentence, or start the new caption on the cut. Fast social edits make that hard to follow strictly, but avoid a caption that flashes up for a few frames just before a cut.
Where to break lines
Break lines where a speaker would naturally pause, ideally after punctuation. The BBC lists pairs that should stay on the same line:
- an article and its noun (the table, a book);
- a preposition and the phrase after it (on the table);
- a conjunction and its clause (and those books);
- a pronoun and its verb (they will come);
- the parts of a complex verb (will have been doing).
Netflix’s English guide says to break after punctuation and before conjunctions or prepositions, and not to separate a name, an adjective from its noun, or a subject from its verb. The DCMP Captioning Key likewise says not to split a person’s name, or a title from the name it belongs to.
Awkward We launched the
new app in March.
Better We launched
the new app in March.Line breaks aren’t the top priority, though. The BBC says that when good line breaks conflict with well-edited text and good timing, the text and the timing matter more.
Placement on vertical video
Traditional captions sit at the bottom center of the frame and move when they would cover something important. DCMP says to move captions to the top when they would obscure a speaker’s face or essential on-screen text, and Netflix places subtitles wherever they’re easier to read when text appears on screen.
Vertical platforms add another obstacle: the app’s own interface. On TikTok, Reels and Shorts, the creator’s name and post text sit over the bottom of the video, and the like, comment and share buttons run down the right edge.
- Raise the captions. The BBC notes that in vertical video it’s common to position subtitles a little higher up, though generally still in the lower third, because faces tend to be in the top half of the frame.
- Leave a right-hand margin so words don’t slide under the buttons, and keep lines centered.
- Keep mouths and on-screen text clear. Move a caption rather than let it overlap a product shot or a title.
- Test on a phone. App layouts change and vary between devices, so check a test upload on a real screen rather than trusting a fixed template.
Contrast, outline and size
A caption has to stay readable over every frame, from a dark room to a white wall behind the speaker. There are three dependable ways to get there:
- A background box. The BBC’s default is white text on a black background. The DCMP Captioning Key prefers a translucent box, so text stays clear on light backgrounds.
- An outline or shadow. DCMP asks for white, medium-weight, sans serif characters with a drop or rim shadow. For bold social styles, a solid dark outline does the same job.
- Measured contrast. WCAG’s minimum contrast ratio for text is 4.5:1 (success criterion 1.4.3). It’s written for text on web pages rather than text in video, but it’s a sensible benchmark for caption text against its box or outline.
For size, the BBC sets subtitle text by line height as a share of the video’s height: 7–8% for landscape, 4:3 and square video, and 3.9–4.5% for vertical 9:16 video. Social caption styles are often larger and bolder than that. That can work, as long as lines stay short enough to fit and the text doesn’t cover the subject.
Finally, letter case. DCMP prefers mixed case for readability and reserves capitals for shouting. All-caps styles are common in short-form video; if you use one, keep your captions even shorter.
Speakers, sounds and music
Viewers who can’t hear the audio — because they’re deaf or hard of hearing, or just scrolling with the sound off — need more than the dialogue.
- Speaker changes. Make it obvious when someone new speaks. The BBC uses a small set of text colors on black — white, yellow, cyan and green — and keeps each speaker’s color consistent. Netflix’s English guide adds a speaker ID in brackets only when viewers can’t see who is talking. DCMP prefers placing the caption beneath the person speaking, with the name in parentheses when placement can’t do the job.
- Sound effects. Caption the sounds that matter to the story or the joke. Conventions differ: the BBC writes sound labels in white capitals as short subject-and-verb phrases, like FLOORBOARDS CREAK; Netflix uses brackets and lowercase, as in [door slams]; DCMP puts descriptions in brackets and names the source of the sound.
- Music. DCMP sets sung lyrics between music-note symbols (♪). For background music, a short bracketed description of the mood is enough when it matters.
- Pick one convention and stick to it. Consistency within a video matters more than which system you choose.
Accuracy: names, numbers and jargon
Automatic speech recognition gives you a fast first draft, and it makes predictable mistakes. YouTube’s help center warns that its automatic captions can misrepresent what was said because of mispronunciations, accents, dialects or background noise. Check these by hand every time:
- Names of people, brands, products and places, including their capitalization.
- Numbers, prices, dates and units. Pick a number style and apply it consistently. DCMP spells out one to ten and uses numerals above ten.
- Jargon and acronyms from your field, which a general-purpose model may not know.
- Words that sound alike, such as their and there, affect and effect, four and for.
- Fillers and false starts. Decide whether to keep “um” and repeated words. Removing them helps reading speed; keeping them preserves the speaker’s tone. Never remove a word that changes the meaning.
The fix should be quick. In AutoCap, every word carries its own start and end time, so you can correct a misheard name, split a long line or nudge a single word without retiming everything around it.
Word-by-word and karaoke styles
Word-by-word captions show one to a few words at a time, or reveal a line one word after another. Karaoke styles show the whole line and highlight each word as it’s spoken. Both are common in short-form video because movement draws the eye. They also change how people read:
- Less text at once means no reading ahead. With a full line on screen, a viewer can take in the phrase and look back at the picture. When words arrive one at a time, they can only read at the speaker’s pace, and their eyes stay on the text.
- Broadcast guidelines treat it as an exception. The BBC says cumulative subtitles — where part of a subtitle appears after the rest — should be used only when there’s a good reason, such as dramatic impact or the rhythm of a song.
- Animation competes with the words. A bounce or color change on every word is more to process than a steady line with one highlight.
Use them for short, punchy clips, hooks in the first seconds, and emphasis on a key word. Switch to a steadier style for tutorials, interviews, long explanations and fast speech. Karaoke highlighting is often the best compromise: the full phrase stays readable while the active word is marked. It only works if the highlight lands on the word as it’s spoken, which is why per-word timing matters — a highlight running half a second early reads as a glitch. You can compare approaches in the caption style gallery.
Checklist
- One or two lines per caption; up to three short lines on vertical video.
- About 25 characters per line or fewer on vertical video; no more than 37–42 on landscape video.
- Reading speed no faster than about 20 characters per second (17 for children’s content), or 160–180 words per minute.
- No caption on screen for less than about a second, or for more than 6–7 seconds.
- Captions start with the speech and avoid straddling cuts.
- Lines broken at natural phrases, never splitting names or article-noun pairs.
- Captions clear of faces, on-screen text and app buttons, checked on a phone.
- High contrast from a box, outline or shadow, tested on light and dark frames.
- Speaker changes and meaningful sounds marked in one consistent style.
- Names, numbers and jargon checked by a person.
- Word-by-word animation saved for short clips.
For when to burn captions in and when to publish a caption file, see open vs closed captions.