A transcript is there to be searched and skimmed
Aliases: full transcript · find-in-page transcript · skimmable transcript
What it is
A transcript is a complete text you can search and skim. Its job is not to roll along with the picture one line at a time. It is to let someone find the quote, the API name, the conclusion inside an hour of material. Captions hand you a few seconds. The transcript turns the session into a document.
Why it happens
Type on a timeline occupies working memory: the next line replaces the last, and the whole cannot be queried. A spatial document lets the eye jump headings, scan speaker turns, and hit a string with find-in-page. The cost gap is orders of magnitude: listen to the episode, versus type a proper name.
Screen-reader users can jump by document structure instead of sitting through the media as the only linear path. Anyone who must quote (journalists, students, support checking a promise) needs copyable text, not a screenshot of captions. Search only holds if the transcript is real text in the page — not a picture trapped in the player that neither the site search nor find-in-page can enter.
Studying it
Give the same material to one group with captions only and one with a full transcript. Set find tasks: where is the price, who first proposed the delay. Do not make “does clicking a sentence seek the player” the main dependent measure; that is alignment.
Independent variables: full text present or not, headings or speaker turns present or not, whether the text is selectable. Dependent variables: time to the target, misses, false hits on a similar but wrong passage.
Bury the target in the second half. Facts visible in the first minutes do not discriminate the two media.
Where it stops holding
A two-second button cue or a CAPTCHA utterance is not worth indexing; visible text already covers it. A live conversation has no final transcript yet; what you search is a draft that will change. Some speech must not leave a searchable copy (privacy, one-time codes). Then persistence conflicts with findability — drop the retained transcript rather than tell people to listen. Very short social voice notes whose words already sit in the chat are duplicates if reattached as a transcript.
Applying it
- Put a complete transcript on the media page as real text, not only inside the player.
- Break long pieces by topic or speaker so skimming has somewhere to land. Find-in-page must hit words in the transcript.
- Do not rasterise the transcript or put it on a canvas; copy and assistive technology both fail.
- How to check: stop the player. Find a number or name buried in the second half using only the text, and time it. If the media has to be heard through, the transcript is not doing search.