How to Format a Video Transcript

A good video transcript shows who is speaking, what they said and where it happens in the video: one speaker label per turn, readable paragraphs and a timestamp at each speaker change. Format it as a document for reading and search; captions and subtitles need short timed segments instead.

Upload a video to get the first draft
or drag & drop it here
MP4, MOV, WEBM and more · first 5 minutes free
Up to 99% Accurate90+ Languages1-Hour Audio in 2 Min30+ File Formats
4.8 ratingTrustpilotG2GDPRSSL

Key facts

Core layoutSpeaker label + timestamp per turn, then the text as a readable paragraph
Transcript vs captionsA transcript is read as a document; captions are short timed segments shown on screen
Subtitle limits (Netflix, English)42 characters per line, 2 lines, up to 20 characters per second
Our SRT/VTT draftsOne cue per transcript segment: typically 1.7 seconds and 31 characters
ExportsTXT and JSON on a free account; DOCX, PDF, SRT and VTT on Standard and Premium

A basic video transcript template

Most video transcripts work best in one simple layout: a timestamp and a speaker label at the start of each turn, followed by what that person said as a normal paragraph. It reads like a document, and every timestamp lets a reader jump back to the exact moment in the video.

  • [00:00:08] Host — Welcome. Today we are going to show how the new reporting workflow works.
  • [00:00:17] Guest — The first change is that every report now keeps a visible revision history.
  • [00:00:28] Host — Where can a user find the earlier versions?

Three rules keep it usable. Start a new paragraph whenever the speaker changes. Use the same label for the same person throughout — "Host" in one place and "Anna" in another confuses readers and search. And correct names and specialist terms before anyone else sees the transcript, because those are the words automatic transcription gets wrong most often.

Video transcript format: speaker labels, timestamps and export options

What to include in a video transcript

Speaker names

Use a real name when it is known and appropriate to publish. Otherwise use a clear role — Host, Guest, Interviewer, Speaker 1. On Standard and Premium, speaker labels are added automatically and can be renamed once for the whole transcript.

Timestamps

Put a timestamp at each speaker change, or every 30–60 seconds in a long monologue. Editors and researchers want more of them; a published transcript reads better with fewer.

Paragraphs and punctuation

Break long speech into paragraphs of a few sentences. Add punctuation for meaning, but do not rewrite what the speaker claimed.

Non-speech information

Add markers such as [music], [laughter] or [inaudible] only where they help a reader follow the video. A transcript full of background sounds is harder to read, not more accurate.

Clean read or verbatim?

Decide this before you edit, because it changes almost every line. A clean-read transcript removes filler words, repetitions and false starts when they add no meaning. It is easier to read and publish, but the editor must not change what the speaker intended. A verbatim transcript keeps the "um", the repeated words and the unfinished sentences.

Clean readVerbatim
Keeps filler wordsNoYes
Keeps false startsNoYes
Best forArticles, show notes, course pages, blog postsResearch analysis, interviews where hesitation matters
Main riskOver-editing that changes meaningHard to read in long passages

Whichever you choose, apply it to the whole transcript and say so at the top. An automatic transcript is a working draft: if yours must meet a legal, medical, research or publication standard, write that standard down and review the draft against it.

Transcript, captions and subtitles are different things

The three are often mixed up, but they are built for different jobs. The W3C accessibility guidance draws the line clearly: captions are for the same language as the spoken audio and include the non-speech sounds needed to understand the content, while subtitles carry speech translated into another language.

OutputMain useStructure
TranscriptReading, search, documentationSpeaker labels, paragraphs, optional timestamps
CaptionsAccessibility in the original languageShort timed segments, speech and important sounds
SubtitlesViewers who speak another languageShort timed segments, translated

A transcript formatted as a document should never be uploaded as captions as it is: it has no timing and its paragraphs are far too long for the screen. Going the other way works — a caption file already holds the words and timing a transcript needs.

What is inside an SRT or VTT file

Both are plain-text files of numbered or unnumbered cues: a time range and the line or two of text shown during it. The differences are small but break uploads when ignored:

  • SRT numbers every cue and writes milliseconds after a comma: 00:00:08,000 --> 00:00:12,500.
  • VTT must start with the word WEBVTT and writes milliseconds after a period: 00:00:08.000 --> 00:00:12.500, as defined in the W3C WebVTT specification.
  • VTT is the format web players read natively; SRT is the one most video editors and platforms accept.

Readability limits come from style guides rather than the formats. Netflix's English timed text style guide, one of the most widely referenced, allows 42 characters per line, two lines per subtitle, and a reading speed of up to 20 characters per second for adult programmes (17 for children's).

How TranscribeThis subtitle drafts are built — and what to fix

Our SRT and VTT exports create one cue per transcript segment, keep the segment's exact start and end times, and do not add speaker names. We measured what that produces: across 333,000 segments from 502 transcripts processed in the last 30 days, the typical cue lasts 1.7 seconds and holds 31 characters; nine in ten are under 4.4 seconds and 65 characters.

So length is rarely the problem — only 3.7% of cues exceed two 42-character lines. Speed is: about half of the cues run faster than 20 characters per second, because the timing follows speech closely and leaves no extra time on screen.

  • Merge neighbouring short cues from the same speaker into one subtitle.
  • Extend cue end times into the pauses so each line stays up long enough to read.
  • Split the few long cues at a natural phrase break.

Which export format to choose

Pick the format by where the transcript goes next. Each export below describes exactly what TranscribeThis puts in the file.

FormatWhat is insidePlan
TXTOne paragraph per speaker turn ("Speaker 1: …"), renamed speakers, no timestampsFree account and up
JSONFull text, language, duration and every segment with start, end, text and speakerFree account and up
DOCXEditable document; with timestamps on, a bold speaker name at each changeStandard, Premium
PDFThe same layout as DOCX, as a fixed documentStandard, Premium
SRTSubtitle draft, one numbered cue per segmentStandard, Premium
VTTSubtitle draft for web players, one cue per segmentStandard, Premium
Timestamped documents
For DOCX or PDF with a timestamp on every segment, turn on "Include Timestamps in Export" in Settings → Transcription → Export Preferences before exporting.

How to create the first draft

Typing a transcript from scratch takes several times the length of the video. Starting from an automatic draft turns the job into editing. The Video to Text converter takes MP4, MOV, WEBM and 30+ other formats; without an account it transcribes the first 5 minutes of files up to 50 MB, and Standard and Premium handle full videos up to 5 and 8 hours.

  1. Upload the video on the Video to Text page and wait for the whole file to be processed.
  2. Rename the speaker labels to real names or clear roles such as Host and Guest.
  3. Correct names, numbers and specialist terms against the video.
  4. Choose clean read or verbatim, and apply that choice to the whole transcript.
  5. Export the format that fits the final use, with timestamps turned on if readers need them.

The transcript appears after the whole video has been processed — this is not live captioning. If the audio is noisy or several people talk over each other, plan extra time for step 3.

Frequently Asked Questions

What is the standard format for a video transcript?

There is no single universal layout. The practical standard is a timestamp and speaker label at each turn, followed by the speech as a readable paragraph. Formal projects such as research studies or broadcast publishing often set their own style guide.

How often should timestamps appear?

At every speaker change, plus every 30–60 seconds during long monologues. More timestamps help editing and verification; fewer make a published transcript easier to read.

Should I remove filler words?

Remove them in a clean-read transcript when they add no meaning. Keep them in a verbatim transcript when hesitation or exact wording matters. In both cases, never change what the speaker meant.

Can a transcript be used as captions?

Not as a document. Captions need short timed segments, so start from an SRT or VTT export, then merge short cues and check line length, reading speed and sync against the video.

What is the difference between SRT and VTT?

Both hold timed text cues. SRT numbers each cue and uses a comma before the milliseconds; VTT starts with the word WEBVTT and uses a period. VTT is native to web players, SRT is accepted by most video editors.

Related resources

Reviewed by the TranscribeThis product team · Last updated: October 2026