How to Transcribe an Audio Recording

To transcribe audio, upload or record your file, let TranscribeThis generate an AI draft, then review it against the recording: check speaker labels and word-level timestamps, correct names, numbers, and specialist terms, and export as TXT, DOCX, PDF, SRT, or VTT. A transcript is a working draft until you verify it — the AI is a fast first pass, not a finished record.

Upload or record audio to transcribe
or drag & drop it here
MP3, WAV, M4A, MP4 and more · first 5 minutes free, no account
Up to 99% Accurate90+ Languages1-Hour Audio in 2 Min30+ File Formats
4.8 ratingTrustpilotG2SOC 2GDPRSSL

At a glance

InputUpload a file, record in your browser (account), or paste a YouTube link
OutputEditable transcript with word-level timestamps; speaker labels on Standard/Premium
Free (no account)First 5 minutes transcribed, files up to 50 MB, preview export
Free account3 transcriptions/day, first 5 minutes each, 150 MB storage
Paid limitsUp to 2 GB / 5 hours (Pro), 5 GB / 8 hours (Business) per file
Formats inMP3, WAV, M4A, AAC, FLAC, OGG, Opus, WebM, MP4, MOV, MKV and more
Formats outTXT, DOCX, PDF, SRT, VTT
ProcessingUpload-first: the file is processed after upload (not live/real-time)

The short version

Transcription with an AI tool is a five-part workflow: pick the cleanest source file, upload or record it, choose language and speaker options, review the draft against the audio, and export the format you need. The draft appears in a few moments, but the transcript is not "done" until you have checked the parts machines get wrong — names, numbers, homophones, and overlapping speech.

Product boundary
TranscribeThis transcribes the spoken audio in a recording. It does not read on-screen text, translate video into captions automatically, or replace a certified human transcript for legal or medical records.
How to Transcribe an Audio Recording

Step 1 — Choose the best source file

Accuracy is decided before you upload. Start from the original recording, not a re-compressed copy shared over chat, and prefer a file where the speakers are close to the microphone with little background noise.

  • Use the original file where possible — every re-export loses detail.
  • Mono or stereo both work; TranscribeThis decodes everything to 16 kHz mono before recognition, so "lossless" formats do not add accuracy on their own.
  • For interviews, a per-speaker channel or a close mic helps speaker separation more than a higher bitrate.

Step 2 — Upload or record

Drop a file into the uploader at the top of this page. Without an account you can transcribe the first 5 minutes of a file up to 50 MB — enough to check quality on your own audio. Recording directly in the browser is available once you sign in.

You can also paste a YouTube link (it uses the video's existing captions where available, so YouTube transcripts are segment-level rather than word-level). Google Drive and Dropbox links are available to signed-in users.

Step 3 — Set language and speaker options

The recognizer auto-detects the spoken language. If your recording has more than one voice and you are on Pro or Business, enable speaker identification (diarization) so the transcript is split by speaker; you can rename "Speaker 1" to real names afterwards. Diarization is not available on the free tier.

Step 4 — Review against the recording

This is the step people skip and regret. Every line carries a word-level timestamp linked to the audio, so you can click a word to hear exactly what was said. Play back the passages that matter and confirm the wording rather than trusting the draft.

  • Scan for [inaudible] or low-confidence passages first.
  • Listen to any sentence you plan to quote or act on.
  • Fix speaker turns where two people talk over each other — overlap is the hardest case for any model.

Step 5 — Correct names, numbers, and terms

AI transcription is strongest on ordinary conversational speech and weakest on things it cannot infer from context: proper nouns, product names, figures, dates, and domain jargon. Correct these deliberately — a mis-heard number or name is the error most likely to cause real damage downstream.

Step 6 — Export the right format

Export once the transcript reads correctly. Choose the format by where the text is going:

FormatUse it for
TXTPlain text for notes, search, or pasting elsewhere
DOCXAn editable document for reports or sharing
PDFA fixed, shareable copy
SRT / VTTTimed captions to review and load into a video editor
Note
Full export is a paid feature; the free tier gives a preview. SRT/VTT are reviewable caption drafts, not finished, styled, accessibility-certified subtitles.

Automatic vs manual transcription

Manual transcription — typing while you listen — is slow but gives you full control. AI transcription inverts that: you get a complete draft in minutes and spend your time reviewing instead of typing. For most meetings, interviews, lectures, and voice notes, an AI draft plus a focused review is faster and more consistent. When a transcript must be certified or verbatim for a court or medical record, a professional human service is the right tool.

Accuracy, and what affects it

No transcription tool is accurate as a single fixed percentage — it depends on the recording. Clear, single-speaker audio transcribes very well; heavy accents, crosstalk, music, and low-quality phone recordings need more review. Treat any "99% accurate" headline (ours or a competitor's) as a best case on clean audio, not a guarantee for your file.

Privacy and consent

Files are encrypted in transit (TLS). You control how long transcripts are kept with a configurable retention setting, and you can delete a file at any time; guest uploads are removed automatically within a couple of hours. Recording other people may require their consent depending on where you are — check the rules for your jurisdiction before recording a call or meeting.

Frequently Asked Questions

How do I transcribe an audio file for free?

Upload it to the tool at the top of this page. Without an account you can transcribe the first 5 minutes of a file up to 50 MB and preview the result. A free account gives 3 transcriptions a day.

What audio formats are supported?

MP3, WAV, M4A, AAC, FLAC, OGG, Opus, WebM, WMA, AMR and more, plus the audio track of video files like MP4, MOV, MKV, and AVI.

Does it transcribe in real time as I speak?

No. TranscribeThis is upload-first: you record or upload a file and it is processed afterwards. It does not stream a live transcript while you speak.

Can it tell speakers apart?

Yes, on the Standard and Premium plans. Speaker identification labels each turn and you can rename speakers. The free tier does not include diarization.

How long can a file be?

The free tier transcribes the first 5 minutes of each file. Paid plans handle up to 5 hours (Pro) or 8 hours (Business) per file.

Are the timestamps accurate to the word?

For uploaded and recorded audio, yes — each word carries its own timestamp linked to the recording. YouTube transcripts are segment-level because they come from the video's captions.

Can I get a Word or PDF file?

Yes. You can export TXT, DOCX, PDF, SRT, and VTT once the transcript is ready (full export is a paid feature).

Is my audio kept private?

Files are encrypted in transit, retention is configurable, and you can delete files yourself. Processing partners operate under no-training terms, so your content is not used to train AI models.

Related resources

Reviewed by the TranscribeThis product team · Last updated: July 2026