How to transcribe audio and video to text

Last updated: by Zendocs Support Team

Zendocs transcription, powered by Whisper speech-to-text, turns meetings, interviews, lectures, and podcasts into editable text faster than real time, with no manual typing.

How to transcribe a recording

Open Transcribe Audio & Video.

  1. Add your recording: upload from your device, import from Google Drive or Dropbox, paste a YouTube link, or dictate live.
  2. Wait while the transcript builds.
  3. Review the transcript and correct names, numbers, or jargon as needed.
  4. Export the result as TXT, DOCX, or PDF.

What you get

  • 99+ languages, with automatic language detection or manual selection.
  • Speaker recognition: each segment is labeled with who is speaking.
  • Timestamps on every segment, so you can jump to any moment and verify quotes against the audio.
  • Export as TXT, DOCX, or PDF.

Supported formats and sources

Supported sources include direct upload, Google Drive, Dropbox, YouTube links, or live dictation. Common audio and video formats include MP3, WAV, M4A, MP4, and MOV.

File size limits

Files up to 100 MB or roughly 3 hours are supported. See Supported file types and size limits.

Need more help?

Our support team is available around the clock: contact support.

Connect With ZenDocs Support

Our support team is available around the clock.

Contact Support