VidTranscriber

VidTranscriberturn video into a transcript you will use

VidTranscriber transcribes YouTube links, uploaded video and audio files, and Zoom recordings into searchable text with speaker labels, AI chapters, and exports in TXT, SRT, VTT, Markdown, and DOCX.

freemium
video transcriptionYouTube transcriptZoom notesspeaker detectionSRT exportmeeting notespodcastsubtitles
Indexed
Updated
3.1 (0 Number of reviews)

Log in to rate the project

Try Now

What VidTranscriber does

VidTranscriber is a web app that converts video and audio into editable text. Users can paste a YouTube link, upload a local file in MP4, MOV, MP3, WAV, M4A, or WEBM format, or point the tool at a Zoom recording. Finished transcripts are searchable, editable in the browser, and enriched with automatic speaker labels for up to six voices.

AI features and languages

Beyond the raw transcript, VidTranscriber generates AI chapter breaks and a summary that highlights the main talking points. The service advertises support for more than 100 languages, so a lecture, meeting, or podcast can be captured in the original language and then handed off to translation or note-taking workflows.

  • Sources: YouTube links, uploaded files, and Zoom recordings.
  • Exports: TXT, SRT, VTT, Markdown, and DOCX.
  • Retention: uploaded files are auto-deleted after 24 hours.

Pricing tiers

A guest run gives 60 minutes one time with no account. The Free tier includes 300 minutes per month and 30-day transcript storage. The Pro tier is listed at USD 9 per month for 1,500 minutes, unlimited storage, and access to the higher accuracy mode.

Pros & Cons

Pros

  • Accepts YouTube links, files, and Zoom recordings in one workflow.
  • Speaker labels and AI chapters make transcripts easier to skim.
  • Free tier of 300 minutes per month is generous for light users.

Cons

  • Uploaded files are removed after 24 hours by design.
  • Speaker detection is capped at six voices.

Frequently Asked Questions

What sources can VidTranscriber handle?

YouTube links, uploaded MP4, MOV, MP3, WAV, M4A, and WEBM files, and Zoom recordings.

What export formats are available?

TXT, SRT, VTT, Markdown, and DOCX.

How much does Pro cost?

Pro is listed at USD 9 per month for 1,500 minutes with unlimited storage and the higher accuracy mode.

Explore More

Open-source Alternatives

Steno: Open-Source AI Note-Taking for High-Confidentiality Conversations

Steno is an open-source AI note-taking tool designed for high-confidentiality conversations. It runs locally or on private servers, automatically generating structured notes from meeting recordings while ensuring data never leaves your controlled environment. Ideal for government, defense, legal, and executive teams. Primary language is TypeScript, license is MIT.

vexa: Open-source meeting transcription API

vexa is an open-source meeting transcription API that seamlessly integrates with Google Meet, Microsoft Teams, and Zoom. It automatically joins meetings and provides real-time transcription via WebSocket. An MCP server enables AI agent integration. Users can self-host or use the SaaS offering. Written in Python, vexa has over 2500 stars on GitHub, making it a robust solution for automating meeting documentation and AI-powered analysis.

typewhisper-mac: on-device speech-to-text for macOS

typewhisper-mac is an open-source macOS app for on-device speech-to-text. It leverages local AI on Apple Silicon or Intel for real-time, privacy-focused transcription without an internet connection. Supports multiple languages including Chinese and English, with an optional cloud mode for enhanced accuracy. Built with Swift and licensed under GPL-3.0, it has over 1500 GitHub stars.

Handy: Turn keyboard shortcut into local speech-to-text

Handy is a cross-platform desktop app that turns a keyboard shortcut into speech-to-text: press, speak, and your words land in whatever text field has focus. Everything runs locally with no cloud upload, using Whisper or Parakeet V3 models and Silero voice activity detection. Built with Rust and Tauri on the backend and React with Tailwind on the front end, it ships for macOS, Windows, and Linux under the MIT license. As of collection time, it had 23,022 GitHub stars.

amical: Local-first AI dictation for offline speech-to-text

amical is an open-source, local-first AI dictation application that leverages open models like Whisper for fast, accurate, and offline speech-to-text. It claims to triple typing speed without a keyboard, supports multiple languages, and prioritizes user privacy by processing all data on-device. Built with TypeScript, it is suitable for both developers and general users.