Speaker separation
Speaker recognition labels who said what, so interviews and meetings stay readable.
Upload an audio or video file for transcription now.
Start a transcript
Drop your file and pick a language.
Upload a file
Drop audio or video here
MP3, MP4, M4A, WAV, MOV and more
Picking the language improves accuracy. Leave it automatic if you are unsure.
By continuing you agree to our Terms of Service and Privacy Policy.
How it works
Audio or video of any length; the free plan transcribes the first 30 minutes of each file. MP3, MP4, M4A, WAV, MOV and a dozen more.
Leave the language on automatic or set it for better accuracy. Turn on speaker recognition if you need to know who said what.
The transcript opens with clickable timestamps — read it and fix a name straight away. A free account keeps it and exports it as TXT; DOCX, PDF, SRT and VTT come with the unlimited plan.
What you get
Not a wall of text. Every line is anchored to the audio, labelled by speaker and ready to export.
Speaker recognition labels who said what, so interviews and meetings stay readable.
Transcribe in the language spoken, then translate the result without uploading again.
Every line is anchored to the recording. Click a sentence to hear exactly that moment.
DOCX, PDF, TXT, SRT and VTT. Subtitles come out ready — no conversion step.
Correct names and jargon in the browser; the fixes carry into every format.
Files are processed for your account only, never used to train models, and deleted when you say so.
FAQ
Short answers. If something is still unclear, the first file is free anyway.
Start a transcriptNot for your first file of the day — upload it and the transcript starts immediately. A free account adds 90 transcript minutes a month, keeps your files and lets you download them; everything you made as a guest comes with you.
The first file each day is free without an account, and a free account gives you 90 transcript minutes a month, up to 30 minutes each, no credit card. Unlimited transcription is $10 a month billed yearly, or $20 month to month.
Audio and video alike: MP3, M4A, WAV, OGG, OPUS, FLAC, AAC, WMA, MP4, MOV, WEBM, WMV and MPEG.
Clear recordings come out close to perfect. Heavy accents, people talking over each other and background noise cost accuracy.
Yes. Turn on speaker recognition and each line is tagged with its speaker — you can rename them in the editor.
It is stored for your account alone and never used to train models. Delete the file and the audio goes with it.
No account, no card. See the transcript before you decide anything.
Start a transcriptAlready have an account? Log in
Your request is open
We sent a confirmation to your email. You will get our reply by email and can read it on the request page.
Open the request