All posts
Speech AI 2 min read

Fast, Balanced, Best: which quality level to pick

Three levels, three different models. What actually changes between them, and when the slowest one earns its wait.

What the three levels are

Every upload runs through one of three levels, and each is a different model:

  • Fast — Whisper large-v3-turbo. A distilled version of the large model: fewer decoder layers, most of the accuracy, a fraction of the time.
  • Balanced — Whisper large-v3. The full model, and the default.
  • Best — GPT-4o Transcribe. A newer architecture that handles accents, crosstalk and noise better than Whisper does.

All three return word-level timing, so the player follows along and subtitle export works the same at every level.

What actually changes

Not much, on clean audio. A single speaker with a decent microphone in a quiet room lands in nearly the same place at all three levels — which is why paying the wait for Best on that file is wasted.

The gap opens where the audio gets hard: overlapping speech, strong accents, background noise, technical vocabulary, proper nouns. Fast starts dropping short interjections and guessing at names. Best keeps them.

Choosing without thinking too hard

Leave it on Balanced. It is the default because it is right most of the time.

Reach for Fast when you want the gist of a long recording quickly, or when the audio is clean and the stakes are low — a voice memo, a solo recording, something you are going to skim.

Reach for Best when a mistake costs you: interviews you will quote, legal or medical recordings, heavy accents, several people talking over each other, or anything you are publishing as subtitles.

You can always re-run a file at a higher level. The original audio stays available for as long as your plan keeps it.

Blog

Turn your own recording into text

Upload a file and read the transcript in minutes — no signup needed.

Transcribe a file