Transcribe audio to text, free and private

MP3, M4A, WAV, OGG, Opus or a video: speech recognition writes it down in your browser. No upload, no minute limit, around a hundred languages.

🔒 Speech recognition runs in your browser. It is downloaded only once; your recording is sent nowhere.

Transcribe a voice message without leaving WhatsApp

Install AnyFileFormat as an app and it appears in your phone's Share menu. The voice message is transcribed on your device and sent nowhere.

  1. Install the app (once).
  2. In WhatsApp or Telegram, press and hold the voice message, then tap Share.
  3. Pick AnyFileFormat: transcription starts by itself.

How to

  1. Drop an audio or video file.
  2. Keep automatic language detection and pick a mode: Fast, Balanced or Accurate.
  3. Transcribe, correct any word in place, then copy the text or download TXT, SRT or VTT.

Drop the audio recording above and you get its text in a few seconds to a few minutes. Speech recognition runs in your browser: the audio is never uploaded, there is no length limit and no account.

What you get

  • This is speech recognition, not a format change: it listens to the recording and writes down what is said, with punctuation and capitals.
  • The text is split into paragraphs at the pauses. Timestamps are kept on screen, and in the SRT and VTT exports.
  • Around a hundred languages, detected automatically. It can also translate the speech into English.
  • Clear speech comes out close to perfect. Background noise, people talking over each other and rare names cause mistakes, which you can correct in place before downloading.
  • Speakers are not labelled: the text says what was said, not who said it.

What it is for

Use plain text for notes, minutes, quotes from an interview, a voice message you cannot listen to right now, or anything you want to search, copy or translate.

Because nothing is uploaded, it is suitable for confidential recordings: medical, legal, HR interviews, or private voice notes.

Frequently asked questions

Is my audio file uploaded?

No. Speech recognition is downloaded to your browser once and runs there. The recording itself never leaves your device: you can go offline once it has loaded and it still works.

How accurate is it?

On a clear voice, Accurate makes very few mistakes; Fast makes more but runs on any phone, and Balanced sits in between. Every line can be corrected before you download.

How long does it take?

On a recent computer, a minute of speech takes a few seconds with Fast or Accurate, and a little longer with Balanced. Phones are slower. Silences are skipped, so a long recording with pauses goes faster than its length suggests.

Why is there a download the first time?

Recognising speech on your device means downloading the recognition itself: 80 MB for Fast, 245 MB for Balanced, 545 MB for Accurate. It is stored in your browser, so later transcriptions start at once.

Which languages are supported?

About a hundred, including English, French, Spanish, German, Italian, Portuguese, Dutch, Arabic, Chinese, Japanese and Russian. The language is detected automatically; you can also set it yourself.

Is it free? Is there a length limit?

Free, with no account and no limit on minutes: your device does the work, so there is nothing to meter. Very long recordings are limited only by your device's memory.