Turn audio, podcasts and recordings into subtitles on your Mac

You have a recording, a podcast episode, an interview or a voice memo, and you want timed subtitles, or text you can copy and search, without sending the audio to someone else's server. Twinsubs is not only for video: drop in an mp3, m4a or wav and you get an .srt beside it just the same. If the episode is in a language you do not understand, the same run can add a translation so you can read along as you listen.

What it can process

Twinsubs main window: files in a queue, with the finished rows listing the subtitle files that were created
The main window: audio files go into the queue just like video, and the finished subtitle files sit beside the originals.

Three steps

  1. Drop it inDrag the audio file into the window.
  2. Wait a momentThe source language is detected for you, or you can set it by hand; you choose the target language. To get the original text only, pick the same language as the target. The first launch needs a connection to download a speech model once (under 1 GB).
  3. Find the subtitles beside the audioThe same folder gains an .srt: a single-language file such as Podcast.en.srt, and for foreign-language audio a bilingual one such as Podcast.en-zh.srt (target first, source second).

What you can do with the SRT

In file mode Twinsubs writes SRT. It does not export txt directly or a transcript with speaker names. For plain text, convert the SRT as described above.

Which audio suits it

When a recording contains other people's voices, make sure you are allowed to process it, and tell them first where that is appropriate.

Limits, stated up front

Privacy: the audio stays on your Mac

Recognition and translation both run on your Mac. There is no backend, so the audio never passes through a server. Only the first launch needs a connection, to download the speech model once; translation language packs come from macOS and work offline afterwards. See the privacy policy.

FAQ

Which audio formats work?

mp3, m4a, wav, aac, aiff, flac and caf. Video files work too, and it is their sound that gets recognized.

Can it export a text transcript directly?

File mode writes SRT. An SRT is plain text, and removing the numbers and timecodes leaves a transcript; the method is in the guide on checking and editing SRT.

Can it tell who is speaking?

No. A conversation comes out as continuous subtitles with no speaker labels.

Can it translate a foreign-language podcast at the same time?

Yes. Pick your target language and you also get a bilingual file with the target on top and the source below. For the original text only, pick the same language as the target.

What about very long audio?

The free version processes the first 10 minutes of each file; the purchase processes the full length. Speed is similar to video, and progress is always visible.

Does it need a connection?

Recognition and translation both run on your Mac, and the audio is never uploaded. The first launch needs a connection to download the speech model once; after that recognition runs on your Mac.

The podcast is playing right now. Can I get subtitles as I listen?

That is what live translation does: it handles the sound playing right now. See the page on live translating videos and streams. File mode handles audio files that already exist.

Coming soon to the Mac App StoremacOS 26 or later, Apple silicon. For early access, write to twinsub@gatheon.com.

Updated 2026-10-12. Back to the Twinsubs home page