When you’re producing a podcast or interview series, turning audio into clean text is a recurring chore. Show notes, social posts, newsletters, even course content all start from a solid transcript. In this tutorial, we’ll walk through setting up a simple macOS workflow that lets you batch-transcribe episodes using the terminal and OpenAI’s Whisper.

Instead of uploading files to different web tools every time, you’ll have a repeatable process: drop your files in a folder, run a single command, and get timestamped transcripts you can quickly clean and label.

The Challenge

There are many ways to get transcripts:

  1. You can subscribe to SaaS tools that transcribe and diarize your audio automatically. They’re polished and convenient, but they add ongoing cost and often lock your data into their platform.
  2. Some DAWs and recording tools include built-in transcription, but formats differ and it’s hard to keep one consistent workflow as your show grows.

In this article, we’ll set up a local, scriptable pipeline on macOS using Whisper. The goal is simple: produce timestamped text files for each episode that are easy to edit and annotate with speaker names.

Creating the workflow: step-by-step guide to batch transcription

We’ll use:

  • Whisper’s command-line interface for high-quality transcription.
  • A small Python script to convert caption files into [hh:mm:ss] text.
  • A batch script to process multiple episodes in one go.

Everything runs locally on your Mac inside a Python virtual environment, so it’s isolated, reproducible, and under your control.

Setting up the tools

First, make sure you have Homebrew installed; it’s the de facto package manager on macOS. Then install ffmpeg, which Whisper uses to handle audio formats, and set up a virtual environment for Python packages.

Open Terminal and run:

Terminal
brew install ffmpeg
cd ~/Podcast/Audio
python3 -m venv .venv
source .venv/bin/activate
pip install -U pip
pip install -U openai-whisper srt

Swap ~/Podcast/Audio for the folder that holds your episodes. If the path has spaces, quote everything after the tilde, for example cd ~/"My Podcast/Audio": a tilde inside quotes isn’t expanded.

The .venv folder keeps your transcription tools separate from the rest of your system. Whenever you want to work on transcripts, activate it with source .venv/bin/activate before running commands.

Transcribing a single episode with Whisper

With the virtual environment active and your .wav or .mp3 file in the folder, you can transcribe one episode like this:

Terminal
whisper "Episode 001 - Guest Name.wav" \
  --model medium \
  --language en \
  --task transcribe \
  --output_format srt

This command:

  • Uses the medium Whisper model for a good balance of accuracy and speed.
  • Assumes English audio with --language en.
  • Produces an .srt subtitle file containing text plus timestamps.

After it finishes, you should see a new file next to your audio:

Output
Episode 001 - Guest Name.srt

This caption file is perfect for video, but a bit clunky to read as plain notes. Next, you’ll convert it into a more podcast-friendly transcript.

Turning subtitles into a readable transcript

To make your transcripts easier to review and annotate, you can flatten the .srt file into simple lines that look like:

Episode 001 - Guest Name.full.txt
[00:00:00] Welcome back to the show…
[00:00:05] Today we’re talking about…

Create a helper script named srt_to_full_text.py in the same folder as your audio:

Terminal
cat > srt_to_full_text.py

Paste this in, then press Enter and Ctrl+D to save it. The scripts in this guide need Python 3.10 or later; check with python3 --version.

srt_to_full_text.py
# srt_to_full_text.py
import sys
import datetime
import srt  # installed earlier with pip

def fmt_time(t: datetime.timedelta) -> str:
    total = int(t.total_seconds())
    h = total // 3600
    m = (total % 3600) // 60
    s = total % 60
    return f"[{h:02d}:{m:02d}:{s:02d}]"

def main(path: str) -> None:
    with open(path, "r", encoding="utf-8") as f:
        subs = list(srt.parse(f.read()))

    for sub in subs:
        text = sub.content.replace("\n", " ").strip()
        if not text:
            continue
        print(f"{fmt_time(sub.start)} {text}")

if __name__ == "__main__":
    if len(sys.argv) != 2:
        print("Usage: python srt_to_full_text.py <file.srt>")
        sys.exit(1)
    main(sys.argv[1])

Now run:

Terminal
python srt_to_full_text.py "Episode 001 - Guest Name.srt" \
  > "Episode 001 - Guest Name.full.txt"

Open the .full.txt file in your editor. You’ll see each segment on its own line with a neatly formatted timestamp. This is the file you’ll use to:

  • Prefix lines with Host: or your guest’s name.
  • Edit for clarity and grammar.
  • Pull quotes and show notes.

Batch-processing multiple episodes

Once the single-file process works, you can scale it to an entire season with a small batch script. Create a file named batch_transcribe.py alongside your audio:

Terminal
cat > batch_transcribe.py

Paste the script below, then press Enter and Ctrl+D. Replace the names in EPISODES with your own files, exactly as they appear in Finder.

batch_transcribe.py
# batch_transcribe.py
#
# Batch:
#  1) Run Whisper with --model medium and output_format srt
#  2) Convert SRT -> [hh:mm:ss] text for manual speaker tagging

import subprocess
import sys
from pathlib import Path
import datetime
import srt

EPISODES = [
    "Episode 001 - Guest Name.wav",
    "Episode 002 - Guest Name.wav",
    "Episode 003 - Guest Name Part 1.wav",
    "Episode 003 - Guest Name Part 2.wav",
]

def fmt_time(t: datetime.timedelta) -> str:
    total = int(t.total_seconds())
    h = total // 3600
    m = (total % 3600) // 60
    s = total % 60
    return f"[{h:02d}:{m:02d}:{s:02d}]"

def srt_to_full_text(srt_path: Path, out_path: Path) -> None:
    with srt_path.open("r", encoding="utf-8") as f:
        subs = list(srt.parse(f.read()))
    with out_path.open("w", encoding="utf-8") as out:
        for sub in subs:
            text = sub.content.replace("\n", " ").strip()
            if not text:
                continue
            out.write(f"{fmt_time(sub.start)} {text}\n")

def run_whisper(audio_path: Path) -> Path | None:
    srt_path = audio_path.with_suffix(".srt")
    cmd = [
        "whisper",
        str(audio_path),
        "--model", "medium",
        "--language", "en",
        "--task", "transcribe",
        "--output_format", "srt",
    ]
    print(f"\n=== Transcribing: {audio_path.name} ===")
    try:
        subprocess.run(cmd, check=True)
    except subprocess.CalledProcessError as e:
        print(f"Whisper failed for {audio_path.name}: {e}", file=sys.stderr)
        return None
    if not srt_path.exists():
        print(f"SRT not found for {audio_path.name}", file=sys.stderr)
        return None
    return srt_path

def main():
    base_dir = Path(".")
    for name in EPISODES:
        audio_path = base_dir / name
        if not audio_path.exists():
            print(f"Skipping (missing file): {name}", file=sys.stderr)
            continue

        srt_path = run_whisper(audio_path)
        if srt_path is None:
            continue

        out_txt = audio_path.with_suffix(".full.txt")
        print(f"Converting {srt_path.name} -> {out_txt.name}")
        srt_to_full_text(srt_path, out_txt)

    print("\nDone. You can now open each *.full.txt and add speaker names manually.")

if __name__ == "__main__":
    main()

Run it with:

Terminal
python batch_transcribe.py

This will:

  • Transcribe each listed episode with Whisper medium.
  • Generate .srt files with timestamps.
  • Convert each .srt into a .full.txt file that’s easy to read and annotate.

Missing files are skipped with a message, and one failed episode doesn’t stop the rest of the batch.

Using your transcripts in your workflow

At this point, every episode has:

  • A caption file you can attach to video or audio players.
  • A plain text transcript with timestamps you can:
    • Edit and polish.
    • Label with speaker names.
    • Mine for quotes, summaries, and social snippets.

Conclusion

By keeping everything in your local project folder, you’ve created a repeatable transcription pipeline that scales with your show and stays under your control.