AI Tech News HubDaily Updates
Developer ToolsSeptember 27, 2026

How to Use OpenAI Whisper: From Installation to Subtitle Output, All in One Guide

A
AI 觀察家
Columnist · 2747 words
How to Use OpenAI Whisper: From Installation to Subtitle Output, All in One Guide

You may have been using ChatGPT or the API for a while, but Whisper is a tool many people have never seriously explored. It's an open-source speech-to-text model released by OpenAI in late 2022, and heading into 2026 it remains one of the most balanced options in terms of accuracy, multilingual support, and cost. In plain terms: you feed it an audio file, it spits out a transcript or subtitle file — and it handles Chinese, Japanese, Taiwanese Hokkien mixed with English without breaking a sweat.

This article has one clear goal: get you to your first transcript from scratch, then show you how to plug it into real-world workflows — whether that's YouTube subtitles, Zoom meeting notes, or podcast post-production.


What You Need to Prepare

Hardware: An NVIDIA GPU will speed things up significantly, but it isn't required. CPU mode works fine, just expect slower processing on longer audio. M1/M2/M3 Macs can also leverage MPS acceleration.

Software prerequisites:

  • Python 3.9 or above (3.11 recommended)
  • ffmpeg: Whisper uses it under the hood for audio decoding — without it, you'll hit an immediate error
  • pip or conda for environment management

The easiest way to install ffmpeg:

  • Mac: brew install ffmpeg
  • Ubuntu/Debian: sudo apt install ffmpeg
  • Windows: download from ffmpeg.org or use winget

Step 1: Create a Virtual Environment and Install Whisper

Don't install directly into your system Python. Do this enough times and you'll understand why.

python -m venv whisper-env
source whisper-env/bin/activate  # Windows: whisper-env\Scripts\activate
pip install openai-whisper

This pulls in Whisper along with its PyTorch dependency — it's not a small download, so the first install will take a moment.


Step 2: Choose a Model Size — Start with small

Whisper comes in five model tiers: tiny, base, small, medium, and large (currently at large-v3).

Model VRAM Required Speed Best For
tiny ~1 GB Fastest Quick drafts
base ~1 GB Fast Primarily English
small ~2 GB Medium Recommended for beginners
medium ~5 GB Slow Mixed multilingual content
large-v3 ~10 GB Slowest High accuracy demands

For your first run, small is the right call — accurate enough without an agonising wait. If you're running on CPU, tiny or base will be considerably less painful.


Step 3: Run Transcription and Output Your Desired Format

The most basic command:

whisper audio.mp3 --model small

By default, Whisper outputs all five formats at once: .txt, .vtt, .srt, .tsv, and .json. If you only want SRT subtitles:

whisper audio.mp3 --model small --output_format srt

For Chinese audio, specify the language explicitly — otherwise it runs language detection first and wastes time:

whisper audio.mp3 --model small --language zh

Output files are saved to the current directory, with the same filename as the audio file and the extension swapped out.


Step 4: Use the Python API to Integrate Into Your Workflow

If you need more than a one-off transcription and want to plug Whisper into an automated pipeline, calling it directly from Python gives you far more flexibility:

import whisper

model = whisper.load_model("small")
result = model.transcribe("audio.mp3", language="zh")

print(result["text"])  # Full transcript
for segment in result["segments"]:
    print(f"[{segment['start']:.1f}s] {segment['text']}")

Each entry in result["segments"] carries a start time, end time, and text — ready to be used for subtitle generation, keyword search, or feeding into an LLM for summarisation.

If you're already using other OpenAI services, the OpenAI API Practical Notes article covers the pricing structure in detail — Whisper also has a cloud API version you can call directly without running anything locally.


Common Errors and How to Avoid Them

"No module named 'whisper'": You forgot to activate your virtual environment, or you installed whisper instead of openai-whisper — these are two different packages.

"ffmpeg not found": ffmpeg is installed but not added to PATH. Installing via Homebrew on Mac usually handles this automatically; on Windows you'll need to set the environment variable manually.

Chinese transcript full of errors: Try upgrading to medium or large-v3. The small model has limited support for Chinese dialects or heavily accented audio. Also make sure you're passing --language zh to skip the language detection step.

Running very slowly: If you have an NVIDIA GPU but it's not being used, check whether your PyTorch build has CUDA support: python -c "import torch; print(torch.cuda.is_available())". If it prints False, PyTorch can't see your GPU.


Advanced Techniques: Timestamps, Audio Segmentation, and Tool Integration

Word-level timestamps: Whisper supports outputting timestamps at the word level. The third-party package whisper-timestamped makes this straightforward and is well-suited for karaoke-style subtitles or precise editing work.

Segmenting long audio: For files longer than 30 minutes, it's worth splitting them first with pydub and processing them in batches. This avoids memory issues and gives you better control over processing speed.

Integrating into automated pipelines: Feeding Whisper transcripts into an LLM for summarisation is a very common pattern. If you're wondering how to give an LLM access to external knowledge rather than relying on what it has memorised, What Is RAG? walks through how to turn meeting notes into a queryable knowledge base — a genuinely practical application.

Batch processing multiple files: A simple Python loop handles this cleanly:

import whisper, os
model = whisper.load_model("small")
for f in os.listdir("./audio_files"):
    if f.endswith(".mp3"):
        result = model.transcribe(f"./audio_files/{f}", language="zh")
        with open(f"{f}.txt", "w") as out:
            out.write(result["text"])

After Your First Run, Check These Points

Once you have your first transcript, verify a few things:

  • Is the SRT timeline offset? (Common when audio starts with a long silence)
  • Are proper nouns or names misrecognised? (Whisper is inherently weak here — post-processing with find-and-replace works)
  • If you're uploading to YouTube, .vtt format is natively supported and requires no conversion

As a next step, consider connecting your transcripts to the RAG architecture described in Fine-tuning vs RAG, turning your meeting records into a queryable knowledge base — this combination is already a well-established workflow heading into 2026.

Frequently Asked Questions

Do you need a GPU to run Whisper?

Not necessarily. Whisper runs on CPU as well, though noticeably slower — the difference becomes very apparent with longer audio files. If you have an NVIDIA GPU with a CUDA-enabled build of PyTorch, performance improves dramatically. M1/M2/M3 Macs can use MPS acceleration, which is also meaningfully faster than pure CPU.

What's the difference between local Whisper and OpenAI's cloud speech API?

Running Whisper locally uses the open-source version — it's free but requires your own hardware. OpenAI also offers a cloud-based Whisper API (/v1/audio/transcriptions), billed per minute, with no local environment to manage. It's the right choice if you'd rather not self-host, though costs can accumulate with heavy usage over time.

How accurate is Whisper for Chinese?

With the medium or large-v3 model and the --language zh parameter, accuracy for Mandarin is quite strong. Taiwanese Hokkien, Cantonese, or heavily accented speech will see more noticeable errors — for those use cases, a larger model or post-processing correction is currently the more reliable path.

What should I do if the SRT subtitle timestamps are out of sync?

The most common cause is a long silence at the start of the audio file, which can confuse Whisper's voice activity detection. The fix is to trim the silent intro with ffmpeg first, or pass the --condition_on_previous_text False parameter to reduce accumulated drift.

Can Whisper process video files directly?

Yes. Since Whisper decodes audio through ffmpeg under the hood, video formats like .mp4, .mov, and .mkv can be passed in directly — no need to manually extract the audio track first. The command is identical to processing an audio file; just swap in the video filename.

Share

Related articles