Whisper

orchestra-research/ai-research-skills/18-multimodal/whisper

by orchestra-research773a52944ba4MIT13K starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 3 months ago

OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.

Instructions onlyAI & Agents
AI-generated overview

Guides use of OpenAI's Whisper speech recognition model for transcription, translation to English, and language identification.

What it does
This skill documents how to run OpenAI's Whisper speech recognition model for speech-to-text, translation to English, and language identification across 99 languages. It covers model sizes, transcription options such as language selection, initial prompts, word-level timestamps, and temperature fallback, plus command-line usage and output formats like txt, srt, vtt, and json. It also describes batch processing, GPU acceleration, subtitle generation, and integration with tools such as LangChain.
When to use it
Use it for speech-to-text transcription, podcast or video transcription, meeting notes automation, translation of audio to English, or multilingual and noisy audio processing. It is intended when robust multilingual ASR is needed rather than real-time streaming or speaker diarization.
Requirements
Requires Python 3.8-3.11, the openai-whisper package, and ffmpeg; the document also mentions transformers and torch as dependencies and faster-whisper for streaming. GPU acceleration is optional. It ships no scripts, only instructions and a language reference file.

Whisper - Robust Speech Recognition

OpenAI's multilingual speech recognition model.

When to use Whisper

Use when:

  • Speech-to-text transcription (99 languages)
  • Podcast/video transcription
  • Meeting notes automation
  • Translation to English
  • Noisy audio transcription
  • Multilingual audio processing

Metrics:

  • 72,900+ GitHub stars
  • 99 languages supported
  • Trained on 680,000 hours of audio
  • MIT License

Use alternatives instead:

  • AssemblyAI: Managed API, speaker diarization
  • Deepgram: Real-time streaming ASR
  • Google Speech-to-Text: Cloud-based

Quick start

Installation

bash
# Requires Python 3.8-3.11pip install -U openai-whisper
# Requires ffmpeg# macOS: brew install ffmpeg# Ubuntu: sudo apt install ffmpeg# Windows: choco install ffmpeg

Basic transcription

python
import whisper
# Load modelmodel = whisper.load_model("base")
# Transcriberesult = model.transcribe("audio.mp3")
# Print textprint(result["text"])
# Access segmentsfor segment in result["segments"]:    print(f"[{segment['start']:.2f}s - {segment['end']:.2f}s] {segment['text']}")

Model sizes

python
# Available modelsmodels = ["tiny", "base", "small", "medium", "large", "turbo"]
# Load specific modelmodel = whisper.load_model("turbo")  # Fastest, good quality
ModelParametersEnglish-onlyMultilingualSpeedVRAM
tiny39M✓✓~32x~1 GB
base74M✓✓~16x~1 GB
small244M✓✓~6x~2 GB
medium769M✓✓~2x~5 GB
large1550M✗✓1x~10 GB
turbo809M✗✓~8x~6 GB

Recommendation: Use turbo for best speed/quality, base for prototyping

Transcription options

Language specification

python
# Auto-detect languageresult = model.transcribe("audio.mp3")
# Specify language (faster)result = model.transcribe("audio.mp3", language="en")
# Supported: en, es, fr, de, it, pt, ru, ja, ko, zh, and 89 more

Task selection

python
# Transcription (default)result = model.transcribe("audio.mp3", task="transcribe")
# Translation to Englishresult = model.transcribe("spanish.mp3", task="translate")# Input: Spanish audio → Output: English text

Initial prompt

python
# Improve accuracy with contextresult = model.transcribe(    "audio.mp3",    initial_prompt="This is a technical podcast about machine learning and AI.")
# Helps with:# - Technical terms# - Proper nouns# - Domain-specific vocabulary

Timestamps

python
# Word-level timestampsresult = model.transcribe("audio.mp3", word_timestamps=True)
for segment in result["segments"]:    for word in segment["words"]:        print(f"{word['word']} ({word['start']:.2f}s - {word['end']:.2f}s)")

Temperature fallback

python
# Retry with different temperatures if confidence lowresult = model.transcribe(    "audio.mp3",    temperature=(0.0, 0.2, 0.4, 0.6, 0.8, 1.0))

Command line usage

bash
# Basic transcriptionwhisper audio.mp3
# Specify modelwhisper audio.mp3 --model turbo
# Output formatswhisper audio.mp3 --output_format txt     # Plain textwhisper audio.mp3 --output_format srt     # Subtitleswhisper audio.mp3 --output_format vtt     # WebVTTwhisper audio.mp3 --output_format json    # JSON with timestamps
# Languagewhisper audio.mp3 --language Spanish
# Translationwhisper spanish.mp3 --task translate

Batch processing

python
import os
audio_files = ["file1.mp3", "file2.mp3", "file3.mp3"]
for audio_file in audio_files:    print(f"Transcribing {audio_file}...")    result = model.transcribe(audio_file)
    # Save to file    output_file = audio_file.replace(".mp3", ".txt")    with open(output_file, "w") as f:        f.write(result["text"])

Real-time transcription

python
# For streaming audio, use faster-whisper# pip install faster-whisper
from faster_whisper import WhisperModel
model = WhisperModel("base", device="cuda", compute_type="float16")
# Transcribe with streamingsegments, info = model.transcribe("audio.mp3", beam_size=5)
for segment in segments:    print(f"[{segment.start:.2f}s -> {segment.end:.2f}s] {segment.text}")

GPU acceleration

python
import whisper
# Automatically uses GPU if availablemodel = whisper.load_model("turbo")
# Force CPUmodel = whisper.load_model("turbo", device="cpu")
# Force GPUmodel = whisper.load_model("turbo", device="cuda")
# 10-20× faster on GPU

Integration with other tools

Subtitle generation

bash
# Generate SRT subtitleswhisper video.mp4 --output_format srt --language English
# Output: video.srt

With LangChain

python
from langchain.document_loaders import WhisperTranscriptionLoader
loader = WhisperTranscriptionLoader(file_path="audio.mp3")docs = loader.load()
# Use transcription in RAGfrom langchain_chroma import Chromafrom langchain_openai import OpenAIEmbeddings
vectorstore = Chroma.from_documents(docs, OpenAIEmbeddings())

Extract audio from video

bash
# Use ffmpeg to extract audioffmpeg -i video.mp4 -vn -acodec pcm_s16le audio.wav
# Then transcribewhisper audio.wav

Best practices

  1. Use turbo model - Best speed/quality for English
  2. Specify language - Faster than auto-detect
  3. Add initial prompt - Improves technical terms
  4. Use GPU - 10-20× faster
  5. Batch process - More efficient
  6. Convert to WAV - Better compatibility
  7. Split long audio - <30 min chunks
  8. Check language support - Quality varies by language
  9. Use faster-whisper - 4× faster than openai-whisper
  10. Monitor VRAM - Scale model size to hardware

Performance

ModelReal-time factor (CPU)Real-time factor (GPU)
tiny~0.32~0.01
base~0.16~0.01
turbo~0.08~0.01
large~1.0~0.05

Real-time factor: 0.1 = 10× faster than real-time

Language support

Top-supported languages:

  • English (en)
  • Spanish (es)
  • French (fr)
  • German (de)
  • Italian (it)
  • Portuguese (pt)
  • Russian (ru)
  • Japanese (ja)
  • Korean (ko)
  • Chinese (zh)

Full list: 99 languages total

Limitations

  1. Hallucinations - May repeat or invent text
  2. Long-form accuracy - Degrades on >30 min audio
  3. Speaker identification - No diarization
  4. Accents - Quality varies
  5. Background noise - Can affect accuracy
  6. Real-time latency - Not suitable for live captioning

Resources

Source and attribution

Source:orchestra-research/ai-research-skillsin18-multimodal/whisperat commit773a529

License: MIT

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal