Azure AI Speech — Speech-to-Text
Transcribe audio to text with Azure Fast Transcription — synchronous,
word-level timestamps, speaker diarization, and multi-language identification.
In OpenMontage this is exposed through the azure_stt tool (capability=analysis,
provider=azure). It is an optional cloud STT provider — when
AZURE_SPEECH_KEY is configured, prefer it for cloud transcription. The local
transcriber tool (faster-whisper) remains the default offline path and the
fallback when Azure is unavailable.
Why Fast Transcription (not Batch)
Azure exposes three STT surfaces. OpenMontage uses Fast Transcription because the pipeline transcribes local audio files:
Fast Transcription needs no Blob storage, no SAS URLs, and no native SDK — just
requests and the two env vars.
Setup
Create a Speech resource in the Azure portal; copy the key and region from its Keys and Endpoint page.
azure_stt reports AVAILABLE once AZURE_SPEECH_KEY plus either
AZURE_SPEECH_REGION or AZURE_SPEECH_ENDPOINT are set.
Using it in a pipeline
Prefer azure_stt over transcriber unless the run must be offline. Its output
matches the transcriber schema exactly, so it is a drop-in for subtitle_gen
and any stage that consumes a transcript.
If azure_stt is unavailable (no key) or errors, fall back to transcriber
(local whisper) — its execute signature and output are identical.
Parameters that matter
language— pass an ISO code ("en") or a full locale ("en-US"). Pin it when you know the language; it is faster and more accurate than auto-ID.candidate_locales— whenlanguageis omitted, Azure runs language identification across this shortlist. Narrow it to the languages you actually expect; a huge list slows detection and invites misclassification.diarize/max_speakers— enable for multi-speaker audio (interviews, podcasts). Setmax_speakersto the real upper bound.profanity_filter—None|Masked(default) |Removed|Tags.
Response shape (mapped to the transcriber schema)
The raw Azure response (phrases[] with offsetMilliseconds / words[]) is
converted to seconds and the OpenMontage transcript schema:
Note: Fast Transcription has no per-word confidence, so each word carries the
phrase confidence in probability.
Limits & tips
- Single file up to ~2 hours / a few hundred MB per request. For longer or bulk jobs, use Azure Batch Transcription instead.
- Send clean audio (16 kHz+ mono is plenty). Transcode video to audio first if you only need speech — smaller upload, same result.
- Verify timing: word timestamps drive subtitle cues in
subtitle_gen. Spot-check the first and last cues against the source audio.


