Azure Speech To Text

calesthio/OpenMontage/.claude/skills/azure-speech-to-text

作者 calesthio9327439db69021ab4b0e2776729bf3b58fdb5a87MIT65K 個星標收錄於 2026年10月9日更新於 2026年10月9日儲存庫5 天前更新

Transcribe audio to text using Azure AI Speech (Fast Transcription REST API). Use when converting audio/video to text, generating subtitles, or processing spoken content in OpenMontage. Optional cloud STT provider — preferred when AZURE_SPEECH_KEY is configured; the local faster-whisper `transcriber` is the default offline path.

僅含說明AI & Agents
AI 產生的概覽

使用 Azure AI Speech 快速轉錄將音訊轉為文字,產生帶時間戳的片段與逐字時間戳。

功能
引導代理使用 azure_stt 工具,透過同步 multipart 請求把本機音訊檔案送往 Azure AI Speech 快速轉錄。傳回包含起訖時間、選用講者標籤、逐字時間戳、辨識語言與時長的轉錄片段,並對應為 OpenMontage 轉錄結構。也說明語言、候選語言清單、講者分離、講者上限與不雅詞過濾等參數,以及退回本機 faster-whisper 轉錄器的方式。
適用情境
適用於把音訊或視訊中的語音轉成文字、產生字幕,或在 OpenMontage 中處理口語內容。設定 Azure Speech 金鑰時,它是偏好的雲端語音轉文字方式;本機轉錄器仍是離線預設路徑。
執行需求
需要網際網路連線,以及一個 Azure AI Speech 資源,提供 AZURE_SPEECH_KEY 和 AZURE_SPEECH_REGION 或 AZURE_SPEECH_ENDPOINT。依賴 requests 套件與 OpenMontage 的 azure_stt 工具;未附帶指令碼。

Azure AI Speech — Speech-to-Text

Transcribe audio to text with Azure Fast Transcription — synchronous, word-level timestamps, speaker diarization, and multi-language identification. In OpenMontage this is exposed through the azure_stt tool (capability=analysis, provider=azure). It is an optional cloud STT provider — when AZURE_SPEECH_KEY is configured, prefer it for cloud transcription. The local transcriber tool (faster-whisper) remains the default offline path and the fallback when Azure is unavailable.

Docs: Fast Transcription · Speech service overview

Why Fast Transcription (not Batch)

Azure exposes three STT surfaces. OpenMontage uses Fast Transcription because the pipeline transcribes local audio files:

SurfaceInputLatencyNeeds
Fast Transcription (used here)local file, multipart POSTsynchronous, sub-real-timekey + region
Batch Transcriptionaudio at a URL (Blob + SAS)async job + pollingBlob storage plumbing
Speech SDK (spx)mic / stream / filestreamingnative azure-cognitiveservices-speech package

Fast Transcription needs no Blob storage, no SAS URLs, and no native SDK — just requests and the two env vars.

Setup

Create a Speech resource in the Azure portal; copy the key and region from its Keys and Endpoint page.

bash
export AZURE_SPEECH_KEY=your_speech_resource_keyexport AZURE_SPEECH_REGION=eastus          # your resource's region# export AZURE_SPEECH_ENDPOINT=https://...  # optional: overrides region

azure_stt reports AVAILABLE once AZURE_SPEECH_KEY plus either AZURE_SPEECH_REGION or AZURE_SPEECH_ENDPOINT are set.

Using it in a pipeline

Prefer azure_stt over transcriber unless the run must be offline. Its output matches the transcriber schema exactly, so it is a drop-in for subtitle_gen and any stage that consumes a transcript.

python
from tools.tool_registry import registryregistry.discover()stt = registry._tools["azure_stt"]
result = stt.execute({    "input_path": "projects/my-video/assets/audio/narration.mp3",    # "language": "en",          # ISO 639-1 or BCP-47 ("en-US"); omit for auto-ID    # "diarize": True,           # speaker labels, no HuggingFace token needed    # "max_speakers": 4,    "output_dir": "projects/my-video/artifacts",})if result.success:    segs = result.data["segments"]          # [{id,start,end,text,words:[...]}]    words = result.data["word_timestamps"]  # flat [{word,start,end,probability}]

If azure_stt is unavailable (no key) or errors, fall back to transcriber (local whisper) — its execute signature and output are identical.

Parameters that matter

  • language — pass an ISO code ("en") or a full locale ("en-US"). Pin it when you know the language; it is faster and more accurate than auto-ID.
  • candidate_locales — when language is omitted, Azure runs language identification across this shortlist. Narrow it to the languages you actually expect; a huge list slows detection and invites misclassification.
  • diarize / max_speakers — enable for multi-speaker audio (interviews, podcasts). Set max_speakers to the real upper bound.
  • profanity_filter — None | Masked (default) | Removed | Tags.

Response shape (mapped to the transcriber schema)

The raw Azure response (phrases[] with offsetMilliseconds / words[]) is converted to seconds and the OpenMontage transcript schema:

json
{  "segments": [    {"id": 0, "start": 0.0, "end": 2.4, "text": "Hello world",     "speaker": 1,     "words": [{"word": "Hello", "start": 0.0, "end": 0.5, "probability": 0.98}]}  ],  "word_timestamps": [{"word": "Hello", "start": 0.0, "end": 0.5, "probability": 0.98}],  "language": "en-US",  "duration_seconds": 2.4,  "provider": "azure"}

Note: Fast Transcription has no per-word confidence, so each word carries the phrase confidence in probability.

Limits & tips

  • Single file up to ~2 hours / a few hundred MB per request. For longer or bulk jobs, use Azure Batch Transcription instead.
  • Send clean audio (16 kHz+ mono is plenty). Transcode video to audio first if you only need speech — smaller upload, same result.
  • Verify timing: word timestamps drive subtitle cues in subtitle_gen. Spot-check the first and last cues against the source audio.

來源與署名

來源:calesthio/OpenMontage位於.claude/skills/azure-speech-to-text提交9327439

授權條款: MIT

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架