Speech Recognition
Transcribe live and pre-recorded audio to text using Apple's Speech framework.
Covers SpeechAnalyzer / SpeechTranscriber (iOS 26+) and
SFSpeechRecognizer (iOS 10+) fallback guidance.
Scope boundary: Use this skill for speech-to-text recognition, speech
authorization, microphone capture plumbing, and result handling. Hand off text
analysis, language identification after transcription, sentiment, embeddings,
and translation to natural-language; hand off audio playback UI to avkit;
hand off summarization or generation over transcripts to apple-on-device-ai.
Contents
- SpeechAnalyzer Strategy (iOS 26+)
- SFSpeechRecognizer Setup
- Authorization
- Live Microphone Transcription
- Pre-Recorded Audio File Recognition
- On-Device vs Server Recognition
- Handling Results
- Common Mistakes
- Review Checklist
- References
SpeechAnalyzer Strategy (iOS 26+)
Use SpeechAnalyzer for modern iOS 26+ speech analysis, especially long-form
recordings, live transcription, time-indexed transcripts, and fully on-device
flows. Keep SFSpeechRecognizer for iOS 10+ deployment targets, server-backed
locale coverage, or existing callback/delegate implementations.
Read SpeechAnalyzer patterns [blocked] when implementing an iOS 26+ transcription pipeline, model asset handling, volatile results, or file/buffer examples.
SpeechAnalyzer setup checklist
- Choose the module:
SpeechTranscriberfor the newer general-purpose on-device model.DictationTranscriberwhenSpeechTranscriberis unavailable for the current device or locale and dictation-compatible support is acceptable.SpeechDetectoronly in conjunction with a transcriber when voice activity detection is worth the accuracy/power tradeoff.
- Check support before creating the session:
SpeechTranscriber.isAvailableSpeechTranscriber.supportedLocale(equivalentTo:)SpeechTranscriber.installedLocales/supportedLocaleswhen showing language choices.
- Pick a documented preset:
.transcriptionfor basic accurate transcription..progressiveTranscriptionfor live UI updates..timeIndexedProgressiveTranscriptionwhen playback highlighting needsaudioTimeRange.
- Install required assets with
AssetInventory.assetInstallationRequest. - Convert live audio buffers to
SpeechAnalyzer.bestAvailableAudioFormat(compatibleWith:)before yieldingAnalyzerInput. - Consume module results from their
AsyncSequencein a separate task. - Finish explicitly with
finalizeAndFinish(through:),finalizeAndFinishThroughEndOfInput(), orcancelAndFinishNow().
Do not use an offlineTranscription preset; Apple does not document one.
Finishing an AsyncStream input sequence does not finish the analyzer session.
SFSpeechRecognizer Setup
Creating a recognizer with locale
Monitoring availability changes
Authorization
Request both speech recognition and microphone permissions before starting
live transcription. Add these keys to Info.plist:
NSSpeechRecognitionUsageDescriptionNSMicrophoneUsageDescription
Live Microphone Transcription
The standard pattern: AVAudioEngine captures microphone audio → buffers are
appended to SFSpeechAudioBufferRecognitionRequest → results stream in.
Pre-Recorded Audio File Recognition
Use SFSpeechURLRecognitionRequest for audio files on disk:
On-Device vs Server Recognition
SFSpeechRecognizer can use on-device recognition for supported locales on
iOS 13+. If supportsOnDeviceRecognition is false, the recognizer requires a
network connection. requiresOnDeviceRecognition only has effect when the
recognizer supports it.
SFSpeechRecognizer requests may still be a poor fit for long-form capture.
Apple documents a roughly one-minute task limit for speech recognition and
other service limits. For long recordings on iOS 26+, prefer SpeechAnalyzer;
otherwise chunk or restart recognition before the limit and preserve transcript
state across tasks.
Handling Results
Partial vs final results
With shouldReportPartialResults, replace the displayed partial transcript until SFSpeechRecognitionResult.isFinal commits it. This is separate from SpeechTranscriber.Result.isFinal, whose volatile attributed range must be replaced by the final result for that range. The live example and references/speechanalyzer-patterns.md [blocked] contain the canonical loops.
Accessing alternative transcriptions and confidence
Adding punctuation (iOS 16+)
Contextual strings
Improve recognition of domain-specific terms:
Common Mistakes
Load references/speechanalyzer-patterns.md [blocked] for complete analyzer finalization and volatile-result code.
Review Checklist
-
NSSpeechRecognitionUsageDescriptionis in Info.plist -
NSMicrophoneUsageDescriptionis in Info.plist (if using live audio) - Authorization is requested before starting recognition
-
SFSpeechRecognizerDelegateis set to handleavailabilityDidChange - Audio engine is stopped and tap removed when recognition ends
-
recognitionRequest.endAudio()is called when done recording - Previous
recognitionTaskis canceled before starting a new one -
supportsOnDeviceRecognitionis checked before requiring on-device mode - Partial results are handled separately from final (
isFinal) results -
SFSpeechRecognizerone-minute/service limits are accounted for - For iOS 26+:
AssetInventoryassets are installed before usingSpeechAnalyzer - For iOS 26+:
SpeechTranscriber.isAvailableand locale support are checked - For iOS 26+: live buffers are converted to the analyzer-compatible format
- For iOS 26+: analyzer sessions are explicitly finalized or canceled
- For iOS 26+: volatile results are replaced by finalized results, not duplicated
References
- Speech framework
- SpeechAnalyzer
- SpeechTranscriber
- SpeechTranscriber.Preset
- DictationTranscriber
- SpeechDetector
- SFSpeechRecognizer
- SFSpeechAudioBufferRecognitionRequest
- SFSpeechURLRecognitionRequest
- SFSpeechRecognitionResult
- SFSpeechRecognitionRequest
- AssetInventory
- Asking Permission to Use Speech Recognition
- Recognizing Speech in Live Audio
- Bring advanced speech-to-text to your app with SpeechAnalyzer
