NaturalLanguage + Translation
Analyze natural language text for tokenization, part-of-speech tagging, named entity recognition, sentiment analysis, language identification, and word/sentence embeddings. Translate text between languages with the Translation framework.
This skill covers two related frameworks: NaturalLanguage (
NLTokenizer,NLTagger,NLEmbedding) for on-device text analysis, and Translation (TranslationSession,LanguageAvailability) for language translation.
Scope boundary: Use this skill after you already have text. It owns
tokenization, language identification, POS/NER tagging, sentiment, embeddings,
custom NLModel classifiers/taggers, and in-app translation. Hand off OCR to
vision-framework, speech-to-text to speech-recognition, UI strings and
locale formatting to ios-localization, and generative summarization or Apple
Intelligence workflows to apple-on-device-ai.
Contents
- Setup
- Tokenization
- Language Identification
- Part-of-Speech Tagging
- Named Entity Recognition
- Sentiment Analysis
- Text Embeddings
- Translation
- Common Mistakes
- Review Checklist
- References
Setup
Import NaturalLanguage for text analysis and Translation for language
translation. No special entitlements or capabilities are required for
NaturalLanguage. Translation has split availability: system translation
presentation is iOS 17.4+ / macOS 14.4+, while TranslationSession,
.translationTask(), LanguageAvailability, and batch translation require
iOS 18+ / macOS 15+.
Direct TranslationSession(installedSource:target:) is the non-UI option, but
only when the source and target languages are already installed on device.
NaturalLanguage classes (NLTokenizer, NLTagger) are not thread-safe.
Use each instance from one thread or dispatch queue at a time.
Tokenization
Segment text into words, sentences, or paragraphs with NLTokenizer.
Token Units
Enumerating with Attributes
Use enumerateTokens(in:using:) to detect numeric or emoji tokens.
Language Identification
Detect the dominant language of a string with NLLanguageRecognizer.
Constrain the recognizer to expected languages for better accuracy on short text.
Part-of-Speech Tagging
Identify nouns, verbs, adjectives, and other lexical classes with NLTagger.
Common Tag Schemes
Named Entity Recognition
Extract people, places, and organizations.
Sentiment Analysis
Score text sentiment from -1.0 (negative) to +1.0 (positive).
Text Embeddings
Measure semantic similarity between words or sentences with NLEmbedding.
Sentence embeddings compare entire sentences.
Translation
System Translation Overlay
Show the built-in translation UI with .translationPresentation().
Programmatic Translation
Use .translationTask() for programmatic translations within a view context.
Batch Translation
Translate multiple strings in a single session.
Checking Language Availability
Common Mistakes
DON'T: Share NLTagger/NLTokenizer across threads
These classes are not thread-safe and will produce incorrect results or crash.
DON'T: Confuse NaturalLanguage with Core ML
NaturalLanguage provides built-in linguistic analysis. Use Core ML for custom
trained models. They complement each other via NLModel.
DON'T: Assume embeddings exist for all languages
Not all languages have word or sentence embeddings available on device.
DON'T: Create a new tagger per token
Creating and configuring a tagger is expensive. Reuse it for the same text.
DON'T: Ignore language hints for short text
Language detection on short strings (under ~20 characters) is unreliable. Set constraints or hints to improve accuracy.
Review Checklist
-
NLTokenizerandNLTaggerinstances used from a single thread - Tagger created once per text, not per token
- Language detection uses constraints/hints for short text
-
NLEmbeddingavailability checked before use (returns nil if unavailable) - Translation
LanguageAvailabilitychecked before attempting translation -
.translationTask()used within a SwiftUI view hierarchy - Batch translation uses
clientIdentifierto match responses to requests - Sentiment scores handled as optional (may return nil for unsupported languages)
-
.joinNamesoption used with NER to keep multi-word names together - Custom ML models loaded via
NLModel, not raw Core ML
References
- Extended patterns (custom models, contextual embeddings, gazetteers): references/translation-patterns.md [blocked]
- Natural Language framework
- NLTokenizer
- NLTagger
- NLEmbedding
- NLLanguageRecognizer
- Translation framework
- TranslationSession
- TranslationSession.Strategy
- LanguageAvailability


