Natural Language

by dpearson26998d90fd121a26No license1.1K starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 2 months ago

Tokenize, tag, and analyze natural language text using Apple's NaturalLanguage framework and translate between languages with the Translation framework. Use when adding language identification, sentiment analysis, named entity recognition, part-of-speech tagging, text embeddings, or in-app translation to iOS/macOS/visionOS apps.

Instructions onlySoftware Development
AI-generated overview

Guides iOS/macOS/visionOS developers in using Apple's NaturalLanguage and Translation frameworks for text analysis and translation.

What it does
This skill provides reference guidance and Swift code patterns for Apple's NaturalLanguage framework, covering tokenization, language identification, part-of-speech tagging, named entity recognition, sentiment analysis, and word/sentence embeddings. It also covers the Translation framework for in-app translation, including system translation overlays, programmatic and batch translation, and language availability checks. It includes common mistakes, a review checklist, and a reference file with extended patterns.
When to use it
Use this skill when adding on-device text analysis or in-app translation to iOS, macOS, or visionOS apps. It is intended for developers who already have text and need tokenization, tagging, sentiment, embeddings, or translation capabilities. It is not for OCR, speech-to-text, UI localization, or generative summarization.
Requirements
Requires Apple platform development (iOS/macOS/visionOS) with Swift and the NaturalLanguage and Translation frameworks. No special entitlements are needed for NaturalLanguage; Translation features require iOS 17.4+/macOS 14.4+ for system presentation and iOS 18+/macOS 15+ for programmatic sessions. Ships no scripts; instructions only, with a reference markdown file.

NaturalLanguage + Translation

Analyze natural language text for tokenization, part-of-speech tagging, named entity recognition, sentiment analysis, language identification, and word/sentence embeddings. Translate text between languages with the Translation framework.

This skill covers two related frameworks: NaturalLanguage (NLTokenizer, NLTagger, NLEmbedding) for on-device text analysis, and Translation (TranslationSession, LanguageAvailability) for language translation.

Scope boundary: Use this skill after you already have text. It owns tokenization, language identification, POS/NER tagging, sentiment, embeddings, custom NLModel classifiers/taggers, and in-app translation. Hand off OCR to vision-framework, speech-to-text to speech-recognition, UI strings and locale formatting to ios-localization, and generative summarization or Apple Intelligence workflows to apple-on-device-ai.

Contents

Setup

Import NaturalLanguage for text analysis and Translation for language translation. No special entitlements or capabilities are required for NaturalLanguage. Translation has split availability: system translation presentation is iOS 17.4+ / macOS 14.4+, while TranslationSession, .translationTask(), LanguageAvailability, and batch translation require iOS 18+ / macOS 15+. Direct TranslationSession(installedSource:target:) is the non-UI option, but only when the source and target languages are already installed on device.

swift
import NaturalLanguageimport Translation

NaturalLanguage classes (NLTokenizer, NLTagger) are not thread-safe. Use each instance from one thread or dispatch queue at a time.

Tokenization

Segment text into words, sentences, or paragraphs with NLTokenizer.

swift
import NaturalLanguage
func tokenizeWords(in text: String) -> [String] {    let tokenizer = NLTokenizer(unit: .word)    tokenizer.string = text
    let range = text.startIndex..<text.endIndex    return tokenizer.tokens(for: range).map { String(text[$0]) }}

Token Units

UnitDescription
.wordIndividual words
.sentenceSentences
.paragraphParagraphs
.documentEntire document

Enumerating with Attributes

Use enumerateTokens(in:using:) to detect numeric or emoji tokens.

swift
let tokenizer = NLTokenizer(unit: .word)tokenizer.string = text
tokenizer.enumerateTokens(in: text.startIndex..<text.endIndex) { range, attributes in    if attributes.contains(.numeric) {        print("Number: \(text[range])")    }    return true // continue enumeration}

Language Identification

Detect the dominant language of a string with NLLanguageRecognizer.

swift
func detectLanguage(for text: String) -> NLLanguage? {    NLLanguageRecognizer.dominantLanguage(for: text)}
// Multiple hypotheses with confidence scoresfunc languageHypotheses(for text: String, max: Int = 5) -> [NLLanguage: Double] {    let recognizer = NLLanguageRecognizer()    recognizer.processString(text)    return recognizer.languageHypotheses(withMaximum: max)}

Constrain the recognizer to expected languages for better accuracy on short text.

swift
let recognizer = NLLanguageRecognizer()recognizer.languageConstraints = [.english, .french, .spanish]recognizer.processString(text)let detected = recognizer.dominantLanguage

Part-of-Speech Tagging

Identify nouns, verbs, adjectives, and other lexical classes with NLTagger.

swift
func tagPartsOfSpeech(in text: String) -> [(String, NLTag)] {    let tagger = NLTagger(tagSchemes: [.lexicalClass])    tagger.string = text
    var results: [(String, NLTag)] = []    let range = text.startIndex..<text.endIndex    let options: NLTagger.Options = [.omitPunctuation, .omitWhitespace]
    tagger.enumerateTags(in: range, unit: .word, scheme: .lexicalClass, options: options) { tag, tokenRange in        if let tag {            results.append((String(text[tokenRange]), tag))        }        return true    }    return results}

Common Tag Schemes

SchemeOutput
.lexicalClassPart of speech (noun, verb, adjective)
.nameTypeNamed entity type (person, place, organization)
.nameTypeOrLexicalClassCombined NER + POS
.lemmaBase form of a word
.languagePer-token language
.sentimentScoreSentiment polarity score

Named Entity Recognition

Extract people, places, and organizations.

swift
func extractEntities(from text: String) -> [(String, NLTag)] {    let tagger = NLTagger(tagSchemes: [.nameType])    tagger.string = text
    var entities: [(String, NLTag)] = []    let options: NLTagger.Options = [.omitPunctuation, .omitWhitespace, .joinNames]
    tagger.enumerateTags(        in: text.startIndex..<text.endIndex,        unit: .word,        scheme: .nameType,        options: options    ) { tag, tokenRange in        if let tag, tag != .other {            entities.append((String(text[tokenRange]), tag))        }        return true    }    return entities}// NLTag values: .personalName, .placeName, .organizationName

Sentiment Analysis

Score text sentiment from -1.0 (negative) to +1.0 (positive).

swift
func sentimentScore(for text: String) -> Double? {    let tagger = NLTagger(tagSchemes: [.sentimentScore])    tagger.string = text
    let (tag, _) = tagger.tag(        at: text.startIndex,        unit: .paragraph,        scheme: .sentimentScore    )    return tag.flatMap { Double($0.rawValue) }}

Text Embeddings

Measure semantic similarity between words or sentences with NLEmbedding.

swift
func wordSimilarity(_ word1: String, _ word2: String) -> Double? {    guard let embedding = NLEmbedding.wordEmbedding(for: .english) else { return nil }    return embedding.distance(between: word1, and: word2, distanceType: .cosine)}
func findSimilarWords(to word: String, count: Int = 5) -> [(String, Double)] {    guard let embedding = NLEmbedding.wordEmbedding(for: .english) else { return [] }    return embedding.neighbors(for: word, maximumCount: count, distanceType: .cosine)}

Sentence embeddings compare entire sentences.

swift
func sentenceSimilarity(_ s1: String, _ s2: String) -> Double? {    guard let embedding = NLEmbedding.sentenceEmbedding(for: .english) else { return nil }    return embedding.distance(between: s1, and: s2, distanceType: .cosine)}

Translation

System Translation Overlay

Show the built-in translation UI with .translationPresentation().

swift
import SwiftUIimport Translation
struct TranslatableView: View {    @State private var showTranslation = false    let text = "Hello, how are you?"
    var body: some View {        Button { showTranslation = true } label: {            Text(text)        }        .buttonStyle(.plain)        .translationPresentation(            isPresented: $showTranslation,            text: text        )    }}

Programmatic Translation

Use .translationTask() for programmatic translations within a view context.

swift
struct TranslatingView: View {    @State private var translatedText = ""    @State private var translationErrorMessage: String?    @State private var configuration: TranslationSession.Configuration?
    var body: some View {        VStack {            Text(translatedText)            Button("Translate") {                configuration = .init(source: Locale.Language(identifier: "en"),                                      target: Locale.Language(identifier: "es"))            }        }        .translationTask(configuration) { session in            do {                let response = try await session.translate("Hello, world!")                await MainActor.run {                    translatedText = response.targetText                    translationErrorMessage = nil                }            } catch {                let message = error.localizedDescription                await MainActor.run {                    translationErrorMessage = message                }            }        }    }}

Batch Translation

Translate multiple strings in a single session.

swift
.translationTask(configuration) { session in    do {        let requests = texts.enumerated().map { index, text in            TranslationSession.Request(sourceText: text,                                       clientIdentifier: "\(index)")        }        let responses = try await session.translations(from: requests)        for response in responses {            print("\(response.sourceText) -> \(response.targetText)")        }    } catch {        // Handle cancellation, unsupported languages, or download refusal.    }}

Checking Language Availability

swift
let availability = LanguageAvailability()let status = await availability.status(    from: Locale.Language(identifier: "en"),    to: Locale.Language(identifier: "ja"))switch status {case .installed: break    // Ready to translate offlinecase .supported: break    // Needs downloadcase .unsupported: break  // Language pair not available}

Common Mistakes

DON'T: Share NLTagger/NLTokenizer across threads

These classes are not thread-safe and will produce incorrect results or crash.

swift
// WRONGlet sharedTagger = NLTagger(tagSchemes: [.lexicalClass])DispatchQueue.concurrentPerform(iterations: 10) { _ in    sharedTagger.string = someText  // Data race}
// CORRECTawait withTaskGroup(of: Void.self) { group in    for _ in 0..<10 {        group.addTask {            let tagger = NLTagger(tagSchemes: [.lexicalClass])            tagger.string = someText            // process...        }    }}

DON'T: Confuse NaturalLanguage with Core ML

NaturalLanguage provides built-in linguistic analysis. Use Core ML for custom trained models. They complement each other via NLModel.

swift
// WRONG: Trying to do NER with raw Core MLlet coreMLModel = try MLModel(contentsOf: modelURL)
// CORRECT: Use NLTagger for built-in NERlet tagger = NLTagger(tagSchemes: [.nameType])
// Or load a custom Core ML model via NLModellet nlModel = try NLModel(mlModel: coreMLModel)tagger.setModels([nlModel], forTagScheme: .nameType)

DON'T: Assume embeddings exist for all languages

Not all languages have word or sentence embeddings available on device.

swift
// WRONG: Force unwraplet embedding = NLEmbedding.wordEmbedding(for: .japanese)!
// CORRECT: Handle nilguard let embedding = NLEmbedding.wordEmbedding(for: .japanese) else {    // Embedding not available for this language    return}

DON'T: Create a new tagger per token

Creating and configuring a tagger is expensive. Reuse it for the same text.

swift
// WRONG: New tagger per wordfor word in words {    let tagger = NLTagger(tagSchemes: [.lexicalClass])    tagger.string = word}
// CORRECT: Set string once, enumeratelet tagger = NLTagger(tagSchemes: [.lexicalClass])tagger.string = fullTexttagger.enumerateTags(in: fullText.startIndex..<fullText.endIndex,                     unit: .word, scheme: .lexicalClass, options: []) { tag, range in    return true}

DON'T: Ignore language hints for short text

Language detection on short strings (under ~20 characters) is unreliable. Set constraints or hints to improve accuracy.

swift
// WRONG: Detect language of a single wordlet lang = NLLanguageRecognizer.dominantLanguage(for: "chat")  // French or English?
// CORRECT: Provide contextlet recognizer = NLLanguageRecognizer()recognizer.languageHints = [.english: 0.8, .french: 0.2]recognizer.processString("chat")

Review Checklist

  • NLTokenizer and NLTagger instances used from a single thread
  • Tagger created once per text, not per token
  • Language detection uses constraints/hints for short text
  • NLEmbedding availability checked before use (returns nil if unavailable)
  • Translation LanguageAvailability checked before attempting translation
  • .translationTask() used within a SwiftUI view hierarchy
  • Batch translation uses clientIdentifier to match responses to requests
  • Sentiment scores handled as optional (may return nil for unsupported languages)
  • .joinNames option used with NER to keep multi-word names together
  • Custom ML models loaded via NLModel, not raw Core ML

References

Source and attribution

Source:dpearson2699/swift-ios-skillsinskills/natural-languageat commit8d90fd1

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal

More from dpearson2699/swift-ios-skills

Widgetkit

dpearson2699

Guides implementing, reviewing, and improving WidgetKit widgets and controls for iOS, iPadOS, watchOS, and CarPlay.

Software Development1.1Kupdated 2 months ago

Weatherkit

dpearson2699

Guides iOS developers in fetching WeatherKit forecasts, alerts, and attribution using WeatherService.

Software Development1.1Kupdated 2 months ago

Vision Framework

dpearson2699

Implement computer vision features including text recognition (OCR), face detection, barcode scanning, image segmentation, object tracking, and document scanning in iOS apps. Covers both the modern Swift-native Vision API (iOS 18+) and legacy VNRequest patterns, VisionKit DataScannerViewController for live camera scanning, and CoreMLRequest/VNCoreMLRequest for custom model inference. Use when adding OCR, barcode scanning, face detection, or custom Core ML model inference with Vision.

Awaiting classification1.1Kupdated 2 months ago

Tipkit

dpearson2699

Implement and review Apple TipKit feature-discovery UI for iOS 17+ apps. Use when adding or auditing in-app tips, contextual help, coach marks, Tip, TipView, popoverTip, rules, events, actions, display frequency, testing overrides, reusable tip identifiers, or iOS 18+ TipGroup and CloudKit tip sync; avoid for generic SwiftUI navigation or layout outside tip presentation.

Awaiting classification1.1Kupdated 2 months ago

Tabletopkit

dpearson2699

Guides building multiplayer spatial board games on visionOS with Apple's TabletopKit and RealityKit.

Software Development1.1Kupdated 2 months ago

Swiftui Webkit

dpearson2699

Guides embedding and controlling web content in SwiftUI apps with WebKit for SwiftUI on iOS 26 and later.

Software Development1.1Kupdated 2 months ago