Axiom Vision

charleswiltgen/axiom/.claude-plugin/plugins/axiom/skills/axiom-vision

作者 charleswiltgene45d98ffbb5fMIT1.1K 個星標收錄於 2026年10月9日更新於 2026年10月8日儲存庫今天更新

Use when implementing ANY computer vision feature — image analysis, pose detection, person segmentation, subject lifting, text recognition, barcode scanning.

AI 產生的概覽

將 Apple Vision 框架的電腦視覺工作分流到實作、API 參考與診斷指南。

功能
此技能是針對 Apple Vision 框架電腦視覺工作的分流指南。它把主體分割、手部與身體姿態偵測、文字辨識、條碼掃描、文件掃描和即時掃描等任務,對應到三份參考文件,分別涵蓋實作模式、API 參考和診斷。它也列出常見故障現象,並指向針對偵測、信心度、效能和座標轉換問題的疑難排解指引。
適用情境
在實作或偵錯任何使用 Vision 框架的電腦視覺功能時使用,包括影像分析、姿態偵測、分割、OCR、條碼掃描和文件掃描。在需要查閱 Vision API 細節,或解決偵測、信心度、效能和座標問題時也適用。跨媒體資料庫的人臉分群則被指向另一個媒體技能。},
執行需求
無需指令碼,僅為說明文件。需要 Apple Vision 框架及相關 Apple 平台 SDK(iOS、iPadOS、macOS、watchOS)以使用所引用的 API。

Computer Vision

You MUST use this skill for ANY computer vision work using the Vision framework.

Quick Reference

Symptom / TaskReference
Subject segmentation, liftingSee skills/vision-framework.md
Hand/body pose detectionSee skills/vision-framework.md
Text recognition (OCR)See skills/vision-framework.md
Barcode/QR code detectionSee skills/vision-framework.md
Document scanningSee skills/vision-framework.md
DataScannerViewControllerSee skills/vision-framework.md
Structured document extraction (iOS 26+)See skills/vision-framework.md
Isolate object excluding handSee skills/vision-framework.md
Tap-to-segment any object OS27See skills/vision-ref.md
Vision on watchOS watchOS27See skills/vision-ref.md
Vision tools for Foundation Models (BarcodeReaderTool, OCRTool) OS27See skills/vision-ref.md
Vision framework API referenceSee skills/vision-ref.md
Visual Intelligence integration (iOS 26+, iPadOS27/macOS27)See skills/vision-ref.md
Sensitive content classification (nudity/gore/violence), categorized via detectedTypes (OS27)See skills/vision-ref.md
Group/cluster faces into people across a library, video highlights/key frames (OS27)Use axiom-media (skills/media-intelligence.md) instead — MediaIntelligence clusters identities; Vision detects faces in one image
Subject not detectedSee skills/vision-diag.md
Hand/body pose missing landmarksSee skills/vision-diag.md
Low confidence observationsSee skills/vision-diag.md
UI freezing during processingSee skills/vision-diag.md
Coordinate conversion bugsSee skills/vision-diag.md
Text not recognized / wrong charsSee skills/vision-diag.md
Barcode not detectedSee skills/vision-diag.md
DataScanner blank / no itemsSee skills/vision-diag.md
Document edges not detectedSee skills/vision-diag.md

Decision Tree

dot
digraph vision {    start [label="Computer vision task" shape=ellipse];    what [label="What do you need?" shape=diamond];
    start -> what;    what -> "skills/vision-framework.md" [label="implement feature"];    what -> "skills/vision-ref.md" [label="API reference"];    what -> "skills/vision-ref.md" [label="Visual Intelligence"];    what -> "skills/vision-ref.md" [label="tap-to-segment / watchOS / FM tools (27)"];    what -> "skills/vision-diag.md" [label="something broken"];}
  1. Implementing (pose, segmentation, OCR, barcodes, documents, live scanning)? → skills/vision-framework.md
  2. Visual Intelligence system integration (camera/screenshot search; iOS 26+, iPadOS27/macOS27)? → skills/vision-ref.md (Visual Intelligence section)
  3. Tap-to-segment, Vision on watchOS, or Vision tools for Foundation Models (27 cycle)? → skills/vision-ref.md
  4. Need API reference / code examples? → skills/vision-ref.md
  5. Debugging issues (detection failures, confidence, coordinates)? → skills/vision-diag.md

Critical Patterns

Implementation (skills/vision-framework.md):

  • Decision tree for choosing the right Vision API
  • Subject segmentation with VisionKit
  • Isolating objects while excluding hands (combining APIs)
  • Hand/body pose detection (21/19 landmarks)
  • Text recognition (fast vs accurate modes)
  • Barcode detection with symbology selection
  • Document scanning and structured extraction (iOS 26+)
  • Live scanning with DataScannerViewController
  • CoreImage HDR compositing

Diagnostics (skills/vision-diag.md):

  • Subject detection failures (edge of frame, lighting)
  • Landmark tracking issues (confidence thresholds)
  • Performance optimization (frame skipping, downscaling)
  • Coordinate conversion (lower-left vs top-left origin)
  • Text recognition failures (language, contrast)
  • Barcode detection issues (symbology, size, glare)
  • DataScanner troubleshooting (availability, data types)

Anti-Rationalization

ThoughtReality
"Vision framework is just a request/handler pattern"Vision has coordinate conversion, confidence thresholds, and performance gotchas. vision-framework.md covers them.
"I'll handle text recognition without the skill"VNRecognizeTextRequest has fast/accurate modes and language-specific settings. vision-framework.md has the patterns.
"Subject segmentation is straightforward"Instance masks have HDR compositing and hand-exclusion patterns. vision-framework.md covers complex scenarios.
"Visual Intelligence is just the camera API"Visual Intelligence is a system-level feature requiring IntentValueQuery and SemanticContentDescriptor. vision-ref.md has the integration section.
"I'll just process on the main thread"Vision blocks UI on older devices. Users on iPhone 12 will experience frozen app. 15 min to add background queue.

Example Invocations

User: "How do I detect hand pose in an image?" → See skills/vision-framework.md

User: "Isolate a subject but exclude the user's hands" → See skills/vision-framework.md

User: "How do I read text from an image?" → See skills/vision-framework.md

User: "Scan QR codes with the camera" → See skills/vision-framework.md

User: "Subject detection isn't working" → See skills/vision-diag.md

User: "Text recognition returns wrong characters" → See skills/vision-diag.md

User: "Show me VNDetectHumanBodyPoseRequest examples" → See skills/vision-ref.md

User: "How do I make my app work with Visual Intelligence?" → See skills/vision-ref.md

User: "Let users tap an object in a photo to cut it out" → See skills/vision-ref.md (Iterative Segmentation)

User: "Can I use Vision in my watchOS app?" → See skills/vision-ref.md (Vision on watchOS)

User: "RecognizeDocumentsRequest API reference" → See skills/vision-ref.md

User: "Group faces into people across my library" / "cluster faces on-device into persons" → Use axiom-media (skills/media-intelligence.md) — identity clustering across assets, not per-image detection

來源與署名

來源:charleswiltgen/axiom位於.claude-plugin/plugins/axiom/skills/axiom-vision提交e45d98f

授權條款: MIT

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架