Image To Text

pascalorg/skills/image-to-text

作者 pascalorg7be87e9292e12fe412d1eb8582fd95fdbd151325無授權條款收錄於 2026年10月9日更新於 2026年10月9日

Extract text from images using OCR. Use when the user shares a screenshot and you need to read the text content, copy UI labels, or extract copy from a design mockup.

包含腳本Data & Analytics
AI 產生的概覽

使用 OCR 從圖片擷取可讀文字,回傳完整文字以及詞級邊界框與信心分數。

功能
此技能對圖片執行 OCR(Tesseract),回傳辨識出的文字,以及詞級與行級的邊界框和信心分數。它附帶 shell 與 JavaScript 指令碼,接受圖片路徑與選用的語言代碼,預設為英文。結果會依行分組呈現,並附上整體信心百分比。
適用情境
當使用者分享含有文字的螢幕截圖、設計稿或照片,需要讀取或重複使用其中文字內容時使用。它也適合擷取介面文案,例如標籤、按鈕與標題,或從設計圖取得文字位置。
執行需求
需要可執行所附指令碼的 shell 環境,以及搭配 Tesseract.js 的 Node.js;首次執行會下載語言資料(英文約 4MB),因此初次使用需要網路連線。必須提供圖片檔案路徑;OCR 語言代碼為選用,預設為 eng。

Image to Text

Extract all readable text from an image using OCR (Tesseract). Returns the full text content along with word-level bounding boxes and confidence scores.

When to Use

  • Reading text content from a screenshot or design mockup
  • Extracting UI copy (labels, buttons, headings) so you don't have to retype it
  • Getting text positions and bounding boxes from a design image

How It Works

  1. The image is passed to Tesseract.js for optical character recognition
  2. Tesseract segments the image into lines and words
  3. Returns the full text plus word-level details (position, confidence)

Usage

bash
bash <skill-path>/scripts/image-to-text.sh <image-path> [language]

Arguments:

  • image-path — Path to the image file (required)
  • language — OCR language code (optional, defaults to eng). Common: eng, fra, deu, spa, chi_sim, jpn

Examples:

bash
# Extract text from a screenshotbash <skill-path>/scripts/image-to-text.sh ./screenshot.png
# Extract French textbash <skill-path>/scripts/image-to-text.sh ./mockup.png fra

Output

json
{  "text": "Request work\nSuggestions\nPlumbing\nHVAC\nCleaning\nElectrical",  "confidence": 87.4,  "words": [    {      "text": "Request",      "confidence": 94.2,      "bbox": { "x0": 142, "y0": 180, "x1": 268, "y1": 204 }    },    {      "text": "work",      "confidence": 96.1,      "bbox": { "x0": 274, "y0": 180, "x1": 332, "y1": 204 }    }  ],  "lines": [    {      "text": "Request work",      "confidence": 95.1,      "bbox": { "x0": 142, "y0": 180, "x1": 332, "y1": 204 }    }  ]}
FieldTypeDescription
textStringFull extracted text, newline-separated
confidenceNumberOverall confidence score (0-100)
wordsArrayEach word with text, confidence, and bounding box
linesArrayEach line with text, confidence, and bounding box

Present Results to User

After extracting text, present the content grouped by lines:

Extracted text (87.4% confidence):
  Request work  Suggestions  Plumbing  HVAC  Cleaning  Electrical
Found 6 lines, 6 words.

Use the extracted text directly when implementing UI copy from a design.

Troubleshooting

Low confidence / garbled text — Tesseract works best with clean, high-contrast text. Screenshots of rendered UI work well. Photos of text at angles or with noise may produce poor results.

Wrong language — Pass the correct language code as the second argument. Tesseract needs the right language model to recognize characters.

First run is slow — Tesseract downloads language data (~4MB for English) on the first run. Subsequent runs are faster.

來源與署名

來源:pascalorg/skills位於image-to-text提交7be87e9

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架