Image To Text

pascalorg/skills/image-to-text

作者 pascalorg7be87e9292e12fe412d1eb8582fd95fdbd151325无许可证收录于 2026年10月9日更新于 2026年10月9日

Extract text from images using OCR. Use when the user shares a screenshot and you need to read the text content, copy UI labels, or extract copy from a design mockup.

包含脚本Data & Analytics
AI 生成的概览

使用 OCR 从图片中提取可读文本,返回完整文本以及词级边界框和置信度分数。

功能
该技能对图片运行 OCR(Tesseract),返回识别出的文本,以及词级和行级的边界框与置信度分数。它附带 shell 和 JavaScript 脚本,接受图片路径和可选的语言代码,默认使用英语。结果按行分组呈现,并给出整体置信度百分比。
适用场景
当用户分享包含文字的截图、设计稿或照片,需要读取或复用其中的文字内容时使用。它也适合提取界面文案,如标签、按钮和标题,或从设计图中获取文字位置。
运行要求
需要可运行所附脚本的 shell 环境,以及带 Tesseract.js 的 Node.js;首次运行会下载语言数据(英语约 4MB),因此初次使用需要网络访问。必须提供图片文件路径;OCR 语言代码可选,默认为 eng。

Image to Text

Extract all readable text from an image using OCR (Tesseract). Returns the full text content along with word-level bounding boxes and confidence scores.

When to Use

  • Reading text content from a screenshot or design mockup
  • Extracting UI copy (labels, buttons, headings) so you don't have to retype it
  • Getting text positions and bounding boxes from a design image

How It Works

  1. The image is passed to Tesseract.js for optical character recognition
  2. Tesseract segments the image into lines and words
  3. Returns the full text plus word-level details (position, confidence)

Usage

bash
bash <skill-path>/scripts/image-to-text.sh <image-path> [language]

Arguments:

  • image-path — Path to the image file (required)
  • language — OCR language code (optional, defaults to eng). Common: eng, fra, deu, spa, chi_sim, jpn

Examples:

bash
# Extract text from a screenshotbash <skill-path>/scripts/image-to-text.sh ./screenshot.png
# Extract French textbash <skill-path>/scripts/image-to-text.sh ./mockup.png fra

Output

json
{  "text": "Request work\nSuggestions\nPlumbing\nHVAC\nCleaning\nElectrical",  "confidence": 87.4,  "words": [    {      "text": "Request",      "confidence": 94.2,      "bbox": { "x0": 142, "y0": 180, "x1": 268, "y1": 204 }    },    {      "text": "work",      "confidence": 96.1,      "bbox": { "x0": 274, "y0": 180, "x1": 332, "y1": 204 }    }  ],  "lines": [    {      "text": "Request work",      "confidence": 95.1,      "bbox": { "x0": 142, "y0": 180, "x1": 332, "y1": 204 }    }  ]}
FieldTypeDescription
textStringFull extracted text, newline-separated
confidenceNumberOverall confidence score (0-100)
wordsArrayEach word with text, confidence, and bounding box
linesArrayEach line with text, confidence, and bounding box

Present Results to User

After extracting text, present the content grouped by lines:

Extracted text (87.4% confidence):
  Request work  Suggestions  Plumbing  HVAC  Cleaning  Electrical
Found 6 lines, 6 words.

Use the extracted text directly when implementing UI copy from a design.

Troubleshooting

Low confidence / garbled text — Tesseract works best with clean, high-contrast text. Screenshots of rendered UI work well. Photos of text at angles or with noise may produce poor results.

Wrong language — Pass the correct language code as the second argument. Tesseract needs the right language model to recognize characters.

First run is slow — Tesseract downloads language data (~4MB for English) on the first run. Subsequent runs are faster.

来源与署名

来源:pascalorg/skills位于image-to-text提交7be87e9

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架