Image To Text

pascalorg/skills/image-to-text

by pascalorg7be87e9292e12fe412d1eb8582fd95fdbd151325No licenseListed Oct 9, 2026Updated Oct 9, 2026

Extract text from images using OCR. Use when the user shares a screenshot and you need to read the text content, copy UI labels, or extract copy from a design mockup.

Includes scriptsData & Analytics
AI-generated overview

Extracts readable text from images with OCR, returning full text plus word-level bounding boxes and confidence scores.

What it does
This skill runs OCR (Tesseract) over an image and returns the recognized text along with word-level and line-level bounding boxes and confidence scores. It ships shell and JavaScript scripts that accept an image path and an optional language code, defaulting to English. Results are presented grouped by line with an overall confidence percentage.
When to use it
Use it when a user shares a screenshot, design mockup, or photo containing text and the text content needs to be read or reused. It is also suited to pulling UI copy such as labels, buttons, and headings, or to obtaining text positions from a design image.
Requirements
Requires a shell environment to run the bundled scripts and Node.js with Tesseract.js, which downloads language data (about 4MB for English) on first run, so network access is needed initially. An image file path is required; the OCR language code is optional and defaults to eng.

Image to Text

Extract all readable text from an image using OCR (Tesseract). Returns the full text content along with word-level bounding boxes and confidence scores.

When to Use

  • Reading text content from a screenshot or design mockup
  • Extracting UI copy (labels, buttons, headings) so you don't have to retype it
  • Getting text positions and bounding boxes from a design image

How It Works

  1. The image is passed to Tesseract.js for optical character recognition
  2. Tesseract segments the image into lines and words
  3. Returns the full text plus word-level details (position, confidence)

Usage

bash
bash <skill-path>/scripts/image-to-text.sh <image-path> [language]

Arguments:

  • image-path — Path to the image file (required)
  • language — OCR language code (optional, defaults to eng). Common: eng, fra, deu, spa, chi_sim, jpn

Examples:

bash
# Extract text from a screenshotbash <skill-path>/scripts/image-to-text.sh ./screenshot.png
# Extract French textbash <skill-path>/scripts/image-to-text.sh ./mockup.png fra

Output

json
{  "text": "Request work\nSuggestions\nPlumbing\nHVAC\nCleaning\nElectrical",  "confidence": 87.4,  "words": [    {      "text": "Request",      "confidence": 94.2,      "bbox": { "x0": 142, "y0": 180, "x1": 268, "y1": 204 }    },    {      "text": "work",      "confidence": 96.1,      "bbox": { "x0": 274, "y0": 180, "x1": 332, "y1": 204 }    }  ],  "lines": [    {      "text": "Request work",      "confidence": 95.1,      "bbox": { "x0": 142, "y0": 180, "x1": 332, "y1": 204 }    }  ]}
FieldTypeDescription
textStringFull extracted text, newline-separated
confidenceNumberOverall confidence score (0-100)
wordsArrayEach word with text, confidence, and bounding box
linesArrayEach line with text, confidence, and bounding box

Present Results to User

After extracting text, present the content grouped by lines:

Extracted text (87.4% confidence):
  Request work  Suggestions  Plumbing  HVAC  Cleaning  Electrical
Found 6 lines, 6 words.

Use the extracted text directly when implementing UI copy from a design.

Troubleshooting

Low confidence / garbled text — Tesseract works best with clean, high-contrast text. Screenshots of rendered UI work well. Photos of text at angles or with noise may produce poor results.

Wrong language — Pass the correct language code as the second argument. Tesseract needs the right language model to recognize characters.

First run is slow — Tesseract downloads language data (~4MB for English) on the first run. Subsequent runs are faster.

Source and attribution

Source:pascalorg/skillsinimage-to-textat commit7be87e9

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal