PDF Extraction Skill
Overview
This skill enables precise extraction of text, tables, and metadata from PDF documents using pdfplumber - the go-to library for PDF data extraction. Unlike basic PDF readers, pdfplumber provides detailed character-level positioning, accurate table detection, and visual debugging.
How to Use
- Provide the PDF file you want to extract from
- Specify what you need: text, tables, images, or metadata
- I'll generate pdfplumber code and execute it
Example prompts:
- "Extract all tables from this financial report"
- "Get text from pages 5-10 of this document"
- "Find and extract the invoice total from this PDF"
- "Convert this PDF table to CSV/Excel"
Domain Knowledge
pdfplumber Fundamentals
PDF Structure
Text Extraction
Basic Text
Advanced Text Options
Character-Level Access
Table Extraction
Basic Table Extraction
Advanced Table Settings
Table Finding
Visual Debugging
Cropping and Filtering
Crop to Region
Filter by Position
Filter by Font
Metadata and Structure
Best Practices
- Debug Visually: Use
to_image()to understand PDF structure - Tune Table Settings: Adjust tolerances for your specific PDF
- Handle Scanned PDFs: Use OCR first (this skill is for native text)
- Process Page by Page: For large PDFs, avoid loading all at once
- Check for Text: Some PDFs are images - verify text exists
Common Patterns
Extract All Tables to DataFrames
Extract Specific Region
Multi-column Layout
Examples
Example 1: Financial Report Table Extraction
Example 2: Invoice Data Extraction
Example 3: Resume/CV Parser
Limitations
- Cannot extract from scanned/image PDFs (use OCR first)
- Complex layouts may need manual tuning
- Some PDF encryption types not supported
- Embedded fonts may affect text extraction
- No direct PDF editing capability



