PDF to Word Skill
Overview
This skill enables conversion from PDF to editable Word documents using pdf2docx - a Python library that preserves layout, tables, images, and text formatting. Unlike OCR-based solutions, pdf2docx extracts native PDF content for accurate conversion.
How to Use
- Provide the PDF file you want to convert
- Optionally specify pages or conversion options
- I'll convert it to an editable Word document
Example prompts:
- "Convert this PDF report to an editable Word document"
- "Turn pages 1-5 of this PDF into Word format"
- "Extract this scanned document as editable text"
- "Convert this PDF contract to Word for editing"
Domain Knowledge
pdf2docx Fundamentals
Conversion Options
Advanced Options
Handling Different PDF Types
Native PDFs (Text-based)
Scanned PDFs (Image-based)
Python Integration
Batch Conversion
Parsing PDF Structure
Best Practices
- Check PDF Type: Native PDFs convert better than scanned
- Preview First: Test with a few pages before full conversion
- Handle Tables: Complex tables may need manual adjustment
- Image Quality: Images are extracted at original resolution
- Font Handling: Some fonts may substitute to system defaults
Common Patterns
Convert with Progress
Extract Tables Only
Examples
Example 1: Contract Conversion
Example 2: Selective Page Conversion
Example 3: PDF Report to Editable Template
Example 4: Bulk Invoice Processing
Limitations
- Scanned PDFs require OCR preprocessing
- Complex layouts may not convert perfectly
- Some fonts may not be available
- Watermarks are included in conversion
- Protected/encrypted PDFs need password



