Pdf

XiaoMaColtAI/math-modeling-skill/dsh-plugin/math-modeling-agent/skills/math-modeling/tools/pdf

by XiaoMaColtAI8fd1f4b6a5f1de927202673bec1953fa0e148718Proprietary. LICENSE.txt has complete termsListed Oct 9, 2026Updated Oct 9, 2026

Use this skill whenever the user wants to do anything with PDF files. This includes reading or extracting text/tables from PDFs, combining or merging multiple PDFs into one, splitting PDFs apart, rotating pages, adding watermarks, creating new PDFs, filling PDF forms, encrypting/decrypting PDFs, extracting images, and OCR on scanned PDFs to make them searchable. If the user mentions a .pdf file or asks to produce one, use this skill.

Instructions onlyDocuments & Office
AI-generated overview

Guide for working with PDF files: reading, extracting, merging, splitting, creating, watermarking, encrypting and OCR.

What it does
Provides instructions and code examples for common PDF operations using Python libraries and command-line tools. It covers reading and extracting text, tables and images, merging and splitting files, rotating pages, adding watermarks, creating new PDFs, filling forms, encrypting and decrypting, and OCR on scanned documents. It also points to reference.md for advanced and JavaScript options and FORMS.md for form filling.
When to use it
Use when a user mentions a .pdf file or asks to produce one, or needs PDF text, tables or images extracted, pages merged, split or rotated, a watermark or password applied, or a scanned PDF made searchable. Also use when a PDF form must be filled.
Requirements
Requires Python with pypdf, pdfplumber, reportlab, pandas, pytesseract and pdf2image for the Python examples, plus command-line tools pdftotext and pdfimages (poppler-utils), qpdf and optionally pdftk. OCR needs Tesseract. Ships no scripts; instructions only, with references to reference.md and FORMS.md.

PDF Processing Guide

Overview

This guide covers essential PDF processing operations using Python libraries and command-line tools. For advanced features, JavaScript libraries, and detailed examples, see reference.md. If you need to fill out a PDF form, read FORMS.md and follow its instructions.

Quick Start

python
from pypdf import PdfReader, PdfWriter
# Read a PDFreader = PdfReader("document.pdf")print(f"Pages: {len(reader.pages)}")
# Extract texttext = ""for page in reader.pages:    text += page.extract_text()

Python Libraries

pypdf - Basic Operations

Merge PDFs
python
from pypdf import PdfWriter, PdfReader
writer = PdfWriter()for pdf_file in ["doc1.pdf", "doc2.pdf", "doc3.pdf"]:    reader = PdfReader(pdf_file)    for page in reader.pages:        writer.add_page(page)
with open("merged.pdf", "wb") as output:    writer.write(output)
Split PDF
python
reader = PdfReader("input.pdf")for i, page in enumerate(reader.pages):    writer = PdfWriter()    writer.add_page(page)    with open(f"page_{i+1}.pdf", "wb") as output:        writer.write(output)
Extract Metadata
python
reader = PdfReader("document.pdf")meta = reader.metadataprint(f"Title: {meta.title}")print(f"Author: {meta.author}")print(f"Subject: {meta.subject}")print(f"Creator: {meta.creator}")
Rotate Pages
python
reader = PdfReader("input.pdf")writer = PdfWriter()
page = reader.pages[0]page.rotate(90)  # Rotate 90 degrees clockwisewriter.add_page(page)
with open("rotated.pdf", "wb") as output:    writer.write(output)

pdfplumber - Text and Table Extraction

Extract Text with Layout
python
import pdfplumber
with pdfplumber.open("document.pdf") as pdf:    for page in pdf.pages:        text = page.extract_text()        print(text)
Extract Tables
python
with pdfplumber.open("document.pdf") as pdf:    for i, page in enumerate(pdf.pages):        tables = page.extract_tables()        for j, table in enumerate(tables):            print(f"Table {j+1} on page {i+1}:")            for row in table:                print(row)
Advanced Table Extraction
python
import pandas as pd
with pdfplumber.open("document.pdf") as pdf:    all_tables = []    for page in pdf.pages:        tables = page.extract_tables()        for table in tables:            if table:  # Check if table is not empty                df = pd.DataFrame(table[1:], columns=table[0])                all_tables.append(df)
# Combine all tablesif all_tables:    combined_df = pd.concat(all_tables, ignore_index=True)    combined_df.to_excel("extracted_tables.xlsx", index=False)

reportlab - Create PDFs

Basic PDF Creation
python
from reportlab.lib.pagesizes import letterfrom reportlab.pdfgen import canvas
c = canvas.Canvas("hello.pdf", pagesize=letter)width, height = letter
# Add textc.drawString(100, height - 100, "Hello World!")c.drawString(100, height - 120, "This is a PDF created with reportlab")
# Add a linec.line(100, height - 140, 400, height - 140)
# Savec.save()
Create PDF with Multiple Pages
python
from reportlab.lib.pagesizes import letterfrom reportlab.platypus import SimpleDocTemplate, Paragraph, Spacer, PageBreakfrom reportlab.lib.styles import getSampleStyleSheet
doc = SimpleDocTemplate("report.pdf", pagesize=letter)styles = getSampleStyleSheet()story = []
# Add contenttitle = Paragraph("Report Title", styles['Title'])story.append(title)story.append(Spacer(1, 12))
body = Paragraph("This is the body of the report. " * 20, styles['Normal'])story.append(body)story.append(PageBreak())
# Page 2story.append(Paragraph("Page 2", styles['Heading1']))story.append(Paragraph("Content for page 2", styles['Normal']))
# Build PDFdoc.build(story)
Subscripts and Superscripts

IMPORTANT: Never use Unicode subscript/superscript characters (₀₁₂₃₄₅₆₇₈₉, ⁰¹²³⁴⁵⁶⁷⁸⁹) in ReportLab PDFs. The built-in fonts do not include these glyphs, causing them to render as solid black boxes.

Instead, use ReportLab's XML markup tags in Paragraph objects:

python
from reportlab.platypus import Paragraphfrom reportlab.lib.styles import getSampleStyleSheet
styles = getSampleStyleSheet()
# Subscripts: use <sub> tagchemical = Paragraph("H<sub>2</sub>O", styles['Normal'])
# Superscripts: use <super> tagsquared = Paragraph("x<super>2</super> + y<super>2</super>", styles['Normal'])

For canvas-drawn text (not Paragraph objects), manually adjust font the size and position rather than using Unicode subscripts/superscripts.

Command-Line Tools

pdftotext (poppler-utils)

bash
# Extract textpdftotext input.pdf output.txt
# Extract text preserving layoutpdftotext -layout input.pdf output.txt
# Extract specific pagespdftotext -f 1 -l 5 input.pdf output.txt  # Pages 1-5

qpdf

bash
# Merge PDFsqpdf --empty --pages file1.pdf file2.pdf -- merged.pdf
# Split pagesqpdf input.pdf --pages . 1-5 -- pages1-5.pdfqpdf input.pdf --pages . 6-10 -- pages6-10.pdf
# Rotate pagesqpdf input.pdf output.pdf --rotate=+90:1  # Rotate page 1 by 90 degrees
# Remove passwordqpdf --password=mypassword --decrypt encrypted.pdf decrypted.pdf

pdftk (if available)

bash
# Mergepdftk file1.pdf file2.pdf cat output merged.pdf
# Splitpdftk input.pdf burst
# Rotatepdftk input.pdf rotate 1east output rotated.pdf

Common Tasks

Extract Text from Scanned PDFs

python
# Requires: pip install pytesseract pdf2imageimport pytesseractfrom pdf2image import convert_from_path
# Convert PDF to imagesimages = convert_from_path('scanned.pdf')
# OCR each pagetext = ""for i, image in enumerate(images):    text += f"Page {i+1}:\n"    text += pytesseract.image_to_string(image)    text += "\n\n"
print(text)

Add Watermark

python
from pypdf import PdfReader, PdfWriter
# Create watermark (or load existing)watermark = PdfReader("watermark.pdf").pages[0]
# Apply to all pagesreader = PdfReader("document.pdf")writer = PdfWriter()
for page in reader.pages:    page.merge_page(watermark)    writer.add_page(page)
with open("watermarked.pdf", "wb") as output:    writer.write(output)

Extract Images

bash
# Using pdfimages (poppler-utils)pdfimages -j input.pdf output_prefix
# This extracts all images as output_prefix-000.jpg, output_prefix-001.jpg, etc.

Password Protection

python
from pypdf import PdfReader, PdfWriter
reader = PdfReader("input.pdf")writer = PdfWriter()
for page in reader.pages:    writer.add_page(page)
# Add passwordwriter.encrypt("userpassword", "ownerpassword")
with open("encrypted.pdf", "wb") as output:    writer.write(output)

Quick Reference

TaskBest ToolCommand/Code
Merge PDFspypdfwriter.add_page(page)
Split PDFspypdfOne page per file
Extract textpdfplumberpage.extract_text()
Extract tablespdfplumberpage.extract_tables()
Create PDFsreportlabCanvas or Platypus
Command line mergeqpdfqpdf --empty --pages ...
OCR scanned PDFspytesseractConvert to image first
Fill PDF formspdf-lib or pypdf (see FORMS.md)See FORMS.md

Next Steps

  • For advanced pypdfium2 usage, see reference.md
  • For JavaScript libraries (pdf-lib), see reference.md
  • If you need to fill out a PDF form, follow the instructions in FORMS.md
  • For troubleshooting guides, see reference.md

Source and attribution

Source:XiaoMaColtAI/math-modeling-skillindsh-plugin/math-modeling-agent/skills/math-modeling/tools/pdfat commit8fd1f4b

License: Proprietary. LICENSE.txt has complete terms

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal