Pdf

XiaoMaColtAI/math-modeling-skill/dsh-plugin/math-modeling-agent/skills/math-modeling/tools/pdf

作者 XiaoMaColtAI8fd1f4b6a5f1de927202673bec1953fa0e148718Proprietary. LICENSE.txt has complete terms收录于 2026年10月9日更新于 2026年10月9日

Use this skill whenever the user wants to do anything with PDF files. This includes reading or extracting text/tables from PDFs, combining or merging multiple PDFs into one, splitting PDFs apart, rotating pages, adding watermarks, creating new PDFs, filling PDF forms, encrypting/decrypting PDFs, extracting images, and OCR on scanned PDFs to make them searchable. If the user mentions a .pdf file or asks to produce one, use this skill.

仅含说明Documents & Office
AI 生成的概览

处理 PDF 文件的指南:读取、提取、合并、拆分、创建、加水印、加密与 OCR。

功能
提供使用 Python 库和命令行工具完成常见 PDF 操作的说明与代码示例。涵盖读取以及提取文本、表格和图片,合并与拆分文件,旋转页面,添加水印,创建新 PDF,填写表单,加密与解密,以及对扫描件进行 OCR。还指引查阅 reference.md 了解进阶与 JavaScript 方案,以及 FORMS.md 了解表单填写。
适用场景
当用户提到 .pdf 文件或要求生成 PDF 时使用;也适用于需要提取 PDF 文本、表格或图片,合并、拆分或旋转页面,添加水印或密码,或让扫描版 PDF 可搜索的场景。需要填写 PDF 表单时同样适用。
运行要求
需要 Python 及 pypdf、pdfplumber、reportlab、pandas、pytesseract、pdf2image 等库以运行示例,并需要命令行工具 pdftotext、pdfimages(poppler-utils)、qpdf,可选 pdftk。OCR 需要 Tesseract。不附带脚本,仅为说明文档,并引用 reference.md 与 FORMS.md。

PDF Processing Guide

Overview

This guide covers essential PDF processing operations using Python libraries and command-line tools. For advanced features, JavaScript libraries, and detailed examples, see reference.md. If you need to fill out a PDF form, read FORMS.md and follow its instructions.

Quick Start

python
from pypdf import PdfReader, PdfWriter
# Read a PDFreader = PdfReader("document.pdf")print(f"Pages: {len(reader.pages)}")
# Extract texttext = ""for page in reader.pages:    text += page.extract_text()

Python Libraries

pypdf - Basic Operations

Merge PDFs
python
from pypdf import PdfWriter, PdfReader
writer = PdfWriter()for pdf_file in ["doc1.pdf", "doc2.pdf", "doc3.pdf"]:    reader = PdfReader(pdf_file)    for page in reader.pages:        writer.add_page(page)
with open("merged.pdf", "wb") as output:    writer.write(output)
Split PDF
python
reader = PdfReader("input.pdf")for i, page in enumerate(reader.pages):    writer = PdfWriter()    writer.add_page(page)    with open(f"page_{i+1}.pdf", "wb") as output:        writer.write(output)
Extract Metadata
python
reader = PdfReader("document.pdf")meta = reader.metadataprint(f"Title: {meta.title}")print(f"Author: {meta.author}")print(f"Subject: {meta.subject}")print(f"Creator: {meta.creator}")
Rotate Pages
python
reader = PdfReader("input.pdf")writer = PdfWriter()
page = reader.pages[0]page.rotate(90)  # Rotate 90 degrees clockwisewriter.add_page(page)
with open("rotated.pdf", "wb") as output:    writer.write(output)

pdfplumber - Text and Table Extraction

Extract Text with Layout
python
import pdfplumber
with pdfplumber.open("document.pdf") as pdf:    for page in pdf.pages:        text = page.extract_text()        print(text)
Extract Tables
python
with pdfplumber.open("document.pdf") as pdf:    for i, page in enumerate(pdf.pages):        tables = page.extract_tables()        for j, table in enumerate(tables):            print(f"Table {j+1} on page {i+1}:")            for row in table:                print(row)
Advanced Table Extraction
python
import pandas as pd
with pdfplumber.open("document.pdf") as pdf:    all_tables = []    for page in pdf.pages:        tables = page.extract_tables()        for table in tables:            if table:  # Check if table is not empty                df = pd.DataFrame(table[1:], columns=table[0])                all_tables.append(df)
# Combine all tablesif all_tables:    combined_df = pd.concat(all_tables, ignore_index=True)    combined_df.to_excel("extracted_tables.xlsx", index=False)

reportlab - Create PDFs

Basic PDF Creation
python
from reportlab.lib.pagesizes import letterfrom reportlab.pdfgen import canvas
c = canvas.Canvas("hello.pdf", pagesize=letter)width, height = letter
# Add textc.drawString(100, height - 100, "Hello World!")c.drawString(100, height - 120, "This is a PDF created with reportlab")
# Add a linec.line(100, height - 140, 400, height - 140)
# Savec.save()
Create PDF with Multiple Pages
python
from reportlab.lib.pagesizes import letterfrom reportlab.platypus import SimpleDocTemplate, Paragraph, Spacer, PageBreakfrom reportlab.lib.styles import getSampleStyleSheet
doc = SimpleDocTemplate("report.pdf", pagesize=letter)styles = getSampleStyleSheet()story = []
# Add contenttitle = Paragraph("Report Title", styles['Title'])story.append(title)story.append(Spacer(1, 12))
body = Paragraph("This is the body of the report. " * 20, styles['Normal'])story.append(body)story.append(PageBreak())
# Page 2story.append(Paragraph("Page 2", styles['Heading1']))story.append(Paragraph("Content for page 2", styles['Normal']))
# Build PDFdoc.build(story)
Subscripts and Superscripts

IMPORTANT: Never use Unicode subscript/superscript characters (₀₁₂₃₄₅₆₇₈₉, ⁰¹²³⁴⁵⁶⁷⁸⁹) in ReportLab PDFs. The built-in fonts do not include these glyphs, causing them to render as solid black boxes.

Instead, use ReportLab's XML markup tags in Paragraph objects:

python
from reportlab.platypus import Paragraphfrom reportlab.lib.styles import getSampleStyleSheet
styles = getSampleStyleSheet()
# Subscripts: use <sub> tagchemical = Paragraph("H<sub>2</sub>O", styles['Normal'])
# Superscripts: use <super> tagsquared = Paragraph("x<super>2</super> + y<super>2</super>", styles['Normal'])

For canvas-drawn text (not Paragraph objects), manually adjust font the size and position rather than using Unicode subscripts/superscripts.

Command-Line Tools

pdftotext (poppler-utils)

bash
# Extract textpdftotext input.pdf output.txt
# Extract text preserving layoutpdftotext -layout input.pdf output.txt
# Extract specific pagespdftotext -f 1 -l 5 input.pdf output.txt  # Pages 1-5

qpdf

bash
# Merge PDFsqpdf --empty --pages file1.pdf file2.pdf -- merged.pdf
# Split pagesqpdf input.pdf --pages . 1-5 -- pages1-5.pdfqpdf input.pdf --pages . 6-10 -- pages6-10.pdf
# Rotate pagesqpdf input.pdf output.pdf --rotate=+90:1  # Rotate page 1 by 90 degrees
# Remove passwordqpdf --password=mypassword --decrypt encrypted.pdf decrypted.pdf

pdftk (if available)

bash
# Mergepdftk file1.pdf file2.pdf cat output merged.pdf
# Splitpdftk input.pdf burst
# Rotatepdftk input.pdf rotate 1east output rotated.pdf

Common Tasks

Extract Text from Scanned PDFs

python
# Requires: pip install pytesseract pdf2imageimport pytesseractfrom pdf2image import convert_from_path
# Convert PDF to imagesimages = convert_from_path('scanned.pdf')
# OCR each pagetext = ""for i, image in enumerate(images):    text += f"Page {i+1}:\n"    text += pytesseract.image_to_string(image)    text += "\n\n"
print(text)

Add Watermark

python
from pypdf import PdfReader, PdfWriter
# Create watermark (or load existing)watermark = PdfReader("watermark.pdf").pages[0]
# Apply to all pagesreader = PdfReader("document.pdf")writer = PdfWriter()
for page in reader.pages:    page.merge_page(watermark)    writer.add_page(page)
with open("watermarked.pdf", "wb") as output:    writer.write(output)

Extract Images

bash
# Using pdfimages (poppler-utils)pdfimages -j input.pdf output_prefix
# This extracts all images as output_prefix-000.jpg, output_prefix-001.jpg, etc.

Password Protection

python
from pypdf import PdfReader, PdfWriter
reader = PdfReader("input.pdf")writer = PdfWriter()
for page in reader.pages:    writer.add_page(page)
# Add passwordwriter.encrypt("userpassword", "ownerpassword")
with open("encrypted.pdf", "wb") as output:    writer.write(output)

Quick Reference

TaskBest ToolCommand/Code
Merge PDFspypdfwriter.add_page(page)
Split PDFspypdfOne page per file
Extract textpdfplumberpage.extract_text()
Extract tablespdfplumberpage.extract_tables()
Create PDFsreportlabCanvas or Platypus
Command line mergeqpdfqpdf --empty --pages ...
OCR scanned PDFspytesseractConvert to image first
Fill PDF formspdf-lib or pypdf (see FORMS.md)See FORMS.md

Next Steps

  • For advanced pypdfium2 usage, see reference.md
  • For JavaScript libraries (pdf-lib), see reference.md
  • If you need to fill out a PDF form, follow the instructions in FORMS.md
  • For troubleshooting guides, see reference.md

来源与署名

来源:XiaoMaColtAI/math-modeling-skill位于dsh-plugin/math-modeling-agent/skills/math-modeling/tools/pdf提交8fd1f4b

许可证: Proprietary. LICENSE.txt has complete terms

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架