Batch Processor

作者 claude-office-skills9c4c7d5cd281MIT499 个星标收录于 2026年10月8日更新于 2026年10月8日仓库8个月前更新

Process multiple documents in bulk with parallel execution

仅含说明Documents & Office
AI 生成的概览

并行批量处理大量文档,进行转换、变换、提取或分析,并跟踪进度。

功能
该技能描述了对大量文档进行批量处理的模式,包括使用工作进程池并行执行、进度反馈,以及通过检查点让长时间任务在失败后可以续跑。它涵盖文件转换、变换、内容提取和分析,并给出批量处理与错误处理的 Python 示例代码。它仅提供说明,产出的是处理后的文件或提取结果,而不是单个文档。
适用场景
当需要对成百上千个文件执行相同操作时使用,例如把 PDF 转成 Word、从图片中提取文字、批量重命名,或批量更新页眉页脚。也适用于需要进度反馈和可续跑检查点的长时间任务。
运行要求
仅提供说明,不附带脚本。示例涉及 Python 的 concurrent.futures、pathlib 和 tqdm,安装部分列出 python-docx、openpyxl、python-pptx、reportlab 和 jinja2。还引用了带 batch_convert 工具的 office-mcp 服务器。

Batch Processor Skill

Overview

This skill enables efficient bulk processing of documents - convert, transform, extract, or analyze hundreds of files with parallel execution and progress tracking.

How to Use

  1. Describe what you want to accomplish
  2. Provide any required input data or files
  3. I'll execute the appropriate operations

Example prompts:

  • "Convert 100 PDFs to Word documents"
  • "Extract text from all images in a folder"
  • "Batch rename and organize files"
  • "Mass update document headers/footers"

Domain Knowledge

Batch Processing Patterns

Input: [file1, file2, ..., fileN]         │         ▼    ┌─────────────┐    │  Parallel   │  ← Process multiple files concurrently    │  Workers    │    └─────────────┘         │         ▼Output: [result1, result2, ..., resultN]

Python Implementation

python
from concurrent.futures import ProcessPoolExecutor, as_completedfrom pathlib import Pathfrom tqdm import tqdm
def process_file(file_path: Path) -> dict:    """Process a single file."""    # Your processing logic here    return {"path": str(file_path), "status": "success"}
def batch_process(input_dir: str, pattern: str = "*.*", max_workers: int = 4):    """Process all matching files in directory."""        files = list(Path(input_dir).glob(pattern))    results = []        with ProcessPoolExecutor(max_workers=max_workers) as executor:        futures = {executor.submit(process_file, f): f for f in files}                for future in tqdm(as_completed(futures), total=len(files)):            file = futures[future]            try:                result = future.result()                results.append(result)            except Exception as e:                results.append({"path": str(file), "error": str(e)})        return results
# Usageresults = batch_process("/documents/invoices", "*.pdf", max_workers=8)print(f"Processed {len(results)} files")

Error Handling & Resume

python
import jsonfrom pathlib import Path
class BatchProcessor:    def __init__(self, checkpoint_file: str = "checkpoint.json"):        self.checkpoint_file = checkpoint_file        self.processed = self._load_checkpoint()        def _load_checkpoint(self):        if Path(self.checkpoint_file).exists():            return json.load(open(self.checkpoint_file))        return {}        def _save_checkpoint(self):        json.dump(self.processed, open(self.checkpoint_file, "w"))        def process(self, files: list, processor_func):        for file in files:            if str(file) in self.processed:                continue  # Skip already processed                        try:                result = processor_func(file)                self.processed[str(file)] = {"status": "success", **result}            except Exception as e:                self.processed[str(file)] = {"status": "error", "error": str(e)}                        self._save_checkpoint()  # Resume-safe

Best Practices

  1. Use progress bars (tqdm) for user feedback
  2. Implement checkpointing for long jobs
  3. Set reasonable worker counts (CPU cores)
  4. Log failures for later review

Installation

bash
# Install required dependenciespip install python-docx openpyxl python-pptx reportlab jinja2

Resources

来源与署名

来源:claude-office-skills/skills位于batch-processor提交9c4c7d5

许可证: MIT

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架