Batch Processor

by claude-office-skills9c4c7d5cd281MIT499 starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 8 months ago

Process multiple documents in bulk with parallel execution

Instructions onlyDocuments & Office
AI-generated overview

Bulk-processes many documents in parallel, converting, transforming, extracting or analyzing files with progress tracking.

What it does
This skill describes patterns for processing large numbers of documents in bulk, including parallel execution with worker pools, progress reporting, and checkpointing so long jobs can resume after failures. It covers converting, transforming, extracting from, and analyzing files, and provides example Python code for batch processing and error handling. It is instruction-only and produces processed files or extraction results rather than a single document.
When to use it
Use it when you need to apply the same operation to hundreds of files, such as converting PDFs to Word, extracting text from images, batch renaming, or mass-updating headers and footers. It also suits long-running jobs that need progress feedback and resumable checkpoints.
Requirements
Instructions only; no scripts are shipped. The examples reference Python with concurrent.futures, pathlib and tqdm, and the installation section lists python-docx, openpyxl, python-pptx, reportlab and jinja2. It also references an office-mcp server with a batch_convert tool.

Batch Processor Skill

Overview

This skill enables efficient bulk processing of documents - convert, transform, extract, or analyze hundreds of files with parallel execution and progress tracking.

How to Use

  1. Describe what you want to accomplish
  2. Provide any required input data or files
  3. I'll execute the appropriate operations

Example prompts:

  • "Convert 100 PDFs to Word documents"
  • "Extract text from all images in a folder"
  • "Batch rename and organize files"
  • "Mass update document headers/footers"

Domain Knowledge

Batch Processing Patterns

Input: [file1, file2, ..., fileN]         │         ▼    ┌─────────────┐    │  Parallel   │  ← Process multiple files concurrently    │  Workers    │    └─────────────┘         │         ▼Output: [result1, result2, ..., resultN]

Python Implementation

python
from concurrent.futures import ProcessPoolExecutor, as_completedfrom pathlib import Pathfrom tqdm import tqdm
def process_file(file_path: Path) -> dict:    """Process a single file."""    # Your processing logic here    return {"path": str(file_path), "status": "success"}
def batch_process(input_dir: str, pattern: str = "*.*", max_workers: int = 4):    """Process all matching files in directory."""        files = list(Path(input_dir).glob(pattern))    results = []        with ProcessPoolExecutor(max_workers=max_workers) as executor:        futures = {executor.submit(process_file, f): f for f in files}                for future in tqdm(as_completed(futures), total=len(files)):            file = futures[future]            try:                result = future.result()                results.append(result)            except Exception as e:                results.append({"path": str(file), "error": str(e)})        return results
# Usageresults = batch_process("/documents/invoices", "*.pdf", max_workers=8)print(f"Processed {len(results)} files")

Error Handling & Resume

python
import jsonfrom pathlib import Path
class BatchProcessor:    def __init__(self, checkpoint_file: str = "checkpoint.json"):        self.checkpoint_file = checkpoint_file        self.processed = self._load_checkpoint()        def _load_checkpoint(self):        if Path(self.checkpoint_file).exists():            return json.load(open(self.checkpoint_file))        return {}        def _save_checkpoint(self):        json.dump(self.processed, open(self.checkpoint_file, "w"))        def process(self, files: list, processor_func):        for file in files:            if str(file) in self.processed:                continue  # Skip already processed                        try:                result = processor_func(file)                self.processed[str(file)] = {"status": "success", **result}            except Exception as e:                self.processed[str(file)] = {"status": "error", "error": str(e)}                        self._save_checkpoint()  # Resume-safe

Best Practices

  1. Use progress bars (tqdm) for user feedback
  2. Implement checkpointing for long jobs
  3. Set reasonable worker counts (CPU cores)
  4. Log failures for later review

Installation

bash
# Install required dependenciespip install python-docx openpyxl python-pptx reportlab jinja2

Resources

Source and attribution

Source:claude-office-skills/skillsinbatch-processorat commit9c4c7d5

License: MIT

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal