Batch Processor

作者 claude-office-skills9c4c7d5cd281MIT499 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫8 個月前更新

Process multiple documents in bulk with parallel execution

僅含說明Documents & Office
AI 產生的概覽

以平行方式批次處理大量文件,進行轉換、變形、擷取或分析,並追蹤進度。

功能
這個技能說明了批次處理大量文件的模式,包括以工作行程池平行執行、進度回報,以及透過檢查點讓長時間工作能在失敗後續跑。內容涵蓋檔案轉換、變形、內容擷取與分析,並提供批次處理與錯誤處理的 Python 範例程式碼。它只提供說明,產出的是處理後的文件或擷取結果,而不是單一文件。
適用情境
當你需要對數百個檔案套用相同操作時使用,例如把 PDF 轉成 Word、從圖片擷取文字、批次重新命名,或大量更新頁首頁尾。也適合需要進度回饋與可續跑檢查點的長時間工作。
執行需求
僅提供說明,未附帶指令碼。範例涉及 Python 的 concurrent.futures、pathlib 與 tqdm,安裝章節列出 python-docx、openpyxl、python-pptx、reportlab 與 jinja2。另引用帶有 batch_convert 工具的 office-mcp 伺服器。

Batch Processor Skill

Overview

This skill enables efficient bulk processing of documents - convert, transform, extract, or analyze hundreds of files with parallel execution and progress tracking.

How to Use

  1. Describe what you want to accomplish
  2. Provide any required input data or files
  3. I'll execute the appropriate operations

Example prompts:

  • "Convert 100 PDFs to Word documents"
  • "Extract text from all images in a folder"
  • "Batch rename and organize files"
  • "Mass update document headers/footers"

Domain Knowledge

Batch Processing Patterns

Input: [file1, file2, ..., fileN]         │         ▼    ┌─────────────┐    │  Parallel   │  ← Process multiple files concurrently    │  Workers    │    └─────────────┘         │         ▼Output: [result1, result2, ..., resultN]

Python Implementation

python
from concurrent.futures import ProcessPoolExecutor, as_completedfrom pathlib import Pathfrom tqdm import tqdm
def process_file(file_path: Path) -> dict:    """Process a single file."""    # Your processing logic here    return {"path": str(file_path), "status": "success"}
def batch_process(input_dir: str, pattern: str = "*.*", max_workers: int = 4):    """Process all matching files in directory."""        files = list(Path(input_dir).glob(pattern))    results = []        with ProcessPoolExecutor(max_workers=max_workers) as executor:        futures = {executor.submit(process_file, f): f for f in files}                for future in tqdm(as_completed(futures), total=len(files)):            file = futures[future]            try:                result = future.result()                results.append(result)            except Exception as e:                results.append({"path": str(file), "error": str(e)})        return results
# Usageresults = batch_process("/documents/invoices", "*.pdf", max_workers=8)print(f"Processed {len(results)} files")

Error Handling & Resume

python
import jsonfrom pathlib import Path
class BatchProcessor:    def __init__(self, checkpoint_file: str = "checkpoint.json"):        self.checkpoint_file = checkpoint_file        self.processed = self._load_checkpoint()        def _load_checkpoint(self):        if Path(self.checkpoint_file).exists():            return json.load(open(self.checkpoint_file))        return {}        def _save_checkpoint(self):        json.dump(self.processed, open(self.checkpoint_file, "w"))        def process(self, files: list, processor_func):        for file in files:            if str(file) in self.processed:                continue  # Skip already processed                        try:                result = processor_func(file)                self.processed[str(file)] = {"status": "success", **result}            except Exception as e:                self.processed[str(file)] = {"status": "error", "error": str(e)}                        self._save_checkpoint()  # Resume-safe

Best Practices

  1. Use progress bars (tqdm) for user feedback
  2. Implement checkpointing for long jobs
  3. Set reasonable worker counts (CPU cores)
  4. Log failures for later review

Installation

bash
# Install required dependenciespip install python-docx openpyxl python-pptx reportlab jinja2

Resources

來源與署名

來源:claude-office-skills/skills位於batch-processor提交9c4c7d5

授權條款: MIT

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架