Doc Pipeline

作者 claude-office-skills9c4c7d5cd281MIT499 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫8 個月前更新

Chain document operations into reusable pipelines

AI 產生的概覽

將擷取、轉換、格式轉換與產生報告等文件操作串成可重複使用的流程。

功能
此技能說明如何建構文件處理流程,把多個操作串接起來,並讓資料在各階段之間流動。它提供流程架構、YAML 流程定義範例,以及用於循序與條件式流程的 Python 類別。它也列出最佳實務與相依套件安裝指令,用來產出 DOCX 報告等成果。
適用情境
當任務包含多個依序執行的文件步驟時使用,例如從 PDF 擷取文字、分析內容並產生報告。它適合想要可重複、可設定的工作流程,而非單次文件操作的使用者。
執行需求
僅為說明內容,未附帶指令碼。文件列出 Python 相依套件(python-docx、openpyxl、python-pptx、reportlab、jinja2),並提到 AI 分析階段,但未說明憑證或網路存取需求。

Doc Pipeline Skill

Overview

This skill enables building document processing pipelines - chain multiple operations (extract, transform, convert) into reusable workflows with data flowing between stages.

How to Use

  1. Describe what you want to accomplish
  2. Provide any required input data or files
  3. I'll execute the appropriate operations

Example prompts:

  • "PDF → Extract Text → Translate → Generate DOCX"
  • "Image → OCR → Summarize → Create Report"
  • "Excel → Analyze → Generate Charts → Create PPT"
  • "Multiple inputs → Merge → Format → Output"

Domain Knowledge

Pipeline Architecture

Stage 1      Stage 2      Stage 3      Stage 4┌──────┐    ┌──────┐    ┌──────┐    ┌──────┐│Extract│ → │Transform│ → │ AI   │ → │Output││ PDF  │    │  Data  │    │Analyze│   │ DOCX │└──────┘    └──────┘    └──────┘    └──────┘     │           │           │           │     └───────────┴───────────┴───────────┘                 Data Flow

Pipeline DSL (Domain Specific Language)

yaml
# pipeline.yamlname: contract-review-pipelinedescription: Extract, analyze, and report on contracts
stages:  - name: extract    operation: pdf-extraction    input: $input_file    output: $extracted_text      - name: analyze    operation: ai-analyze    input: $extracted_text    prompt: "Review this contract for risks..."    output: $analysis      - name: report    operation: docx-generation    input: $analysis    template: templates/review_report.docx    output: $output_file

Python Implementation

python
from typing import Callable, Anyfrom dataclasses import dataclass
@dataclassclass Stage:    name: str    operation: Callable    class Pipeline:    def __init__(self, name: str):        self.name = name        self.stages: list[Stage] = []        def add_stage(self, name: str, operation: Callable):        self.stages.append(Stage(name, operation))        return self  # Fluent API        def run(self, input_data: Any) -> Any:        data = input_data        for stage in self.stages:            print(f"Running stage: {stage.name}")            data = stage.operation(data)        return data
# Example usagepipeline = Pipeline("contract-review")pipeline.add_stage("extract", extract_pdf_text)pipeline.add_stage("analyze", analyze_with_ai)pipeline.add_stage("generate", create_docx_report)
result = pipeline.run("/path/to/contract.pdf")

Advanced: Conditional Pipelines

python
class ConditionalPipeline(Pipeline):    def add_conditional_stage(self, name: str, condition: Callable,                                if_true: Callable, if_false: Callable):        def conditional_op(data):            if condition(data):                return if_true(data)            return if_false(data)        return self.add_stage(name, conditional_op)
# Usagepipeline.add_conditional_stage(    "ocr_if_needed",    condition=lambda d: d.get("has_images"),    if_true=run_ocr,    if_false=lambda d: d)

Best Practices

  1. Keep stages focused (single responsibility)
  2. Use intermediate outputs for debugging
  3. Implement stage-level error handling
  4. Make pipelines configurable via YAML/JSON

Installation

bash
# Install required dependenciespip install python-docx openpyxl python-pptx reportlab jinja2

Resources

來源與署名

來源:claude-office-skills/skills位於doc-pipeline提交9c4c7d5

授權條款: MIT

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架