Doc Pipeline

作者 claude-office-skills9c4c7d5cd281MIT499 个星标收录于 2026年10月8日更新于 2026年10月8日仓库8个月前更新

Chain document operations into reusable pipelines

AI 生成的概览

将提取、转换、转换格式和生成报告等文档操作串联成可复用流水线。

功能
该技能说明如何构建文档处理流水线,把多个操作串联起来,并让数据在各阶段之间流动。它给出流水线架构、YAML 流水线定义示例,以及用于顺序和条件流水线的 Python 类。它还列出最佳实践和依赖安装命令,用于产出 DOCX 报告等结果。
适用场景
当任务包含多个按顺序执行的文档步骤时使用,例如从 PDF 提取文本、分析内容并生成报告。它适合希望获得可重复、可配置工作流,而不是单次文档操作的用户。
运行要求
仅为说明性内容,不附带脚本。文档列出 Python 依赖(python-docx、openpyxl、python-pptx、reportlab、jinja2),并提到 AI 分析阶段,但未说明凭据或网络访问要求。

Doc Pipeline Skill

Overview

This skill enables building document processing pipelines - chain multiple operations (extract, transform, convert) into reusable workflows with data flowing between stages.

How to Use

  1. Describe what you want to accomplish
  2. Provide any required input data or files
  3. I'll execute the appropriate operations

Example prompts:

  • "PDF → Extract Text → Translate → Generate DOCX"
  • "Image → OCR → Summarize → Create Report"
  • "Excel → Analyze → Generate Charts → Create PPT"
  • "Multiple inputs → Merge → Format → Output"

Domain Knowledge

Pipeline Architecture

Stage 1      Stage 2      Stage 3      Stage 4┌──────┐    ┌──────┐    ┌──────┐    ┌──────┐│Extract│ → │Transform│ → │ AI   │ → │Output││ PDF  │    │  Data  │    │Analyze│   │ DOCX │└──────┘    └──────┘    └──────┘    └──────┘     │           │           │           │     └───────────┴───────────┴───────────┘                 Data Flow

Pipeline DSL (Domain Specific Language)

yaml
# pipeline.yamlname: contract-review-pipelinedescription: Extract, analyze, and report on contracts
stages:  - name: extract    operation: pdf-extraction    input: $input_file    output: $extracted_text      - name: analyze    operation: ai-analyze    input: $extracted_text    prompt: "Review this contract for risks..."    output: $analysis      - name: report    operation: docx-generation    input: $analysis    template: templates/review_report.docx    output: $output_file

Python Implementation

python
from typing import Callable, Anyfrom dataclasses import dataclass
@dataclassclass Stage:    name: str    operation: Callable    class Pipeline:    def __init__(self, name: str):        self.name = name        self.stages: list[Stage] = []        def add_stage(self, name: str, operation: Callable):        self.stages.append(Stage(name, operation))        return self  # Fluent API        def run(self, input_data: Any) -> Any:        data = input_data        for stage in self.stages:            print(f"Running stage: {stage.name}")            data = stage.operation(data)        return data
# Example usagepipeline = Pipeline("contract-review")pipeline.add_stage("extract", extract_pdf_text)pipeline.add_stage("analyze", analyze_with_ai)pipeline.add_stage("generate", create_docx_report)
result = pipeline.run("/path/to/contract.pdf")

Advanced: Conditional Pipelines

python
class ConditionalPipeline(Pipeline):    def add_conditional_stage(self, name: str, condition: Callable,                                if_true: Callable, if_false: Callable):        def conditional_op(data):            if condition(data):                return if_true(data)            return if_false(data)        return self.add_stage(name, conditional_op)
# Usagepipeline.add_conditional_stage(    "ocr_if_needed",    condition=lambda d: d.get("has_images"),    if_true=run_ocr,    if_false=lambda d: d)

Best Practices

  1. Keep stages focused (single responsibility)
  2. Use intermediate outputs for debugging
  3. Implement stage-level error handling
  4. Make pipelines configurable via YAML/JSON

Installation

bash
# Install required dependenciespip install python-docx openpyxl python-pptx reportlab jinja2

Resources

来源与署名

来源:claude-office-skills/skills位于doc-pipeline提交9c4c7d5

许可证: MIT

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架