Doc Pipeline

by claude-office-skills9c4c7d5cd281MIT499 starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 8 months ago

Chain document operations into reusable pipelines

AI-generated overview

Chains document operations such as extract, transform, convert and report into reusable pipelines.

What it does
This skill describes how to build document processing pipelines that chain multiple operations, with data flowing between stages. It presents a pipeline architecture, a YAML pipeline definition example, and Python classes for sequential and conditional pipelines. It also lists best practices and dependency installation commands for producing outputs such as DOCX reports.
When to use it
Use it when a task involves several document steps in sequence, such as extracting text from a PDF, analyzing it, and generating a report. It suits users who want a repeatable, configurable workflow rather than a single document operation.
Requirements
Instructions only; no scripts are shipped. The document lists Python dependencies (python-docx, openpyxl, python-pptx, reportlab, jinja2) and references an AI analysis stage, but no credentials or network access are specified.

Doc Pipeline Skill

Overview

This skill enables building document processing pipelines - chain multiple operations (extract, transform, convert) into reusable workflows with data flowing between stages.

How to Use

  1. Describe what you want to accomplish
  2. Provide any required input data or files
  3. I'll execute the appropriate operations

Example prompts:

  • "PDF → Extract Text → Translate → Generate DOCX"
  • "Image → OCR → Summarize → Create Report"
  • "Excel → Analyze → Generate Charts → Create PPT"
  • "Multiple inputs → Merge → Format → Output"

Domain Knowledge

Pipeline Architecture

Stage 1      Stage 2      Stage 3      Stage 4┌──────┐    ┌──────┐    ┌──────┐    ┌──────┐│Extract│ → │Transform│ → │ AI   │ → │Output││ PDF  │    │  Data  │    │Analyze│   │ DOCX │└──────┘    └──────┘    └──────┘    └──────┘     │           │           │           │     └───────────┴───────────┴───────────┘                 Data Flow

Pipeline DSL (Domain Specific Language)

yaml
# pipeline.yamlname: contract-review-pipelinedescription: Extract, analyze, and report on contracts
stages:  - name: extract    operation: pdf-extraction    input: $input_file    output: $extracted_text      - name: analyze    operation: ai-analyze    input: $extracted_text    prompt: "Review this contract for risks..."    output: $analysis      - name: report    operation: docx-generation    input: $analysis    template: templates/review_report.docx    output: $output_file

Python Implementation

python
from typing import Callable, Anyfrom dataclasses import dataclass
@dataclassclass Stage:    name: str    operation: Callable    class Pipeline:    def __init__(self, name: str):        self.name = name        self.stages: list[Stage] = []        def add_stage(self, name: str, operation: Callable):        self.stages.append(Stage(name, operation))        return self  # Fluent API        def run(self, input_data: Any) -> Any:        data = input_data        for stage in self.stages:            print(f"Running stage: {stage.name}")            data = stage.operation(data)        return data
# Example usagepipeline = Pipeline("contract-review")pipeline.add_stage("extract", extract_pdf_text)pipeline.add_stage("analyze", analyze_with_ai)pipeline.add_stage("generate", create_docx_report)
result = pipeline.run("/path/to/contract.pdf")

Advanced: Conditional Pipelines

python
class ConditionalPipeline(Pipeline):    def add_conditional_stage(self, name: str, condition: Callable,                                if_true: Callable, if_false: Callable):        def conditional_op(data):            if condition(data):                return if_true(data)            return if_false(data)        return self.add_stage(name, conditional_op)
# Usagepipeline.add_conditional_stage(    "ocr_if_needed",    condition=lambda d: d.get("has_images"),    if_true=run_ocr,    if_false=lambda d: d)

Best Practices

  1. Keep stages focused (single responsibility)
  2. Use intermediate outputs for debugging
  3. Implement stage-level error handling
  4. Make pipelines configurable via YAML/JSON

Installation

bash
# Install required dependenciespip install python-docx openpyxl python-pptx reportlab jinja2

Resources

Source and attribution

Source:claude-office-skills/skillsindoc-pipelineat commit9c4c7d5

License: MIT

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal