Translate Book Parallel

作者 reason-machines2384a003145a无许可证83 个星标收录于 2026年10月8日更新于 2026年10月8日仓库3个月前更新

Translate entire books (PDF/DOCX/EPUB) into any language using Claude Code parallel subagents with resumable chunked pipeline

AI 生成的概览

使用并行子代理和可续传的分块流程,将 PDF、DOCX 或 EPUB 整本书翻译成其他语言。

功能
该技能编排一条图书翻译流水线:先通过 Calibre 把 PDF、DOCX 或 EPUB 输入转换为 Markdown,再切分成约 6000 字符的分块,并用 manifest 中的 SHA-256 哈希进行跟踪,然后由并行子代理在各自独立的上下文中翻译每个分块。之后它会校验每个源分块都有对应的非空输出,合并结果并生成 output.md 以及 HTML、DOCX、EPUB 和 PDF 版本。翻译可按分块粒度续传,重新运行时已翻译的分块会被跳过。
适用场景
当你需要把整本书或长文档翻译成其他语言,并希望得到 EPUB、DOCX 或 PDF 等成品阅读格式时使用。它适合单次会话翻译会截断或丢失上下文的长文本,以及需要中断后继续的任务。
运行要求
需要 Calibre(ebook-convert)、Pandoc,以及 Python 包 pypandoc 和 beautifulsoup4,并且都应在 PATH 中可用。该技能仅为说明文档,本包中不含脚本,但会引用 convert.py、merge_and_build.py 等脚本;它还依赖代理的并行子代理及其所需的模型访问。

Translate Book (Parallel Subagents)

Skill by ara.so — Daily 2026 Skills collection.

A Claude Code skill that translates entire books (PDF/DOCX/EPUB) into any language using parallel subagents. Each chunk gets an isolated context window — preventing truncation and context accumulation that plague single-session translation.

Pipeline Overview

Input (PDF/DOCX/EPUB)  │  ▼Calibre ebook-convert → HTMLZ → HTML → Markdown  │  ▼Split into chunks (~6000 chars each)  │  manifest.json tracks SHA-256 hashes  ▼Parallel subagents (8 concurrent by default)  │  each: read chunk → translate → write output_chunk*.md  ▼Validate (manifest hash check, 1:1 source↔output match)  │  ▼Merge → Pandoc → HTML (with TOC) → Calibre → DOCX / EPUB / PDF

Prerequisites

bash
# 1. Calibre (provides ebook-convert)# macOSbrew install --cask calibre# Linuxsudo apt-get install calibre# Or download from https://calibre-ebook.com/
# 2. Pandocbrew install pandoc        # macOSsudo apt-get install pandoc # Linux
# 3. Python dependenciespip install pypandoc beautifulsoup4

Verify all tools are available:

bash
ebook-convert --versionpandoc --versionpython3 -c "import pypandoc; print('pypandoc ok')"

Installation

Option A: npx (recommended)

bash
npx skills add deusyu/translate-book -a claude-code -g

Option B: ClawHub

bash
clawhub install translate-book

Option C: Git clone

bash
git clone https://github.com/deusyu/translate-book.git ~/.claude/skills/translate-book

Usage in Claude Code

Once the skill is installed, use natural language inside Claude Code:

translate /path/to/book.pdf to Chinese
translate ~/Downloads/mybook.epub to Japanese
/translate-book translate /path/to/book.docx to French

The skill orchestrates the full pipeline automatically.

Supported Languages

CodeLanguage
zhChinese
enEnglish
jaJapanese
koKorean
frFrench
deGerman
esSpanish

Language codes are extensible — add new ones in the skill definition.

Running Pipeline Steps Manually

Step 1: Convert to Markdown Chunks

bash
python3 scripts/convert.py /path/to/book.pdf --olang zh

This produces inside {book_name}_temp/:

  • chunk0001.md, chunk0002.md, ... (source chunks, ~6000 chars each)
  • manifest.json (SHA-256 hashes for validation)
bash
# For EPUB inputpython3 scripts/convert.py /path/to/book.epub --olang ja
# For DOCX inputpython3 scripts/convert.py /path/to/book.docx --olang fr

Step 2: Translate (Parallel Subagents)

The skill handles this step — it launches 8 concurrent subagents per batch, each translating one chunk independently:

# Each subagent receives exactly this task:Read chunk0042.md → translate to target language → write output_chunk0042.md

Resumable: Already-translated chunks (valid output_chunk*.md files) are skipped on re-run.

Step 3: Merge and Build All Formats

bash
python3 scripts/merge_and_build.py \  --temp-dir book_name_temp \  --title "《Book Title in Target Language》"

Before merging, validation checks:

  • Every source chunk has a matching output file (1:1)
  • Source chunk hashes match manifest.json (no stale outputs)
  • No output files are empty

Outputs produced:

FileDescription
output.mdMerged translated Markdown
book.htmlWeb version with floating TOC
book.docxWord document
book.epubE-book format
book.pdfPrint-ready PDF

Project Structure

translate-book/├── SKILL.md                    # Claude Code skill definition (orchestrator)├── scripts/│   ├── convert.py              # PDF/DOCX/EPUB → Markdown chunks via Calibre HTMLZ│   ├── manifest.py             # SHA-256 chunk tracking and merge validation│   ├── merge_and_build.py      # Merge chunks → HTML → DOCX/EPUB/PDF│   ├── calibre_html_publish.py # Calibre wrapper for format conversion│   ├── template.html           # Web HTML template with floating TOC│   └── template_ebook.html     # Ebook HTML template└── README.md

How Manifest Validation Works

python
# scripts/manifest.py (conceptual usage)
# During convert.py — records source hashesmanifest = {    "chunk0001.md": "sha256:abc123...",    "chunk0002.md": "sha256:def456...",    # ...}
# During merge_and_build.py — validates before merging# 1. Check every chunk has a corresponding output_chunk# 2. Re-hash source chunks and compare against manifest# 3. Reject if any hash mismatches (stale/corrupt output)# 4. Reject if any output file is empty

If validation fails, the script auto-deletes stale output.md and re-merges from valid chunk outputs.

Real-World Example: Translate a Technical Book

bash
# 1. Install the skillnpx skills add deusyu/translate-book -a claude-code -g
# 2. Open Claude Code in your working directorycd ~/books
# 3. Say in Claude Code:# "translate clean-code.pdf to Chinese"
# Claude Code will:# - Run convert.py to split into chunks# - Launch 8 parallel subagents per batch# - Each subagent translates one chunk# - Validate all outputs via manifest# - Merge and build all formats
# 4. Outputs appear in:ls clean-code_temp/# chunk0001.md  chunk0002.md  ...  (source)# output_chunk0001.md  ...         (translated)# manifest.json# output.md# book.html# book.docx# book.epub# book.pdf

Resuming an Interrupted Translation

bash
# If translation is interrupted, just re-run the same command:# "translate clean-code.pdf to Chinese"
# The skill detects existing output_chunk*.md files# and skips already-translated chunks automatically.# Only missing or failed chunks are retried.

Changing Output Metadata After Translation

If you need to update the title, author, template, or image assets without re-translating:

bash
# Delete only the final artifacts (keeps translated chunks)cd book_name_temp/rm -f output.md book*.html book.docx book.epub book.pdf
# Re-run merge steppython3 ../scripts/merge_and_build.py \  --temp-dir . \  --title "《New Title》"

Do NOT delete chunk files — those are your translated content. Only delete final artifacts when changing metadata.

Troubleshooting

ProblemSolution
Calibre ebook-convert not foundInstall Calibre; ensure ebook-convert is in $PATH
Manifest validation failedSource chunks changed — re-run convert.py
Missing source chunkSource file deleted — re-run convert.py to regenerate
Incomplete translationRe-run the skill — resumes from last valid chunk
Changed title/template but output unchangedDelete output.md, book*.html, book.docx, book.epub, book.pdf then re-run merge_and_build.py
output.md exists but manifest invalidScript auto-deletes stale output and re-merges
PDF generation failsVerify Calibre has PDF output support; try ebook-convert --help
Empty output chunksRetry failed chunks; check API rate limits

Diagnosing Chunk Issues

bash
# Check which chunks are missing translationls book_temp/chunk*.md | wc -l          # total source chunksls book_temp/output_chunk*.md | wc -l   # translated chunks so far
# Find missing output chunksfor f in book_temp/chunk*.md; do  base=$(basename "$f" .md)  out="book_temp/output_${base}.md"  if [ ! -f "$out" ] || [ ! -s "$out" ]; then    echo "Missing: $out"  fidone
# Check manifestcat book_temp/manifest.json | python3 -m json.tool | head -30

Configuration Tips

  • Chunk size: ~6000 chars per chunk is the default. Smaller chunks = more parallelism but more API calls.
  • Concurrency: Default is 8 parallel subagents per batch. Adjust in SKILL.md if hitting rate limits.
  • Languages: Add new language codes to the skill triggers and translation prompt in SKILL.md.
  • Templates: Customize scripts/template.html and scripts/template_ebook.html for different HTML/ebook styling.

Key Design Principles

  1. Isolated context per chunk — each subagent starts fresh, preventing context overflow on long books
  2. Hash-based integrity — SHA-256 tracking catches stale or corrupt translated chunks before merging
  3. Resumable at chunk granularity — never re-translate what's already done
  4. Format-agnostic input — Calibre handles PDF/DOCX/EPUB normalization before the pipeline begins
  5. Multiple output formats — single pipeline produces HTML, DOCX, EPUB, and PDF simultaneously

来源与署名

来源:reason-machines/trending-skills位于skills/translate-book-parallel提交2384a00

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架