Convert Documents To Markdown

by firecrawl261fc257d17cMIT22K starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 5 weeks ago

Convert Word (.doc, .docx), PowerPoint (.ppt, .pptx), Excel (.xls, .xlsx), OpenDocument (.odt, .ods, .odp), RTF, EPUB, CSV, and PDF files to GitHub-Flavored Markdown. Use when a task needs the contents of an office document, spreadsheet, presentation, ebook, or PDF you cannot read directly.

Instructions onlyDocuments & Office
AI-generated overview

Converts office documents, spreadsheets, presentations, ebooks and PDFs into GitHub-Flavored Markdown using the anydoc CLI.

What it does
This skill instructs an agent to run the anydoc command-line tool to turn Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV and PDF files into GitHub-Flavored Markdown. Output can be printed to stdout or written to a file with the -o flag, and input can also be read from stdin with an explicit format. It documents supported file extensions, exit codes, and the option to send scanned or image-only PDF pages to a hosted OCR service.
When to use it
Use it when a task requires the contents of an office document, spreadsheet, presentation, ebook or PDF that the agent cannot read directly. It is also useful for large documents where the Markdown should be written to a file and read in parts.
Requirements
Node 20 or higher and network access to fetch the anydoc package via npx; no installation is required. Hosted OCR for scanned or image-only PDFs requires network access to Firecrawl Parse and optionally an API key via --api-key or the FIRECRAWL_API_KEY environment variable. The skill ships no scripts; it is instructions only.

Convert documents to Markdown

Run the anydoc CLI. It needs Node 20+ and no install:

bash
npx -y @firecrawl/anydoc <file>              # Markdown to stdoutnpx -y @firecrawl/anydoc <file> -o out.md    # write to a filenpx -y @firecrawl/anydoc - --format csv < f  # read stdin

Rules:

  1. Supported inputs: .doc, .docx, .docm, .odt, .rtf, .epub, .pdf, .ppt, .pps, .pot, .pptx, .pptm, .ppsx, .ppsm, .odp, .xls, .xlsx, .xlsm, .xlsb, .ods, .csv.
  2. The format is detected from the file content. Pass --format <name> only when detection cannot work: CSV from stdin, or a missing or wrong extension.
  3. Exit codes: 0 success, 1 the document could not be converted, 2 usage error, 3 pages of a PDF need OCR. Failures print one anydoc: <message> line to stderr. The CLI never prompts.
  4. For a large document, write to a file with -o and read the parts you need instead of streaming everything into context.
  5. Scanned and image-only pages need OCR, which anydoc does not do, so the document exits 3. Rerun with --ocr hosted to send it to Firecrawl Parse. No signup needed. Pass --api-key or set FIRECRAWL_API_KEY for higher limits.
  6. Inside a Node, Python, or Rust codebase, prefer the library over shelling out: @firecrawl/anydoc on npm, firecrawl-anydoc on PyPI, anydoc on crates.io. Each exposes the same to_markdown / toMarkdown API.

Source and attribution

Source:firecrawl/anydocinskills/convert-documents-to-markdownat commit261fc25

License: MIT

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal