PDF Text, Table and OCR Extractor

io.github.bkeller-researchv0.1.0Updated Oct 7, 2026

PDF URLs to per-page text, tables as rows, Markdown, metadata and OCR for scanned pages.

VerifiedStreamable HTTPWeb executableFiles & Storage

Overview

AI-generated overview

Extracts per-page text, tables, Markdown, metadata and OCR from PDF URLs through a remote MCP endpoint.

What it does
This remote MCP server takes PDF URLs and returns extracted content: per-page text, tables converted into rows, Markdown output, document metadata, and OCR for scanned pages. It is hosted at a remote endpoint, so the assistant connects to it rather than running anything locally.
When to use it
Useful when an assistant needs to read or summarize PDFs found online, pull tables out of reports, or handle scanned documents that require OCR. It fits document-processing workflows where the PDF is reachable by URL.
Requirements
A remote MCP endpoint at mcp.apify.com over streamable HTTP. No packages, environment variables or headers are declared, and no authentication is declared. The PDFs must be reachable by URL.
Before you install
The server fetches PDFs from URLs you provide, so those documents are sent to a third-party service. No credentials are declared, but confirm what the endpoint does with fetched files before using sensitive or private documents.

Installation

In SourceWeft

  1. Open PDF Text, Table and OCR Extractor in the dashboard and add it to a workspace.
  2. Enable the server for the chats that should use its tools.

Web executable via Streamable HTTP. Remote servers run from the web runtime once configured in a workspace.

Other MCP clients

Add this to your client's mcpServers config.

{
  "mcpServers": {
    "pdf-text-table-ocr-extractor": {
      "type": "http",
      "url": "https://mcp.apify.com/?tools=brenton8907/pdf-text-table-ocr-extractor"
    }
  }
}

Tools

0
Tool metadata has not been indexed yet.

Version history

1
  1. v0.1.0LatestOct 7, 2026