
PDF Text Extractor
io.github.fernandoguiraud16-coderv1.0.0Updated Oct 6, 2026
Extract text, pages and metadata from PDFs in bulk, with OCR for scanned files.
Overview
Extracts text, pages and metadata from PDFs in bulk, including OCR for scanned files, through a hosted Apify MCP endpoint.
- What it does
- This remote MCP server exposes Apify's PDF Text Extractor so an assistant can pull text, page content and metadata out of PDF documents in bulk. It also applies OCR to scanned or image-only PDFs. It is part of a wider family of Apify data tools served through the same MCP endpoint, selected per tool.
- When to use it
- Use it when you need an assistant to read or convert many PDFs at once, including scanned documents that require OCR, without running extraction software locally. It suits document-processing and data-preparation workflows rather than interactive single-file reading.
- Requirements
- A remote MCP endpoint at mcp.apify.com; no local package is declared. An Apify account and an API token are needed, supplied as the APIFY_TOKEN environment variable in the documented examples. Network access to the Apify platform is required.
Installation
In SourceWeft
- Open PDF Text Extractor in the dashboard and add it to a workspace.
- Enable the server for the chats that should use its tools.
Web executable via Streamable HTTP. Remote servers run from the web runtime once configured in a workspace.
Other MCP clients
Add this to your client's mcpServers config.
{
"mcpServers": {
"pdf-text-extractor": {
"type": "http",
"url": "https://mcp.apify.com/?tools=fguiraud/pdf-text-extractor"
}
}
}README
Data Tools: clean web data for developers and AI agents
Ready-to-run Python examples for pay-per-use Apify Actors that turn search trends, news, competitor ads, documents, websites, YouTube, TikTok and Instagram videos and audio into clean JSON, Markdown, CSV and Excel.
🌐 Website: https://fernandoguiraud16-coder.github.io/data-tools/
Each example is a single short script plus a sample-output.json from a real cloud run, so you can see exactly what you get before running anything.
[Sample output of the TikTok Transcript Scraper: real results from a cloud run]
Examples
Also in the family (same engine as Audio & Video Transcription, so the audio-transcription example works with them too):
For AI agents
- MCP: every tool is listed in the official MCP registry as
io.github.fernandoguiraud16-coder/<tool>and served athttps://mcp.apify.com/?tools=fguiraud/<tool>. Seeexamples/mcp-agents. - Agent skills:
npx skills add fernandoguiraud16-coder/data-toolsinstalls 4 skills for Claude Code, Codex, Cursor and other coding agents. - llms.txt: a short index of all tools for LLMs.
Quick start
-
Create a free Apify account and copy your API token from Console → Settings → API & Integrations.
-
Install the client and set the token:
-
Run any example:
Pricing
All Actors are pay per result: you pay a fraction of a cent per item (term, article, ad, document, domain, video or audio minute), and failed or empty items are not charged. Apify's free plan includes monthly credits, enough to try every example. Exact prices are on each Actor's page.
Support
Found a bug or need a feature? Open an issue on the Actor's page in the Apify Store (replies within 48 hours), or an issue in this repository for problems with the examples.
License
The example code is MIT licensed: copy it into your own projects.
Source: README.md at commit 2c4f1e0
Tools
0Version history
1- v1.0.0LatestOct 6, 2026
