PulpMiner

io.github.pulpminercomv0.1.0Updated Oct 11, 2026

Turn any public webpage into JSON with AI, save it as a live JSON API, and call it from your agent.

VerifiedSTDIODesktop onlyDeveloper ToolsWeb Search & Scraping

Overview

AI-generated overview

Lets an assistant scrape a public webpage, convert it to structured JSON with AI, save that setup as a reusable API, and fetch data from it.

What it does
PulpMiner exposes five tools: generate JSON from a URL with an optional JSON format sample, CSS selector and extra instructions; save a URL plus JSON format as a reusable REST API, optionally with a dynamic URL template and caching; fetch data from a saved API, passing variables for dynamic ones; list saved APIs; and validate the API key. Scraping and the AI conversion run on PulpMiner's backend, and options select the scraper and the LLM used.
When to use it
Useful when an assistant needs structured data from public web pages that have no API, or when a page-to-JSON setup should be saved once and reused as a callable endpoint. Also handy for checking which saved APIs exist and what variables they take.
Requirements
Runs locally over stdio via npx pulpminer-mcp (or a local Node.js build), or as a local Streamable HTTP endpoint with the --http flag. Requires a PulpMiner account and the PULPMINER_API_KEY environment variable, sent as the apikey header. Optional PULPMINER_BASE_URL and PULPMINER_TIMEOUT_MS. Network access to the PulpMiner API is needed.
Before you install
The API key is the only credential and is sent to PulpMiner's backend; keep it out of shared configs. Tool calls consume PulpMiner credits (generate 0.25, fetch 0.4, plus 0.1 for JavaScript rendering), so usage costs money. Saved APIs are created on the account and fetched data comes from third-party pages. The HTTP mode can accept a per-request x-pulpminer-api-key header and should be placed behind HTTPS if exposed.

Installation

In SourceWeft

  1. Open PulpMiner in the dashboard and add it to a workspace.
  2. Enable the server for the chats that should use its tools.

Desktop only via STDIO. STDIO servers start a local process, so they need the SourceWeft desktop host.

Other MCP clients

Follow the launch instructions in the repository.

README

PulpMiner MCP Server

Source: github.com/pulpminercom/pulpminer-mcp · Issues: github.com/pulpminercom/pulpminer-mcp/issues

An MCP server for PulpMiner. It lets AI assistants (Cursor, Claude Desktop, and any other MCP client):

  • Generate JSON from any web page with AI. PulpMiner scrapes the page and an LLM turns it into structured JSON, optionally in the exact shape you give it.
  • Save a page-to-JSON setup as a reusable REST API (static, or dynamic with {{ variables }} in the URL).
  • Fetch data from saved APIs and list them.

Tools

ToolWhat it doesBackend endpointCredits
pulpminer_generate_jsonScrape a URL and convert it to JSON with AI (optional jsonFormat sample, CSS selector, extra instructions)POST /external/ai/json0.25 (+0.1 renderJs)
pulpminer_save_as_apiSave URL + JSON format (+ optional dynamic URL template, cache) as an API; returns the API id and endpointPOST /external/ai/json/savefree
pulpminer_fetch_api_dataCall a saved API (GET, or POST with variables for dynamic APIs)GET/POST /external/{apiId}0.4 (+0.1 renderJs)
pulpminer_list_saved_apisList saved APIs (id, url, dynamic URL template and its variables)GET /external/api/urlfree
pulpminer_check_api_keyValidate the API key and show which credentials are configuredGET /external/n8n/authfree

Option values match the backend: scraper is scraper_1 (default) or scraper_2. llm is llm_2 (OpenAI GPT, default), llm_1 (Google Gemini), or llm_3 (Groq, only if enabled for your account).

Configuration

Env varRequiredDescription
PULPMINER_API_KEYyesYour API key from the PulpMiner dashboard → API Keys. Sent as the apikey header. It is the only credential any tool needs.
PULPMINER_BASE_URLnoDefault https://api.pulpminer.com
PULPMINER_TIMEOUT_MSnoDefault 300000 (scraping plus the LLM step can take a while)

Install & build

bash
git clone https://github.com/pulpminercom/pulpminer-mcp.gitcd pulpminer-mcpnpm installnpm run build

Run it:

bash
PULPMINER_API_KEY=... node dist/index.js          # stdioPULPMINER_API_KEY=... node dist/index.js --http   # Streamable HTTP on http://127.0.0.1:3333/mcp

Once published to npm, this becomes npx -y pulpminer-mcp.

Client configuration

Cursor (~/.cursor/mcp.json or .cursor/mcp.json)

json
{  "mcpServers": {    "pulpminer": {      "command": "npx",      "args": ["-y", "pulpminer-mcp"],      "env": { "PULPMINER_API_KEY": "your-api-key" }    }  }}

Claude Desktop (claude_desktop_config.json)

json
{  "mcpServers": {    "pulpminer": {      "command": "npx",      "args": ["-y", "pulpminer-mcp"],      "env": { "PULPMINER_API_KEY": "your-api-key" }    }  }}

Before it is published, use a local build: "command": "node", "args": ["/absolute/path/to/pulpminer-mcp/dist/index.js"].

Streamable HTTP clients

Start with --http (PORT, HOST env; it binds to 127.0.0.1:3333 by default) and point the client at http://127.0.0.1:3333/mcp. The server is stateless. Each request can send its own API key in the x-pulpminer-api-key header, which overrides PULPMINER_API_KEY. That makes it usable as a hosted multi-user endpoint (put it behind HTTPS).

json
{  "mcpServers": {    "pulpminer": {      "url": "http://127.0.0.1:3333/mcp",      "headers": { "x-pulpminer-api-key": "your-api-key" }    }  }}

Typical flow

  1. pulpminer_generate_json { url, jsonFormat? }: preview the JSON.
  2. pulpminer_save_as_api { url, jsonFormat, cache?, dynamicUrl? } returns apiId and endpoint.
  3. pulpminer_fetch_api_data { apiId, variables? }

Dynamic API example: dynamicUrl: "https://example.com/search?q={{ query }}", then fetch it with variables: { "query": "laptops" }.

Development

bash
npm test        # builds, then runs scripts/smoke-test.mjs against a local mock API (no real calls)npm run inspect # opens the MCP Inspector UInpx @modelcontextprotocol/inspector --cli node dist/index.js --method tools/list

Source: README.md at commit 5554217

Tools

0
Tool metadata has not been indexed yet.

Version history

1
  1. v0.1.0LatestOct 11, 2026