
PulpMiner
io.github.pulpminercomv0.1.0Updated Oct 11, 2026
Turn any public webpage into JSON with AI, save it as a live JSON API, and call it from your agent.
Overview
Lets an assistant scrape a public webpage, convert it to structured JSON with AI, save that setup as a reusable API, and fetch data from it.
- What it does
- PulpMiner exposes five tools: generate JSON from a URL with an optional JSON format sample, CSS selector and extra instructions; save a URL plus JSON format as a reusable REST API, optionally with a dynamic URL template and caching; fetch data from a saved API, passing variables for dynamic ones; list saved APIs; and validate the API key. Scraping and the AI conversion run on PulpMiner's backend, and options select the scraper and the LLM used.
- When to use it
- Useful when an assistant needs structured data from public web pages that have no API, or when a page-to-JSON setup should be saved once and reused as a callable endpoint. Also handy for checking which saved APIs exist and what variables they take.
- Requirements
- Runs locally over stdio via npx pulpminer-mcp (or a local Node.js build), or as a local Streamable HTTP endpoint with the --http flag. Requires a PulpMiner account and the PULPMINER_API_KEY environment variable, sent as the apikey header. Optional PULPMINER_BASE_URL and PULPMINER_TIMEOUT_MS. Network access to the PulpMiner API is needed.
Installation
In SourceWeft
- Open PulpMiner in the dashboard and add it to a workspace.
- Enable the server for the chats that should use its tools.
Desktop only via STDIO. STDIO servers start a local process, so they need the SourceWeft desktop host.
Other MCP clients
Follow the launch instructions in the repository.
README
PulpMiner MCP Server
Source: github.com/pulpminercom/pulpminer-mcp · Issues: github.com/pulpminercom/pulpminer-mcp/issues
An MCP server for PulpMiner. It lets AI assistants (Cursor, Claude Desktop, and any other MCP client):
- Generate JSON from any web page with AI. PulpMiner scrapes the page and an LLM turns it into structured JSON, optionally in the exact shape you give it.
- Save a page-to-JSON setup as a reusable REST API (static, or dynamic with
{{ variables }}in the URL). - Fetch data from saved APIs and list them.
Tools
Option values match the backend: scraper is scraper_1 (default) or scraper_2. llm is llm_2 (OpenAI GPT, default), llm_1 (Google Gemini), or llm_3 (Groq, only if enabled for your account).
Configuration
Install & build
Run it:
Once published to npm, this becomes npx -y pulpminer-mcp.
Client configuration
Cursor (~/.cursor/mcp.json or .cursor/mcp.json)
Claude Desktop (claude_desktop_config.json)
Before it is published, use a local build: "command": "node", "args": ["/absolute/path/to/pulpminer-mcp/dist/index.js"].
Streamable HTTP clients
Start with --http (PORT, HOST env; it binds to 127.0.0.1:3333 by default) and point the client at http://127.0.0.1:3333/mcp. The server is stateless. Each request can send its own API key in the x-pulpminer-api-key header, which overrides PULPMINER_API_KEY. That makes it usable as a hosted multi-user endpoint (put it behind HTTPS).
Typical flow
pulpminer_generate_json { url, jsonFormat? }: preview the JSON.pulpminer_save_as_api { url, jsonFormat, cache?, dynamicUrl? }returnsapiIdandendpoint.pulpminer_fetch_api_data { apiId, variables? }
Dynamic API example: dynamicUrl: "https://example.com/search?q={{ query }}", then fetch it with variables: { "query": "laptops" }.
Development
Source: README.md at commit 5554217
Tools
0Version history
1- v0.1.0LatestOct 11, 2026


