
HTTrack MCP Server
io.github.TechVentures-Studiov1.0.1Updated Oct 1, 2026
AI-agent website mirroring and offline archive browsing over MCP, powered by HTTrack.
Overview
Wraps HTTrack so an assistant can mirror websites into a local archive and then browse, read, and search the downloaded files.
- What it does
- This server exposes HTTrack's website-copying ability as MCP tools. An assistant can start a background mirror with mirror_site, poll progress with mirror_status, and then work with the archive using catalog, list_files, read_file, and search_files. Mirrors are stored per project under a shared archive directory, and text files can be read or searched by substring.
- When to use it
- Use it when an assistant needs an offline copy of a site for later reading, searching, or answering questions from, rather than live fetching. It suits research or documentation archiving where the same downloaded content is queried repeatedly.
- Requirements
- Runs as a local process; the documented setup uses Docker, with the container exposing the MCP server over streamable HTTP on a mapped port. A host directory must be mounted read-write as the archive. No authentication is declared.
Installation
In SourceWeft
- Open HTTrack MCP Server in the dashboard and add it to a workspace.
- Enable the server for the chats that should use its tools.
Desktop only via STDIO. STDIO servers start a local process, so they need the SourceWeft desktop host.
Other MCP clients
Follow the launch instructions in the repository.
README
HTTrack MCP Server
A FastMCP (Model Context Protocol) server that wraps the classic HTTrack website copier, so AI agents can mirror websites into a shared archive and then browse, read, and search the downloaded content — all over MCP.
Copyright (C) 2026 Tech Ventures VCC. Licensed under the GNU Affero General Public License v3.0 (see LICENSE).
Why
HTTrack is battle-tested for offline website copying, but it has no scripting API — only a
GUI and a CLI. This server makes it agent-native: an LLM agent can call mirror_site to
start a download in the background, poll mirror_status, and later answer questions from
the archive via catalog, list_files, read_file, and search_files.
Tools
Quick start
Prerequisites: Docker.
The container exposes the MCP server over streamable HTTP at http://<host>:9050/mcp.
Register in an MCP client
Example for MCPHub-style gateways (streamable-http):
Clients that speak streamable HTTP directly can connect to /mcp with no extra auth
(nobody is authenticated; run this on a trusted network, or put it behind your gateway's
auth). See HELP.md for full usage, examples, and troubleshooting.
Concurrency model
- One
httrackprocess per project, isolated by its own output directory — no shared state between projects, so multiple agents can mirror different sites in parallel. mirror_siteguards against double-mirroring the same project (in-process job table + a/procscan for httrack processes writing to that project directory).
CI
GitHub Actions runs on every PR and push: Python syntax check of server.py,
Dockerfile sanity assertions, and a Docker image build.
License
This program is free software: you can redistribute it and/or modify it under the terms of the GNU Affero General Public License as published by the Free Software Foundation, either version 3 of the License, or (at your option) any later version.
If you run a modified version of this server as a network service, AGPL §13 requires you to offer your users the corresponding source of that modified version.
Copyright (C) 2026 Tech Ventures VCC. HTTrack itself is GPL-licensed by Xavier Roche and contributors; this project only shells out to it and is not affiliated with it.
Source: README.md at commit 7feccd3
Tools
0Version history
1- v1.0.1LatestOct 1, 2026

