
Itamos MCP Tools
io.github.itamos-technologiav1.1.0更新于 Oct 9, 2026
Code-map tools for AI agents: see a repo's structure first, then edit only what matters.
概览
让编码助手以图结构感知代码库,从而只导航和编辑真正相关的代码片段。
- 功能
- 四个工具用结构化代码感知取代直接倾倒文件。master_architect 构建代码库图,并暴露拓扑、单文件骨架结构和片段级寻址(R14、R15)。read_file 默认返回骨架,可读取或编辑指定片段,并采用先验证后提交的方式(R17-R20);write_file 创建新文件并自动放置到项目中,且拒绝覆盖(R22、R23)。沙箱化 git 工具支持以 depth 1 克隆远程 URL、查看状态和提交,但不支持 push(R24)。web_skeleton 提供网页的 search、skeleton、read 和 click 操作(R28)。
- 适用场景
- 当助手处理无法放入上下文的大型代码库,需要从架构入手再缩小到模块,而不是用 grep 并读取整个文件时,适合添加(R6、R11)。也适合以更少 token 阅读长网页(R27、R45、R46)。README 建议用自己的模型和代码库在托管沙箱中试用(R52)。
- 运行要求
- 远程方式:在任意 MCP 客户端中添加托管端点作为连接器,会打开页面一键创建沙箱,无需账号(R59、R60)。自托管需要带 ZFS 的 Linux、Node.js 20 或更新版本及构建工具、git,以及 Chrome 或 Chromium(R64、R66、R67)。设置项包括 SANDBOX_ROOT、SANDBOX_TOTAL_SLOTS、SANDBOX_PORT、SANDBOX_TTL_MINUTES 和 SANDBOX_DATA_DIR(R96-R100)。可选摘要模型通过 ITAMOS_SUMMARIZER_URL 配置(R83)。
安装
在 SourceWeft 中
- 打开 控制台中的 Itamos MCP Tools,将其添加到工作区。
- 为需要使用其工具的对话启用该服务。
Web executable,通过 Streamable HTTP。 远程服务在工作区中配置后即可从网页运行时运行。
其他 MCP 客户端
把它添加到你客户端的 mcpServers 配置中。
{
"mcpServers": {
"itamos-mcp-tools": {
"type": "http",
"url": "https://mcp.itamos-technologia.com/mcp"
}
}
}README
Itamos MCP Tools
Built in Greece, for the world. Free and open source under AGPL-3.0. Commercial licences for closed-source use.
The problem
Every LLM coding assistant faces the same wall: codebases are too large to fit in context. The standard response is to dump files — grep for something promising, cat it, hope for the best. At 50k files this breaks. At 240k files it never worked.
We built four tools that give an LLM structured perception of a codebase instead of raw file access, plus a sandboxed git to bring code in. The result is a model that navigates code the way a senior engineer does — starting from the architecture, narrowing to the module, reading only the segment it needs.
Read the full paper: docs/PAPER.md: how each tool works, why it was designed that way, and how the hosted sandbox runs, with diagrams.
Tools
At least 97% fewer tokens than the usual shell workflow (grep, cat, run_cmd), measured against the best case where the model fixes the bug in one attempt. Real sessions save more, because the failed attempts common with raw shell access aren't counted.
How it works
Session model
Every client gets an isolated workspace. No shared state, no cross-session leakage. Workspaces are wiped after 10 minutes of inactivity.
master_architect navigation flow
The model is guided through four phases: scan to index the repo, topology to see the data-flow graph, bones to inspect a file, navigate or read_file to read the specific segment. The model never reads a file it has not first located in the graph.
read_file segment addressing
Files are parsed into named segments. A 5000-line file might have 40 segments. The model reads the skeleton, picks the segment it needs, reads that segment. Total context used is roughly 120 lines instead of 5000.
web_skeleton token reduction
Raw HTML of a modern web page runs 50,000 to 200,000 tokens. web_skeleton output is 500 to 3,000 tokens. The model reads the skeleton, picks the section id it needs, reads that section only.
Benchmarks
Verified in live use with Claude Opus 5.5 and Claude Sonnet. Any MCP-capable model can use the tools.
Don't take our word for it. Test the tools free in the hosted sandbox with your own model and your own repository, then post your results in Discussions → Benchmarks: model, task, tokens, time.
Found a bug? Open an issue with the steps to reproduce it. Every report makes the tools better.
Connecting
The sandbox server speaks standard MCP over HTTP POST with SSE support.
Compatible with any MCP client. Used in practice through Claude.ai and through direct API integration (our benchmark harness).
Hosted sandbox (free live alpha): add https://mcp.itamos-technologia.com/mcp as a connector in any MCP client. A page opens where you create your sandbox with one click, no account needed. Sandboxes are deleted after 10 minutes of inactivity, so don't use them for sensitive data.
Self-hosting
Requirements
- Linux with ZFS (required). Every sandbox is its own ZFS dataset with a hard size limit (quota) and compression, so no user can fill the disk for everyone else. The server checks this at startup and refuses to start without it.
- Node.js 20 or newer, plus build tools for the native modules (Debian/Ubuntu:
apt install build-essential python3). - git for the git tool, and Chrome or Chromium for web_skeleton.
Install
One npm install sets up everything, including the tools folder.
On Ubuntu 26.04 you can install the package from Releases instead: sudo apt install ./itamos-mcp-tools_1.1.0_all.deb. It sets up a service user, a systemd service, settings in /etc/itamos-mcp-tools/env and the command itamos-mcp-create-slots, then tells you the next two steps.
Create the sandbox slots
This creates slot_001 to slot_500 under the ZFS dataset tank/sandboxes (use your own pool name), each with a 2 GB quota and lz4 compression, owned by the user the server runs as. It is safe to run again.
Start
From the same machine, add http://localhost:4200/mcp as a connector in your MCP client. Local connections are identified by IP address and need no sign-in.
Going public
Firewall port 4200 and put an HTTPS reverse proxy (for example nginx) in front of it that sets X-Forwarded-For and X-Forwarded-Proto. Requests that arrive through the proxy use the anonymous one-click sign-in, so users on shared addresses (such as claude.ai) each get their own sandbox. Direct connections to port 4200 skip the sign-in, which is why the port must not be reachable from outside.
Text models (optional)
read_file and web_skeleton give untitled paragraphs and page sections a short title (5 to 10 words) from a small summarizer model, so the agent can tell parts apart without reading them. Without a summarizer everything still works; those parts are then named by their first line.
Any OpenAI-compatible chat endpoint works, set with ITAMOS_SUMMARIZER_URL (default http://127.0.0.1:8090): a single llama-server, Ollama, or the included router in front of several GPUs. The whole paragraph is always sent, never cut.
The hosted sandbox runs Gemma 4 E4B (Q4_0) as summarizer and EmbeddingGemma 300M as embedder on three GPUs (one 16 GB MI50 and two 8 GB V340 dies), each summarizer started like this:
--kv-unified: one shared KV pool, so each request uses only the tokens it needs.--cache-type-k q4_0 --cache-type-v q4_0(needs-fa on): a 4-bit KV cache, about 4.5 KB per token for this model.-cis the pool in tokens: 524,288 on the MI50, 262,144 on each V340 die (with 32 slots).
The router (router/llm-router.js, service itamos-llm-router) listens on 8090 (summarizer) and 8091 (embedder). Before sending a request it counts its tokens with the backend's own tokenizer, sends it to a GPU whose pool has room (preferring faster GPUs by weight), and makes it wait in line when none has. A GPU that fails is skipped for 15 seconds and the request is retried on another. Configure it with ROUTES (in the package: /etc/itamos-mcp-tools/router.env); GET /router/status shows each GPU's pool, reserved tokens and queue.
The embedder on 8091 serves other Itamos services; the tools don't use embeddings yet.
Settings
License
Dual licensed. Free under AGPL-3.0: use, modify and share, with source published for modified versions, including network use. Building it into a closed product? A commercial licence removes the AGPL obligations. Contact [email protected].
About
Designed and built by Konstantinos Karamperis, founder of Itamos Technologia in Trikala, Greece. A systems architect with a background in industrial engineering and infrastructure, building AI-native developer tools on AMD hardware with open-source inference stacks.
来源:README.md,提交 7a9ed90
工具
0版本历史
1- v1.1.0最新Oct 9, 2026

