
ImageStep
dev.imagestepv0.1.0更新于 Oct 3, 2026
The image pipeline for agents: presets, assets and jobs. Answers with asset ids and stable URLs.
概览
让助手通过 ImageStep 的预设和作业生成、转换和批量处理图片,返回素材 ID 和 URL 而不是图片字节。
- 功能
- 把 ImageStep 的图片流水线暴露为 MCP 工具:根据提示词生成图片、执行单个原子转换(去背景、放大、上色、缩放等)、运行已保存的版本化预设、查询作业进度、搜索素材、保存预设和提交反馈。工具接受素材 ID、公开 URL 或本地文件路径,返回素材 ID、路径或签名 URL,从不返回图片字节。转换可对单个确定性输入同步执行,也可作为作业处理批量或 AI 操作。
- 适用场景
- 适合助手在工作流中需要创建或编辑图片的场景:产品图、社交尺寸变体、去背景、放大,或通过预设保持角色和产品批次一致。也适合希望结果不进入上下文窗口、仅以 ID 或 URL 引用的场景。
- 运行要求
- 需要 ImageStep API 密钥:本地 npm 包通过环境变量 IMAGESTEP_API_KEY 提供,托管端点通过 Authorization Bearer 请求头提供。本地包需要 Node.js 和 npx;远程端点无需安装。本地文件路径仅在 stdio 下可用;托管服务器接受 URL 或素材 ID。需要能访问 ImageStep API 的网络。
安装
在 SourceWeft 中
- 打开 控制台中的 ImageStep,将其添加到工作区。
- 为需要使用其工具的对话启用该服务。
Web executable,通过 Streamable HTTP。 远程服务在工作区中配置后即可从网页运行时运行。
其他 MCP 客户端
把它添加到你客户端的 mcpServers 配置中。
{
"mcpServers": {
"mcp": {
"type": "http",
"url": "https://mcp.imagestep.dev/mcp"
}
}
}README
@imagestep/mcp
The image pipeline for agents: presets + assets + jobs, as MCP tools. Images never enter the context window — every tool takes and returns asset references (ids, sizes, public URLs).
Inputs are asset_ids, urls (public http(s), fetched and ingested by the service — never downloaded
by this server) or — when the server runs on
your machine over stdio — file_paths. Every tool that spends money has dry_run (exact price,
nothing created or uploaded — asset_ids are priced as themselves, file_paths / urls by how many)
and returns structured errors with retryable, param and the service's
requestId (quote it when reporting a failure), so a parameter mistake is refused before any
credit is charged.
The write tools (generate · transform · run_preset · save_preset · send_feedback) take an optional idempotency_key.
An MCP client that times out makes the agent call the tool again; with the same key
and the same arguments that second call returns what the first one answered — the job it created, the preset it saved — instead
of submitting (and charging for) another. The same key with different arguments is refused with
idempotency_key_reuse. Without a key, every call is a new submission.
Every tool carries MCP annotations, the hints a client reads to decide what needs your confirmation:
job_status and search_assets are readOnlyHint; the write tools are not idempotent (without a
key), and none is destructiveHint: every job writes new assets and never overwrites the one it read.
transform has two modes and the agent does not choose
Synchronous when all of these hold: exactly ONE local file or URL, a deterministic op, and
wait: true. Nothing is stored in the account — no asset id exists afterwards, because none
was created. Sub-second, and the account stays clean.
Where the result goes depends on which of the two servers you are talking to, because the two do not share a filesystem:
Both carry mimeType, width and height, so the next tool does not have to guess. The hosted
answer is a URL for the obvious reason and it took a bug to notice it: your agent is not on the
machine the server runs on, so a path into that container's /tmp is an answer it cannot read.
Under the hood that is the API's own ?response=url, which writes the bytes to a short-lived temp
object — still nothing in your asset catalogue.
As a job for everything else: several images, asset_ids, any AI op, or wait: false. You get
asset ids and public URLs, plus progress, cancellation and webhooks.
Which ops may go the first way is read from GET /api/v1/ops (syncEndpoint), never hard-coded —
a deterministic op added to the API is fast here without this package changing. If the catalogue
cannot be reached, the job path is used: slower is a better failure than wrong.
A urls input is judged the same on both paths: it must resolve to a public address, and it is
refused here — before anything leaves the process — not only by the service that would fetch it.
One input never gets two answers because of which transport happened to carry it.
Neither mode ever returns image bytes. An image in the context window is tokens the agent pays for and cannot read; the answer is always a path, an id or a URL.
Resources
The catalogue resources exist because an MCP-only client (Claude Desktop) cannot curl the API or open
the console, and "never invent a parameter" is a rule it can only follow if it can read them.
transform also carries a one-line-per-op summary — name(type=default), plus the values a parameter
takes when the catalogue lists them, what a required one means, and for an AI op the model's own
parameters object — generated from the same catalogue; when that could not be read, it falls back to
a built-in summary and says so. All five are
read live: a read that cannot reach the service fails rather than answering from a copy.
imagestep://agent-guidelines is the operating contract — read the op catalogue instead of
carrying a list, price with a dry run before spending, branch on retryable rather than the status
code, keep a batch consistent with a preset, and report a missing capability instead of working
around it. It is a resource rather than a tool because a tool is something an agent decides to
call and the rules are something it should have read, and it is fetched from the service on every
read, so a copy of this package released months ago still serves today's contract.
Connect
Get an API key at https://imagestep.dev/keys.
Claude Desktop (claude_desktop_config.json) — local, can upload files from disk:
Cursor (.cursor/mcp.json) — same shape:
Claude Code — local:
Remote (any client that speaks Streamable HTTP) — nothing to install:
The remote server is stateless: each request is served by a fresh server bound to the API key in
the Authorization header (Bearer or ApiKey). file_paths is disabled there; pass urls or
asset_ids.
The remote server waits at most 90 s per tool call (wait_seconds above that is clamped; the
default there is 90). It answers each call as one JSON body, so nothing reaches your client until the
tool returns, and Cloudflare's edge drops a response that has not started after 100 s — you would get
an HTML 524 with no job id while the job kept running. Past 90 s you get the job handle with
timedOut: true instead; poll job_status. Over stdio the wait is what you ask for (default 180 s,
max 600 s).
Try it
Upload ./product.jpg, remove the background, upscale it 2×, and give me the URL.
The agent calls transform (remove_bg, file_paths), then transform (upscale,
asset_ids = the previous output, parameters: {"scaleFactor": 2}), and reads publicUrl from
the last result. Three calls, zero image bytes in context.
Run it yourself
IMAGESTEP_BASE_URL points the server at another API host.
--http listens on 127.0.0.1 and answers only to Host: localhost:<port> / 127.0.0.1:<port> / [::1]:<port> —
anything else is a 403 before a key is looked at, so a web page that rebinds a hostname of its own to your machine cannot
drive it. --host 0.0.0.0 opens it to the network (the hosted image does that); add
--allowed-hosts mcp.example.com,… to hold it to the names you serve it under.
On SIGTERM / SIGINT --http stops taking connections, lets the calls in flight finish and exits 0 — at most 25 s, under
the hosted pod's 30 s grace. A call still waiting for its job answers at once with the job handle and timedOut: true, as
a wait that runs out does — poll job_status, don't submit again; for that the hosted server lets the service hold a
submit for at most 15 s before reading on. In the image node runs under tini, never as PID 1, which would ignore the signal.
Does an agent pick the right tool?
eval/ measures it: a minimal agent loop over this server's own tools/list, against a fake account
(no image is made, only model tokens are spent), tasks from the docs and recipes — a fifth of them
things ImageStep cannot do, where the right call is send_feedback. Tool choice, argument validity and
steps per model, with a spend cap:
It is the regression signal for a tool description, an annotation or an error message, and it spends
money, so it is run on purpose rather than by habit. Not a test and not in CI. When to run it, what it
costs, how to measure a change step by step, how to read and write tasks, and the numbers so far:
eval/README.md.
来源:packages/mcp/README.md,提交 42fa993
工具
0版本历史
1- v0.1.0最新Oct 3, 2026


