
ImageStep
dev.imagestepv0.1.0Updated Oct 3, 2026
The image pipeline for agents: presets, assets and jobs. Answers with asset ids and stable URLs.
Overview
Lets an assistant generate, transform and batch-process images through ImageStep presets and jobs, returning asset ids and URLs instead of image bytes.
- What it does
- Exposes ImageStep's image pipeline as MCP tools: generate images from prompts, run single atomic transforms (background removal, upscaling, colorizing, resizing and more), run saved versioned presets, check job progress, search assets, save presets and send feedback. Tools take asset ids, public URLs or local file paths and return asset ids, paths or signed URLs, never image bytes. A transform can run synchronously for one deterministic input or as a job for batches and AI operations.
- When to use it
- Useful when an assistant needs to create or edit images as part of a workflow: product photos, social-size variants, background removal, upscaling, or consistent character and product batches via presets. Also useful when results should stay out of the context window and be referenced by id or URL.
- Requirements
- An ImageStep API key, supplied as the IMAGESTEP_API_KEY environment variable for the local npm package or as an Authorization Bearer header for the hosted endpoint. The local package needs Node.js and npx; the remote endpoint needs no install. Local file paths work only over stdio; the hosted server accepts URLs or asset ids. Network access to the ImageStep API is required.
Installation
In SourceWeft
- Open ImageStep in the dashboard and add it to a workspace.
- Enable the server for the chats that should use its tools.
Web executable via Streamable HTTP. Remote servers run from the web runtime once configured in a workspace.
Other MCP clients
Add this to your client's mcpServers config.
{
"mcpServers": {
"mcp": {
"type": "http",
"url": "https://mcp.imagestep.dev/mcp"
}
}
}README
@imagestep/mcp
The image pipeline for agents: presets + assets + jobs, as MCP tools. Images never enter the context window — every tool takes and returns asset references (ids, sizes, public URLs).
Inputs are asset_ids, urls (public http(s), fetched and ingested by the service — never downloaded
by this server) or — when the server runs on
your machine over stdio — file_paths. Every tool that spends money has dry_run (exact price,
nothing created or uploaded — asset_ids are priced as themselves, file_paths / urls by how many)
and returns structured errors with retryable, param and the service's
requestId (quote it when reporting a failure), so a parameter mistake is refused before any
credit is charged.
The write tools (generate · transform · run_preset · save_preset · send_feedback) take an optional idempotency_key.
An MCP client that times out makes the agent call the tool again; with the same key
and the same arguments that second call returns what the first one answered — the job it created, the preset it saved — instead
of submitting (and charging for) another. The same key with different arguments is refused with
idempotency_key_reuse. Without a key, every call is a new submission.
Every tool carries MCP annotations, the hints a client reads to decide what needs your confirmation:
job_status and search_assets are readOnlyHint; the write tools are not idempotent (without a
key), and none is destructiveHint: every job writes new assets and never overwrites the one it read.
transform has two modes and the agent does not choose
Synchronous when all of these hold: exactly ONE local file or URL, a deterministic op, and
wait: true. Nothing is stored in the account — no asset id exists afterwards, because none
was created. Sub-second, and the account stays clean.
Where the result goes depends on which of the two servers you are talking to, because the two do not share a filesystem:
Both carry mimeType, width and height, so the next tool does not have to guess. The hosted
answer is a URL for the obvious reason and it took a bug to notice it: your agent is not on the
machine the server runs on, so a path into that container's /tmp is an answer it cannot read.
Under the hood that is the API's own ?response=url, which writes the bytes to a short-lived temp
object — still nothing in your asset catalogue.
As a job for everything else: several images, asset_ids, any AI op, or wait: false. You get
asset ids and public URLs, plus progress, cancellation and webhooks.
Which ops may go the first way is read from GET /api/v1/ops (syncEndpoint), never hard-coded —
a deterministic op added to the API is fast here without this package changing. If the catalogue
cannot be reached, the job path is used: slower is a better failure than wrong.
A urls input is judged the same on both paths: it must resolve to a public address, and it is
refused here — before anything leaves the process — not only by the service that would fetch it.
One input never gets two answers because of which transport happened to carry it.
Neither mode ever returns image bytes. An image in the context window is tokens the agent pays for and cannot read; the answer is always a path, an id or a URL.
Resources
The catalogue resources exist because an MCP-only client (Claude Desktop) cannot curl the API or open
the console, and "never invent a parameter" is a rule it can only follow if it can read them.
transform also carries a one-line-per-op summary — name(type=default), plus the values a parameter
takes when the catalogue lists them, what a required one means, and for an AI op the model's own
parameters object — generated from the same catalogue; when that could not be read, it falls back to
a built-in summary and says so. All five are
read live: a read that cannot reach the service fails rather than answering from a copy.
imagestep://agent-guidelines is the operating contract — read the op catalogue instead of
carrying a list, price with a dry run before spending, branch on retryable rather than the status
code, keep a batch consistent with a preset, and report a missing capability instead of working
around it. It is a resource rather than a tool because a tool is something an agent decides to
call and the rules are something it should have read, and it is fetched from the service on every
read, so a copy of this package released months ago still serves today's contract.
Connect
Get an API key at https://imagestep.dev/keys.
Claude Desktop (claude_desktop_config.json) — local, can upload files from disk:
Cursor (.cursor/mcp.json) — same shape:
Claude Code — local:
Remote (any client that speaks Streamable HTTP) — nothing to install:
The remote server is stateless: each request is served by a fresh server bound to the API key in
the Authorization header (Bearer or ApiKey). file_paths is disabled there; pass urls or
asset_ids.
The remote server waits at most 90 s per tool call (wait_seconds above that is clamped; the
default there is 90). It answers each call as one JSON body, so nothing reaches your client until the
tool returns, and Cloudflare's edge drops a response that has not started after 100 s — you would get
an HTML 524 with no job id while the job kept running. Past 90 s you get the job handle with
timedOut: true instead; poll job_status. Over stdio the wait is what you ask for (default 180 s,
max 600 s).
Try it
Upload ./product.jpg, remove the background, upscale it 2×, and give me the URL.
The agent calls transform (remove_bg, file_paths), then transform (upscale,
asset_ids = the previous output, parameters: {"scaleFactor": 2}), and reads publicUrl from
the last result. Three calls, zero image bytes in context.
Run it yourself
IMAGESTEP_BASE_URL points the server at another API host.
--http listens on 127.0.0.1 and answers only to Host: localhost:<port> / 127.0.0.1:<port> / [::1]:<port> —
anything else is a 403 before a key is looked at, so a web page that rebinds a hostname of its own to your machine cannot
drive it. --host 0.0.0.0 opens it to the network (the hosted image does that); add
--allowed-hosts mcp.example.com,… to hold it to the names you serve it under.
On SIGTERM / SIGINT --http stops taking connections, lets the calls in flight finish and exits 0 — at most 25 s, under
the hosted pod's 30 s grace. A call still waiting for its job answers at once with the job handle and timedOut: true, as
a wait that runs out does — poll job_status, don't submit again; for that the hosted server lets the service hold a
submit for at most 15 s before reading on. In the image node runs under tini, never as PID 1, which would ignore the signal.
Does an agent pick the right tool?
eval/ measures it: a minimal agent loop over this server's own tools/list, against a fake account
(no image is made, only model tokens are spent), tasks from the docs and recipes — a fifth of them
things ImageStep cannot do, where the right call is send_feedback. Tool choice, argument validity and
steps per model, with a spend cap:
It is the regression signal for a tool description, an annotation or an error message, and it spends
money, so it is run on purpose rather than by habit. Not a test and not in CI. When to run it, what it
costs, how to measure a change step by step, how to read and write tasks, and the numbers so far:
eval/README.md.
Source: packages/mcp/README.md at commit 42fa993
Tools
0Version history
1- v0.1.0LatestOct 3, 2026


