AInotate

io.github.dontpayfullv0.1.7Updated Oct 9, 2026

Annotated screenshots for AI agents: steps, Skitch-style arrows, boxes, labels, solid redaction.

Overview

AI-generated overview

Lets an assistant capture screens or web pages and return annotated screenshots with numbered steps, arrows, boxes and redactions.

What it does
AInotate captures a web page, window, screen or existing image, locates the elements the assistant refers to, and draws numbered steps, arrows, boxes, labels, highlights and magnifiers on them. It can also black out emails, tokens and other sensitive text before the image is shared, and it verifies the result before sending it. Outputs include single annotated images, step-by-step guides in Markdown, HTML or PDF, before/after plates and animations. It runs as a local MCP server over stdio and is also usable from the command line or as a Python library.
When to use it
Use it when an answer depends on where something is on screen: settings, menus, buttons behind icons, permission or payment screens, bug reports, tickets, docs and changelogs. It is meant for showing a control instead of describing its location, and for producing step-by-step visual guides.
Requirements
Runs locally as a Python package (Python 3.10+) via uvx or pipx, or through a Claude plugin or Desktop extension. Web capture needs a Playwright browser such as Chromium, which a doctor command can install. macOS is verified; Windows and Linux are experimental. No accounts, API keys or environment variables are declared.
Before you install
Automatic redaction of emails, phone numbers, card numbers, IBANs, API keys, tokens and JWTs is best-effort and can miss text drawn in images, canvas, unusual formats or embedded frames; always inspect the image before sharing. Blur and pixelation are not treated as redaction. Captures may include whatever is visible on screen or in a logged-in browser session, and images are saved locally by default.

Installation

In SourceWeft

  1. Open AInotate in the dashboard and add it to a workspace.
  2. Enable the server for the chats that should use its tools.

Desktop only via STDIO. STDIO servers start a local process, so they need the SourceWeft desktop host.

Other MCP clients

Follow the launch instructions in the repository.

README

[AInotate]

Annotated screenshots for AI agents: the agent shows you the button instead of describing where it is.

[Tests] [PyPI] [Python] [License: AGPL-3.0-or-later] [MCP server]

[Hacker News with Skitch-style arrows: Discussion, Post a link and a numbered Sign in step, in a browser window on a sunset gradient]

Ask an agent how to post a link on Hacker News and you usually get directions: the orange bar at the top, the last link after jobs. With AInotate it sends the picture above instead and says "click submit, where the blue arrow points".

It captures a web page, a window or a screenshot you already have, finds the elements you mean, draws numbered steps, arrows, boxes and labels where they hide the least, blacks out emails and tokens, and looks at the result before sending it. It works in Claude Code, Cowork, Claude Desktop, Codex, Gemini CLI, Cursor and any other MCP client, and from the command line or Python.

We built it at DontPayFull because our own agents kept answering "where is it?" with a paragraph. A picture with a number on the right button settles the question in a second, in a chat, a ticket or a guide.

What it is for

  • "Where is it?" Settings, menus, buttons that hide behind icons. The agent answers with the screen itself, the control boxed and numbered.
  • "Walk me through it." One image per step, in the order to follow; ainotate guide turns them into a Markdown, HTML or PDF guide.
  • "Your turn." A permission switch, a 2FA code, a payment or consent screen: the agent does not click for you. It shows exactly what to press and says what the click does.
  • Bug reports and tickets. What is wrong, outlined in red, what is right in green, emails and tokens blacked out before anyone sees them.
  • Docs and changelogs. Screenshots that point at the thing the text talks about, re-shot from the same spec when the UI changes; before and after plates for a change.

Set it up in your agent

1. Install, one of:

WithCommand
Homebrew (macOS)brew install dontpayfull/tap/ainotate
pipx or uv (Python 3.10+)pipx install "ainotate[all]" or uv tool install "ainotate[all]"
Claude plugin or Desktop extensionnothing to install first: step 2 brings AInotate along

Then ainotate doctor checks the machine and prints the exact fix for anything missing, including the one command that downloads Chromium for web capture. To update later: brew upgrade ainotate, pipx upgrade ainotate or uv tool upgrade ainotate; the plugin with /plugin marketplace update ainotate; the Desktop extension by opening the newer .mcpb from Releases.

2. Connect it to the agent you use:

AgentHow
Claude Code/plugin marketplace add dontpayfull/AInotate, then /plugin install ainotate@ainotate: skill and MCP server in one step
CoworkCustomize > Plugins > Add marketplace > dontpayfull/AInotate, then install AInotate
Claude Desktopdownload ainotate-<version>.mcpb from Releases and open it: a one-click extension
Codexcodex plugin marketplace add dontpayfull/AInotate, then codex plugin add ainotate@ainotate: skill and MCP server
Gemini CLIgemini extensions install https://github.com/dontpayfull/AInotate: skill and MCP server
Cursor[Add AInotate to Cursor] (needs uv), plus the skill below
VS Code, Cline, other MCP clientscommand ainotate, arguments mcp (stdio); also listed in the MCP Registry as io.github.dontpayfull/ainotate; agents can follow llms-install.md
json
{"mcpServers": {"ainotate": {"command": "ainotate", "args": ["mcp"]}}}

AInotate runs on your own computer: Cowork reaches it through the Claude desktop app, so keep the app open while a task uses it. The plugin starts the installed ainotate, or runs it with uvx when it is not installed.

The skill teaches an agent when and how to annotate: pick a source, find exact positions, annotate, read the image to verify, deliver. The plugin brings it along; without the plugin it ships with the package:

bash
ainotate install-skill             # Claude Code and Codex (~/.claude/skills, ~/.agents/skills)ainotate install-skill --zip ~/Desktop   # a ZIP for Claude Desktop, Cowork, claude.ai:                                         # Customize > Skills > Upload a skillnpx skills add dontpayfull/AInotate      # any agent the skills CLI knows (Cursor, Windsurf, ...)

Config paths for every OS, permissions and the tool list: skills/ainotate/references/mcp.md.

3. Ask. "Show me where to switch Wikipedia to dark mode." The agent captures the page, annotates it, checks it and sends the image. With the skill it also does this on its own whenever an answer depends on where something is on the screen: a "how do I" question, a visual bug, a before/after of a change.

Gallery

Every image below was made in one shoot call against a public site: open the page, find the elements, place the labels, draw, frame. The specs are in docs/images/specs.

[The same two arrows drawn in the four arrow styles: skitch, curved, straight and line]

Four arrow styles: skitch (default: straight when the way is clear, a gentle bend when it is long and diagonal, a curve around text), curved, straight, line

[Wikipedia on a phone-sized viewport with the menu and search buttons numbered]
Phone preset: mobile viewport, touch and user agent

Labels that stay off the content

[Left: labels at a fixed spot under each link cover the first headline. Right: AInotate puts the same labels in free space and points at the links with arrows]

A label dropped at a fixed spot next to its target usually lands on the text the reader needs. AInotate places every label itself:

  • It maps the page content first (text, icons, lines, images) and tries hundreds of spots per label. Empty space near the target wins; a spot over text is used only when the frame has no free room.
  • Arrows are priced too: a long arrow, or one that crosses text, another label or another mark, loses to a shorter, cleaner one. When the straight path crosses text, the arrow curves around it.
  • All labels are planned together, so the first one cannot take the only clean spot a later one needs.
  • Step badges sit on the corner of their box that covers the least.

The image is never enlarged to make room. If a label still has to cover something, AInotate says so in a warning, and label_at pins a label exactly where you want it.

From the command line

Save this as hn.json:

json
{"url": "https://news.ycombinator.com", "frame": {"chrome": "browser"}, "marks": [  {"type": "step", "n": 1, "target": {"text": "new", "exact": true},   "label": "Newest"},  {"type": "step", "n": 2, "target": {"text": "login"},   "label": "Sign in"}]}
bash
ainotate shoot hn.json --draft    # temp file; drop --draft to save

The output path is printed on stdout. Images are saved to ~/Pictures/AInotate unless you configure another folder. ainotate --help lists every command (capture, annotate, locate by OCR, grid, zoom, window capture with UI element targets, elements, guide, compare, animate, copy, install-skill); the full spec is in skills/ainotate/references/spec.md. Exit codes: 0 ok, 1 cannot save, 2 invalid spec, 3 cannot draw, 4 capture failed or an optional dependency is missing, 5 ambiguous text target, 6 target not found.

Python. The same pipeline is importable:

python
from ainotate.shoot import shootres = shoot({"url": "https://en.wikipedia.org/wiki/Screenshot",             "marks": [{"type": "box",                        "target": {"text": "View history"},                        "label": "Past edits"}]}, draft=True)print(res.paths[0], res.redactions)

Features

  • Marks: numbered step, box, arrow, click ripple, keys (keycaps), magnify (loupe), highlight, spotlight, text, solid redact, and decorative blur / pixelate.
  • Placement that reads well: labels go to free space near their target and are planned together; arrows are tapered Skitch-style shapes with a soft shadow that curve around text (see Labels that stay off the content). Warnings for overlapping labels, crossing arrows, more than 6 marks or labels over 4 words.
  • Web capture in one call: ainotate shoot opens a page (laptop, wide or phone preset), runs actions (click, fill, wait, scroll), measures each target, redacts, annotates and saves.
  • Any browser: its own Playwright Chromium, Firefox or WebKit, or attach over CDP to a Chromium you already use and are logged into (Chrome, Edge, Brave, Arc, BrowserOS). The preset is applied to that one tab and restored afterwards.
  • Any image: locate finds text by OCR (Apple Vision, Windows OCR or Tesseract); grid and zoom locate icons at full resolution. Screen, window and clipboard capture are built in.
  • Desktop apps on macOS: ainotate window ID --target measures buttons, switches and fields through the Accessibility API, so marks sit exactly on the control, icons included; ainotate elements --app NAME lists what it can find.
  • Frames: gradient backgrounds (auto from the screenshot, or presets), rounded corners, shadow, browser or window chrome, social aspect ratios.
  • Sharing: guides in Markdown, HTML and PDF; before/after plates; APNG and GIF animations; copy to the clipboard.
  • Strict specs: a typo stops the run with a list of every problem instead of saving a wrong or unredacted image. Clear exit codes.
  • Diacritics: bundled Inter font, so Romanian, German, French and other Latin-script labels render correctly everywhere.

Privacy

  • Redaction is solid. redact paints an opaque block. AInotate never uses blur to hide data: blur and pixelation can be reversed. blur and pixelate exist only to de-emphasize clutter, and a mark flagged "sensitive": true is refused for them.
  • Automatic redaction is on for web shots. shoot scans the live page (text and input values) for emails, phone numbers, card numbers, IBANs, API keys, tokens, JWTs, password and personal-data fields, and blacks them out. On images it runs on OCR when you pass --privacy auto. You can allow-list your own addresses or add regexes.
  • It is best-effort. Detection misses things: text drawn in images or canvas, unusual formats, data split across elements, embedded frames from other sites (reported, not redacted). Always look at the image before you share it. The skill makes this check mandatory for agents.
  • Local. Capture, OCR and rendering run on your machine. AInotate makes no network requests of its own beyond loading the pages you ask it to capture.

Platform support

macOSWindowsLinux
Annotate, frames, exportverifiedexperimentalexperimental
Web capture (Playwright, CDP)verifiedexperimentalexperimental
Screen and window captureverifiedexperimentalexperimental (X11)
UI elements (window --target, elements)verified (Accessibility)not yetnot yet
OCR (locate, text targets)verified (Vision)experimental (Windows OCR)experimental (Tesseract)
Clipboardverifiedexperimentalexperimental
MCP serververifiedexperimentalexperimental

Experimental means implemented and unit-tested with mocks, not yet verified on real Windows or Linux machines. Reports are welcome.

License

Copyright © 2026 DontPayFull. AInotate is free software under the GNU Affero General Public License v3.0 or later. If you run a modified AInotate as a network service, the AGPL requires you to offer its source to the users of that service.

Bundled and derived third-party work (the Inter font, arrow proportions from Arrowshot, a label-placement idea from github/awesome-copilot) and dependency licenses are listed in NOTICE.


Made with ❤️ by the DontPayFull team
Coupons & discount codes for 20,000+ stores

Source: README.md at commit a922a23

Tools

0
Tool metadata has not been indexed yet.

Version history

1
  1. v0.1.7LatestOct 9, 2026