Skill: Chrome Automation (agent-browser)
Automate browser tasks in the user's real Chrome session via the agent-browser CLI.
Prerequisite: agent-browser must be installed and Chrome must have remote debugging enabled. See
references/agent-browser-setup.mdif unsure.
Core Principle: Reuse the User's Existing Chrome
This skill operates on a single Chrome process — the user's real browser. There is no session management, no separate profiles, no launching a fresh Playwright browser.
Always Start by Listing Tabs
Before opening any new page, always list existing tabs first:
This returns all open tabs with their index numbers, titles, and URLs. Check if the page you need is already open:
- If the target page is already open → switch to that tab directly instead of opening a new one. The user likely has it open because they are already logged in and the page is in the right state.
- If the target page is NOT open → open it in the current tab or a new tab.
Why This Matters
- The user's Chrome has their cookies, login sessions, and browser state
- Opening a new page when one is already available wastes time and may lose login state
- Many marketing platforms (social media dashboards, ad managers, CMS tools) require login — reusing an existing logged-in tab avoids re-authentication
Connection
Always use --auto-connect to connect to the user's running Chrome instance:
This auto-discovers Chrome with remote debugging enabled. If connection fails, guide the user through enabling remote debugging (see references/agent-browser-setup.md).
Chrome 144+ WebSocket-Only Fallback
Chrome 144+ can expose remote debugging from chrome://inspect/#remote-debugging as a WebSocket-only endpoint. In that state the page shows Server running at: 127.0.0.1:9222, but the traditional discovery URLs return 404:
Older agent-browser versions such as 0.27.x may fail with No running Chrome instance found even though Chrome is ready. First try the latest CLI without changing the global install:
If this works, use npx -y agent-browser@latest <command> for the rest of the browser task. If it fails with an engine warning or install error, upgrade Node to 24+ or install the latest agent-browser globally.
Common Workflows
1. Navigate and Interact
2. Extract Data from a Page
3. Replay a Chrome DevTools Recording
The user may provide a recording exported from Chrome DevTools Recorder (JSON, Puppeteer JS, or @puppeteer/replay JS format). See Replaying Recordings below.
Step-by-Step Interaction Guide
Taking Snapshots
Use snapshot -i to see all interactive elements with refs (@e1, @e2, ...):
The output lists each interactive element with its role, text, and ref. Use these refs for subsequent actions.
Step Type Mapping
How to Distinguish Input Types
- Standard input/textarea → use
fill - Contenteditable div / rich text editor (LinkedIn message box, Gmail compose, Slack, CMS editors) → click/focus first, then use
keyboard inserttext
Ref Lifecycle
Refs (@e1, @e2, ...) are invalidated when the page changes. Always re-snapshot after:
- Clicking links or buttons that trigger navigation
- Submitting forms
- Triggering dynamic content loads (AJAX, SPA navigation)
Verification
After each significant action, verify the result:
Replaying Recordings
Accepted Formats
-
JSON (recommended) — structured, can be read progressively:
-
@puppeteer/replay JS (
import { createRunner }) -
Puppeteer JS (
require('puppeteer'),page.goto,Locator.race)
How to Replay
- Parse the recording — understand the full intent before acting. Summarize what the recording does.
- List tabs first — check if the target page is already open.
- Navigate — execute
navigatesteps, reusing existing tabs when possible. - For each interaction step:
- Take a snapshot (
snapshot -i) to see current interactive elements - Match the recording's
aria/...selectors against the snapshot - Fall back to
text/..., then CSS class hints, then screenshot - Do not rely on ember IDs, numeric IDs, or exact XPaths — these change every page load
- Take a snapshot (
- Verify after each step — snapshot or screenshot to confirm
Iframe-Heavy Sites
snapshot -i operates on the main frame only and cannot penetrate iframes. Sites like LinkedIn, Gmail, and embedded editors render content inside iframes.
Detecting Iframe Issues
snapshot -ireturns unexpectedly short or empty results- Recording references elements not appearing in snapshot output
get text bodycontent doesn't match what a screenshot shows
Workarounds
-
Use
evalto access iframe content:Note: Only works for same-origin iframes.
-
Use
keyboardfor blind input: If the iframe element has focus,keyboard inserttext "..."sends text regardless of frame boundaries. -
Use
get text bodyto read full page content including iframes. -
Use
screenshotfor visual verification when snapshot is unreliable.
When to Ask the User
If workarounds fail after 2 attempts on the same step, pause and explain:
- The page uses iframes that cannot be accessed via snapshot
- Which element you need and what you expected
- Ask the user to perform that step manually, then continue
Handling Unexpected Situations
Handle Automatically (do not stop):
- Popups or banners → dismiss them (
find text "Dismiss" clickorfind text "Close" click) - Cookie consent dialogs → accept or dismiss
- Tooltip overlays → close them first
- Element not in snapshot → try
find text "..." click, or scroll to reveal withscroll down 300
Pause and Ask the User:
- Login / authentication is required
- A CAPTCHA appears
- Page structure is completely different from expected
- A destructive action is about to happen (deleting data, sending real content) — confirm first
- Stuck for more than 2 attempts on the same step
- All iframe workarounds have failed
When pausing, explain clearly: what step you are on, what you expected, and what you see.
Key Commands Reference
Known Limitations
- Iframe blindness:
snapshot -icannot see inside iframes. See Iframe-Heavy Sites. find textstrict mode: Fails when multiple elements match. Usesnapshot -ito locate the specific ref instead.fillvs contenteditable:fillonly works on<input>and<textarea>. For rich text editors, usekeyboard inserttext.evalis main-frame only: To interact with iframe content, traverse viadocument.querySelector('iframe').contentDocument...
Multi-Platform Operations
When the user requests an action across multiple platforms (e.g., "publish this article to Dev.to, LinkedIn, and X"), do NOT attempt all platforms in a single conversation. Instead, launch sequential subagents/workers, one per platform. Selecting this skill is explicit permission to use subagents/workers for this multi-platform workflow.
Why Subagents
Each platform operation consumes ~25-40K tokens (reference file + snapshots + interactions). Running 3-5 platforms in one context risks hitting the context limit and degrading late-platform accuracy. Each subagent/worker gets its own fresh context window.
How to Execute
- Prepare the content — confirm the post text, title, tags, and any platform-specific adaptations with the user.
- For each platform, launch a subagent/worker with a prompt that includes:
- The full content to publish
- Instructions to read the relevant reference file (e.g.,
Read /path/to/skills/chrome-automation/references/x.md) - Instructions to read the agent-browser skill file for command reference
- The specific task (post, comment, reply, etc.)
- Any platform-specific instructions (e.g., "use these hashtags on LinkedIn")
- Run subagents/workers sequentially (one at a time), because they all share the same Chrome browser via
--auto-connect. Parallel subagents/workers would cause tab conflicts. - After each subagent/worker completes, report the result to the user before launching the next one.
Prompt Template for Subagents
When NOT to Use Subagents
- Single platform — just do it directly in the current conversation.
- Read-only tasks (browsing, searching, extracting data) — context usage is lighter; a single conversation can handle 2-3 platforms.
Platform References
When automating tasks on specific platforms, consult the relevant reference document for page structure details, common operations, and known quirks:
For installation and Chrome setup instructions, see
references/agent-browser-setup.md.

