Browser

by ruvnet6051f6702b61No license74K starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated today

Web browser automation with AI-optimized snapshots for claude-flow agents

Instructions onlyProductivity & Workflow
AI-generated overview

Automates web browsers through the agent-browser CLI, using AI-optimized accessibility snapshots and element refs.

What it does
This skill documents a command-line workflow for driving a web browser: opening URLs, taking accessibility-tree snapshots, and interacting with pages via clicks, form fills, typing, scrolling and waits. It covers navigation, snapshots, interaction, information retrieval, waiting, sessions and selector strategies such as element refs, CSS selectors and semantic locators. It also describes integration points with Claude Flow, including MCP tools, memory storage and hooks.
When to use it
Use it when an agent needs to navigate websites, fill and submit forms, extract page data, or capture screenshots. It also suits parallel browsing tasks that require isolated sessions or shared authentication state.
Requirements
Requires the agent-browser CLI to be installed and available, plus network access to reach target sites. Optional Claude Flow integration uses npx @claude-flow/cli for memory and hooks. The skill ships no scripts; it is instructions only.

Browser Automation Skill

Web browser automation using agent-browser with AI-optimized snapshots. Reduces context by 93% using element refs (@e1, @e2) instead of full DOM.

Core Workflow

bash
# 1. Navigate to pageagent-browser open <url>
# 2. Get accessibility tree with element refsagent-browser snapshot -i    # -i = interactive elements only
# 3. Interact using refs from snapshotagent-browser click @e2agent-browser fill @e3 "text"
# 4. Re-snapshot after page changesagent-browser snapshot -i

Quick Reference

Navigation

CommandDescription
open <url>Navigate to URL
backGo back
forwardGo forward
reloadReload page
closeClose browser

Snapshots (AI-Optimized)

CommandDescription
snapshotFull accessibility tree
snapshot -iInteractive elements only (buttons, links, inputs)
snapshot -cCompact (remove empty elements)
snapshot -d 3Limit depth to 3 levels
screenshot [path]Capture screenshot (base64 if no path)

Interaction

CommandDescription
click <sel>Click element
fill <sel> <text>Clear and fill input
type <sel> <text>Type with key events
press <key>Press key (Enter, Tab, etc.)
hover <sel>Hover element
select <sel> <val>Select dropdown option
check/uncheck <sel>Toggle checkbox
scroll <dir> [px]Scroll page

Get Info

CommandDescription
get text <sel>Get text content
get html <sel>Get innerHTML
get value <sel>Get input value
get attr <sel> <attr>Get attribute
get titleGet page title
get urlGet current URL

Wait

CommandDescription
wait <selector>Wait for element
wait <ms>Wait milliseconds
wait --text "text"Wait for text
wait --url "pattern"Wait for URL
wait --load networkidleWait for load state

Sessions

CommandDescription
--session <name>Use isolated session
session listList active sessions

Selectors

Element Refs (Recommended)

bash
# Get refs from snapshotagent-browser snapshot -i# Output: button "Submit" [ref=e2]
# Use ref to interactagent-browser click @e2

CSS Selectors

bash
agent-browser click "#submit"agent-browser fill ".email-input" "[email protected]"

Semantic Locators

bash
agent-browser find role button click --name "Submit"agent-browser find label "Email" fill "[email protected]"agent-browser find testid "login-btn" click

Examples

Login Flow

bash
agent-browser open https://example.com/loginagent-browser snapshot -iagent-browser fill @e2 "[email protected]"agent-browser fill @e3 "password123"agent-browser click @e4agent-browser wait --url "**/dashboard"

Form Submission

bash
agent-browser open https://example.com/contactagent-browser snapshot -iagent-browser fill @e1 "John Doe"agent-browser fill @e2 "[email protected]"agent-browser fill @e3 "Hello, this is my message"agent-browser click @e4agent-browser wait --text "Thank you"

Data Extraction

bash
agent-browser open https://example.com/productsagent-browser snapshot -i# Iterate through product refsagent-browser get text @e1  # Product nameagent-browser get text @e2  # Priceagent-browser get attr @e3 href  # Link

Multi-Session (Swarm)

bash
# Session 1: Navigatoragent-browser --session nav open https://example.comagent-browser --session nav state save auth.json
# Session 2: Scraper (uses same auth)agent-browser --session scrape state load auth.jsonagent-browser --session scrape open https://example.com/dataagent-browser --session scrape snapshot -i

Integration with Claude Flow

MCP Tools

All browser operations are available as MCP tools with browser/ prefix:

  • browser/open
  • browser/snapshot
  • browser/click
  • browser/fill
  • browser/screenshot
  • etc.

Memory Integration

bash
# Store successful patternsnpx @claude-flow/cli memory store --namespace browser-patterns --key "login-flow" --value "snapshot->fill->click->wait"
# Retrieve before similar tasknpx @claude-flow/cli memory search --query "login automation"

Hooks

bash
# Pre-browse hook (get context)npx @claude-flow/cli hooks pre-edit --file "browser-task.ts"
# Post-browse hook (record success)npx @claude-flow/cli hooks post-task --task-id "browse-1" --success true

Tips

  1. Always use snapshots - They're optimized for AI with refs
  2. Prefer -i flag - Gets only interactive elements, smaller output
  3. Use refs, not selectors - More reliable, deterministic
  4. Re-snapshot after navigation - Page state changes
  5. Use sessions for parallel work - Each session is isolated

Source and attribution

Source:ruvnet/rufloin.claude/skills/browserat commit6051f67

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal