Core Agent Browser

by actionbook5c40d3ad7851No license1.5K starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 6 weeks ago

Internal support skill for agent-browser CLI workflows used by rust-learner, docs-researcher, and crate-researcher. Use only when browser automation is explicitly required.

Instructions onlyAI & Agents
AI-generated overview

Internal support skill documenting agent-browser CLI commands for browser automation workflows.

What it does
This skill is an internal reference for driving the agent-browser command-line tool. It lists commands for navigation, accessibility-tree snapshots, element interaction by reference, form filling, information retrieval, screenshots, waits, and semantic locators, plus a worked form-submission example. It also sets a priority order that places direct browser automation last, after other fetching approaches.
When to use it
Use it only when browser automation is explicitly required, such as when no pre-computed selectors exist for a target site, when interactive browser testing is needed, or when screenshots or form filling are required. It is not user-invocable and is not meant to be triggered automatically.
Requirements
Requires the agent-browser CLI to be installed and available, plus network access to reach target pages. It ships no scripts; it is instructions only.

Browser Automation with agent-browser

Priority Note

For fetching Rust/crate information, use this priority order:

  1. rust-learner skill - Orchestrates actionbook + browser-fetcher
  2. actionbook MCP - Pre-computed selectors for known sites
  3. agent-browser CLI - Direct browser automation (last resort)

Use agent-browser directly only when:

  • actionbook has no pre-computed selectors for the target site
  • You need interactive browser testing/automation
  • You need screenshots or form filling

Quick start

bash
agent-browser open <url>        # Navigate to pageagent-browser snapshot -i       # Get interactive elements with refsagent-browser click @e1         # Click element by refagent-browser fill @e2 "text"   # Fill input by refagent-browser close             # Close browser

Core workflow

  1. Navigate: agent-browser open <url>
  2. Snapshot: agent-browser snapshot -i (returns elements with refs like @e1, @e2)
  3. Interact using refs from the snapshot
  4. Re-snapshot after navigation or significant DOM changes

Commands

Navigation

bash
agent-browser open <url>      # Navigate to URLagent-browser back            # Go backagent-browser forward         # Go forwardagent-browser reload          # Reload pageagent-browser close           # Close browser

Snapshot (page analysis)

bash
agent-browser snapshot        # Full accessibility treeagent-browser snapshot -i     # Interactive elements only (recommended)agent-browser snapshot -c     # Compact outputagent-browser snapshot -d 3   # Limit depth to 3

Interactions (use @refs from snapshot)

bash
agent-browser click @e1           # Clickagent-browser dblclick @e1        # Double-clickagent-browser fill @e2 "text"     # Clear and typeagent-browser type @e2 "text"     # Type without clearingagent-browser press Enter         # Press keyagent-browser press Control+a     # Key combinationagent-browser hover @e1           # Hoveragent-browser check @e1           # Check checkboxagent-browser uncheck @e1         # Uncheck checkboxagent-browser select @e1 "value"  # Select dropdownagent-browser scroll down 500     # Scroll pageagent-browser scrollintoview @e1  # Scroll element into view

Get information

bash
agent-browser get text @e1        # Get element textagent-browser get value @e1       # Get input valueagent-browser get title           # Get page titleagent-browser get url             # Get current URL

Screenshots

bash
agent-browser screenshot          # Screenshot to stdoutagent-browser screenshot path.png # Save to fileagent-browser screenshot --full   # Full page

Wait

bash
agent-browser wait @e1                     # Wait for elementagent-browser wait 2000                    # Wait millisecondsagent-browser wait --text "Success"        # Wait for textagent-browser wait --load networkidle      # Wait for network idle

Semantic locators (alternative to refs)

bash
agent-browser find role button click --name "Submit"agent-browser find text "Sign In" clickagent-browser find label "Email" fill "[email protected]"

Example: Form submission

bash
agent-browser open https://example.com/formagent-browser snapshot -i# Output shows: textbox "Email" [ref=e1], textbox "Password" [ref=e2], button "Submit" [ref=e3]
agent-browser fill @e1 "[email protected]"agent-browser fill @e2 "password123"agent-browser click @e3agent-browser wait --load networkidleagent-browser snapshot -i  # Check result

Source and attribution

Source:actionbook/rust-skillsinskills/core-agent-browserat commit5c40d3a

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal