Agent Device

by callstack5ee37f30a8dfNo license4.9K starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated today

Automates Apple-platform apps (iOS, tvOS, macOS), Android devices, and Amazon Vega OS TV apps in Vega Virtual Devices. Use when navigating apps, taking snapshots/screenshots where supported, driving TV remotes, tapping, typing, scrolling, extracting UI info, collecting evidence, or planning agent-device CLI commands.

Instructions onlySoftware Development
AI-generated overview

Drives iOS, tvOS, macOS, Android and Vega OS apps via the agent-device CLI, using snapshots, refs and actions.

What it does
This skill instructs an agent to drive apps on Apple platforms, Android devices and Amazon Vega OS TV apps through the agent-device CLI. It covers opening an app, acting with press, click, fill, longpress, scroll and back, reading snapshot diffs, verifying expectations with wait, is, get or find, and closing the session. It also explains ref handling, selector preferences, screenshot fallback when accessibility data is sparse, and when to consult help topics.
When to use it
Use it for tasks that involve navigating apps, tapping, typing, scrolling, driving TV remotes, taking snapshots or screenshots where supported, extracting UI information, collecting evidence, or planning agent-device CLI commands.
Requirements
Requires the agent-device CLI to be available and a connected or reachable target device or virtual device. No scripts are shipped; the skill is instructions only.

agent-device

For a normal app-driving task, start immediately. Do not probe first with --help, --version, devices, appstate, snapshot, or screenshot:

bash
agent-device open <app> --foreground

That starts the session and returns the initial interactive snapshot with @refs.

Loop: act with press|click|fill|longpress <target> ... --settle, scroll <direction> --settle, or back --settle; continue from the printed diff, verify the named expectation (wait text "...", is, get, or find), then run agent-device close.

Reaching an off-screen target is one command, not a scroll-and-check loop: scroll down --until <selector> scrolls until that element is on screen, and scroll bottom runs to the end of the content. Repeated bare scroll down calls are the slow way to find something.

Copy refs byte-for-byte: @e12, @e12~s4 — keep the @ and any ~sN. Prefer current refs, then id/label/role selectors; coordinates are a last resort. If snapshot reports sparse/AX-unavailable, its refs and selectors are invalid: run agent-device screenshot, inspect the image, use coordinates, then retry snapshot -i after navigating. Otherwise run snapshot -i only when the diff lacks the next target.

Error output includes corrective hints; follow them instead of re-planning. Only when the task is specialized (for example gestures, scripting, TV, macOS, remote, or debugging) or a command shape is unclear, run agent-device help <topic>. agent-device --help lists topics, but is not a startup step.

Source and attribution

Source:callstack/agent-deviceinskills/agent-deviceat commit5ee37f3

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal