Wonda CLI
Wonda CLI is a content creation toolkit for terminal-based agents. Use it to generate images, videos, music, and audio; edit and compose media; publish to social platforms; and research/automate across LinkedIn, Reddit, and X/Twitter.
Install
If wonda is not found on PATH, install it first. The recommended installs are the signed desktop installers (CLI + tray icon + always-on relay, zero extra steps): macOS brew install --cask degausai/tap/wonda-app (or the wonda-macos.pkg from releases), Windows winget/wonda-windows-setup.exe. The CLI-only channels below work everywhere and can add the desktop app later with wonda app install:
Setup
- Auth:
wonda auth login(opens browser, recommended) or setWONDA_API_KEYenv var - Verify:
wonda auth check
OAuth connector auth
Claude web and Cowork connectors use Wonda's OAuth 2.1 flow instead of a CLI
API key field. The connector signs in through Wonda in the browser, grants the
requested account access, and receives OAuth tokens bound to the Wonda API
resource. The server swaps those tokens to the account's internal API key only
inside Wonda, so agents and connector hosts never see the sk_... key. For the
CLI and local stdio MCP path, keep using wonda auth login or
WONDA_API_KEY.
Claude Cowork local relay
Claude Cowork (the desktop app) runs local MCP servers on the host, so it can
load the .mcpb bundle or a local stdio wonda-mcp config directly, WAB
writes included (verified 2026-07-07). Claude web cannot. The Wonda local
relay is the alternative path: it lets the REMOTE connector (web or Cowork)
run actions on the user's own Mac and residential IP without any local MCP
config:
- Open
https://wonda.sh/downloadwhile signed in and install the notarized Mac package. - Pair the relay with
wonda relay pairor the first-run browser handoff. This uses the existingcli-authflow with a relay-scopedwrelay_...credential stored in the macOS Keychain. Do not ask the user to paste an API key or device code. - Open
https://wonda.sh/setup, connect LinkedIn, X, and Reddit through the headful local WAB, then approve the Wonda connector once in Claude.
The engine policy is auto | my_machine | cloud. auto uses the local relay
when it is online and cloud otherwise. my_machine must not silently fall back:
if the relay is offline, ask whether to switch to cloud.
Organizations & spend context
Wondercat orgs are shared wallets with their own seats and billing. Members can spend from the org wallet (instead of their personal credits) by switching context:
wonda organizations list(aliases:wonda orgs list,wonda org list) — see every org you belong to with your role and seat plan in each.wonda use --org <slug>— sticky org context for this machine. SetsX-Wonda-Orgon every request; holds, charges, andwonda balanceroute through the org wallet.wonda use --personal— back to personal.wonda usage— spend-only usage summary (total + per-model + per-project breakdown) for a period (--month 2026-05, or--from/--to; defaults to the current month, UTC).--project <name>restricts the report to one project. In org context it reports org-wide usage including a per-member breakdown — admin/owner role required. Admins can also download a full Excel report from the org page on the web.
Projects (spend tagging)
Projects attribute spend to a named workstream for monitoring. Agents
should check the active project at task start (wonda use prints it) and
set one per task when the operator monitors spend by project:
wonda use --project <name>— sticky: every subsequent charge carries the project (inwonda usage, the API, and the org Excel report).wonda use --no-projectstops tagging; switching org/personal context clears the project automatically (projects are per-scope).--project <name>on any command — one-off override for that invocation.wonda project list|create|delete— manage the registry in the active scope. Org projects are created by org admins/owners only; personal projects are self-service. Tagging against a name that doesn't exist fails withunknown_project(no silent new buckets, so typos can't split the monitoring data).
wonda topup always tops up your personal wallet, regardless of
context. Topping up the org wallet (and configuring auto top-up) is
admin-only and happens on the web at /organizations/<slug>. If a
member runs out of org credits, the error tells them to ask an admin or
switch back to personal — they cannot top up the org wallet from CLI.
Roles inside an org are separate from the seat plan:
- Owner: the original creator. Cannot be demoted or kicked. Can transfer ownership to another member from the org page (rare).
- Admin: can invite (single or bulk via paste), kick, change roles, change seats, top up, configure auto top-up, change monthly limits.
- User: can only spend within the org wallet (subject to a per-member monthly limit if the admin set one).
A paid org seat (WONDA / WONDA_PREMIUM) grants the same paid feature access (skills, etc.) as a personal paid plan, but only while in org context. wonda use --personal falls back to the user's personal account plan.
Access tiers
Wonda is paid-only: every product surface requires a paid plan, including API calls, local platform reads, browser automation, media composition, file editing, diagnostics, generation, publishing, scraping, analysis, skills, relay operation, and cloud twin. Before any product command does work, the CLI performs a live entitlement check against Wonda. The check fails closed when access cannot be verified. New accounts get no product access until they subscribe, and a positive credit balance does not substitute for a plan.
If a command returns a 403 (paid_plan_required), subscribe at https://wonda.sh/account.
This applies to the local stdio MCP path too because MCP executes the same CLI. Platform cookies and data remain on-device, but the command must first receive a successful paid-access decision from GET /api/v1/auth/access. A relay-only installation may use its scoped wrelay_... credential for that entitlement probe. The scoped credential still cannot call ordinary account APIs.
Voice cloning
Clone a voice from a 10s+ audio clip and use it in TTS. Hard limit: 20 cloned voices per account. Cost: $1.50 per clone.
Use a cloned voice in TTS by passing the providerVoiceId from voice get as voiceId to /audio/speech:
7-day expiry: cloned voices that haven't been used in TTS within 7 days are automatically expired. Running TTS with a cloned voice automatically refreshes its expiry. Idle voices that lapse must be re-cloned ($1.50 again).
Credentials vault
Persist logins created on external platforms (Instagram, TikTok, Twitter, etc.) so they can be reused on the next run. Passwords are AES-256-GCM encrypted with a server-side key and only decrypted on get.
Fields: website (required — typed input like insta is canonicalized to instagram.com), username, email, password (required), metadata (arbitrary JSON). At least one of username / email must be present. Multiple records per (website, username) are allowed — dedupe on your side if you need to.
Event log: every credentials get/use, create, password rotate, and other updates are recorded as events on the credential (actor: cli | web | system). Use credentials events <id> or the web UI's history icon to audit. The event log is append-only and cascades on credential delete.
Global output flags
All commands support these output control flags:
--json— Force JSON output (auto-enabled when stdout is piped)--quiet— Only output the primary identifier (job ID, media ID, etc.) — ideal for scripting-o <path>— Download output to file (implies--wait)--fields status,outputs— Select specific JSON fields--jq '.outputs[0].media.url'— Filter JSON output with a jq expression
CLI announcements & deprecation warnings
On every command the CLI polls GET /api/v1/updates (anonymous, 1h cache in ~/.wonda/state.json) for active announcements: deprecation notices, incident heads-ups, upgrade prompts. Messages are printed to stderr only, so stdout/JSON stays clean for piping.
Per-request deprecation hints arrive as the standard Warning: 299 - "<message>" HTTP header and are surfaced to stderr by the CLI's HTTP client as [deprecated METHOD /path] <message>.
Silence both channels with WONDA_QUIET=1 (env var) or --quiet (flag). Disable just the network checks with WONDA_NO_UPDATE_CHECK=1.
WAB / Wonda Automation Browser (wonda wab)
Claude/Codex and other agent launches default to background. Use wonda wab start <persona> and normal platform commands without --visible; this matches the Dock menu's Send to background state. Show a WAB only for an intentional visible inspection or manual login. wab show/hide affects the current session and never changes the next launch default. wab config set <persona> visible true is a deliberate persistent opt-in. Routine reuse preserves visibility and the selected tab, and showing the window preserves that tab too.
1Password is included in local WAB profiles. WAB downloads the official stable 1Password browser extension, verifies its publisher signature, and pins it to the toolbar on first setup. Open the WAB with wonda wab show <persona>, then click 1Password to sign in. Each persona keeps its own 1Password session. The desktop app may require you to approve WAB as an additional browser; signing in directly in the extension also works. Cached extension files work offline. WAB checks cached extensions for updates at most once a day when a persona starts, including failed checks; changing the Chromium version triggers another compatibility check. If the extension has never installed successfully, each start retries installation. An already-running browser picks up an extension update on its next start. When the bundled extension changes, WAB refreshes its background service worker so Chromium cannot reuse code from the previous version. This preserves the persona’s 1Password sign-in and vault data. Startup and wab install prune unused cached versions older than seven days; versions used by a live WAB process are kept. Set WAB_ONEPASSWORD=0 before starting WAB to skip the bundled extension. Anonymous captures, scratch sessions, and cloud twins do not load it. A Wonda data path containing a comma also skips 1Password with a driver-log warning; use a WONDA_HOME path without commas to enable the extension. Enhanced Safe Browsing or administrator policies can also block automatic extension loading; WAB reports Chromium’s rejection in its driver log and preserves your browser settings.
The Wonda Automation Browser (WAB) is a premium stealth antidetect browser, hardened so platforms cannot fingerprint it as automation. wonda wab is the one command for the antidetect Chromium stack (an undetected Playwright fork). It has two faces:
- Authenticated sessions. One persistent headful Chromium per persona that holds signed-in sessions for LinkedIn, X, Reddit, and friends. The CLI spawns it on demand, lets it idle out, and routes platform reads/writes through it whenever a command runs
--via wab. Cookies live in the persona's Chromium profile, not in~/.wonda/config.json. - Anonymous capture.
wonda wab screenshot <url>,wonda wab record <url>, andwonda brand extractdrive an ephemeral Chromium with a fresh fingerprint, no persona, and no cookies. Screenshot supports responsive single, batch, and manifest capture. See the screenshot and record blocks below.
The mental model: you have accounts (one identity per platform). Each platform command routes to that account's cookies via either the flat JSON store (--via cookies, fast, no Chromium) or the account's persona (--via wab, live antidetect Chromium). A persona is the Chromium envelope that can hold multiple accounts under one fingerprint. In almost every case the persona is auto-created on first --via wab use, named after the account, so you never type a persona name.
The local wonda.mcpb Desktop Extension uses this same local WAB path from Claude Desktop or Claude Code: platform cookies stay on-device, reads use local cookies, and writes use the local WAB. Claude web and Cowork need the remote MCP connector instead.
Native login is the default for a new persona. wonda wab login <persona> <platform> opens a headful WAB window and you log in there. The session is minted INSIDE the WAB, so it is independent (logging out of the same account in an unrelated Chrome cannot revoke it) and the cookies are born under the WAB's own fingerprint, so session and browser identity stay coherent. A brand-new persona auto-created on first --via wab use chains straight into this flow on a TTY. After an X login, Wonda detects the signed-in screen_name and records it as the persona's X account binding. Existing bindings are never silently changed; a different detected handle produces a warning. Pasting cookies from another browser (wonda linkedin auth set, wonda x auth set, ...) still works and is the explicit fallback, but a hand-pasted li_at on a novel WAB fingerprint is the highest-risk shape.
Keep cron personas warm. A persona that continuously backs cookie-only read crons can use wonda wab config set <persona> idle-timeout off, followed by wonda wab start <persona>. The WAB then stays up and its existing 10-minute cookie sync keeps the flat files current. Use always-on only for cron-backing personas; wonda wab config unset <persona> idle-timeout restores the default 30-minute idle shutdown.
Local browser proxy (proxy_url). By default the local WAB dials direct (your own IP). Set wonda wab config set <persona> proxy_url managed to route the LOCAL browser through your account's minted twin proxy, so it shares the same egress as the cloud twin (useful for IP continuity or a VPN/office/CGNAT network). A literal socks5://…/https://… value is a manual override instead; unset clears it back to direct. The proxy is optional: if minting is disabled for the environment or unavailable, the browser falls back to a direct dial.
Lifecycle commands take an --account (e.g. wonda wab login <account> linkedin); the persona is auto-derived from the account name. wonda wab bind is the one place a persona is named explicitly: use it when one Chromium must host accounts that have different names per platform.
Drive any page (snapshot, act on @ref, snapshot). Click, type, scroll, and read ANY site in the persona's logged-in WAB, offscreen, without an action script. Use it for every site or step the platform commands do not cover; never drive the WAB with computer-use or screen clicks. The loop:
wonda wab goto <url> --persona <p>(or start from whatever the tab already shows).wonda wab snapshot -i --persona <p>prints the actionable elements as YAML, one per line, each with a ref:- Act on a ref:
wonda wab click @e7,wonda wab type @e41 "rust" --submit. - Snapshot again. Refs expire on every snapshot and navigation of the tab; a stale ref fails fast with
ref_not_foundand a hint to re-snapshot. Add--snapshotto any page-changing action (click, type, fill, press, hover, select, scroll, goto, back, forward, reload, upload) to print the fresh interactive snapshot in the same call.
Target grammar (one argument; quote it when it has spaces): @e5 / e5 / ref=e5 (snapshot ref; @f1e3 inside a frame), label=Email, placeholder=Search, 640,360 (viewport x,y, click and hover only), anything else is a Playwright selector ('text=Sign in', 'role=button[name="Post"]', CSS, xpath=//a). Prefer refs.
Every page command takes --persona (default: configured default account; a persona that does not exist on this machine is an error, never a fresh blank profile) and --tab (default default, the shared tab wonda wab start --open and MCP wab_open show; any other name is created on first use, so --tab research keeps side work off the page the user watches). Output is one short plain-text line per action even when piped (clicked @e5 (url: ...)); snapshot prints YAML, get prints the raw value, eval prints JSON. --json prints the driver response as one JSON line (e.g. {"ok":"click","url":...}), plus a snapshot field with --snapshot. Errors go to stderr with exit 1 and usually a Hint: line naming the next step. Text starting with - goes after --: wonda wab type @e4 -- "-5 degrees".
typevsfill:typetypes keystroke by keystroke with human timing, so use it wherever the site watches input (search with autocomplete, post and chat composers, login forms).fillsets the value instantly; use it for long text in plain form fields. Neither submits: add--submittotype, orwonda wab press Enter.evalruns in an isolated world by default (sees the DOM, not the page's JS globals, invisible to the site);--main-worldreads page globals but is detectable. The result must be JSON. Prefersnapshot/getfor reading.gotoinjects the stored LinkedIn/X/Reddit/Instagram session into a cold profile first and warns on stderr (without failing) onsession_revoked,rate_limited, orstale_cookies; a captcha or a page that never loads exits 1. Page-changing commands that land on those platforms flush the refreshed cookies to disk like every other WAB action.- Platform writes (posts, DMs, connects) still go through the platform commands, which carry recipient guards, rate limits, and audit logs. Use page commands for everything around them.
Scroll a page like a person (browse). wonda wab browse [url] --persona <persona> loads a page in the persona's WAB and scrolls it: a pause to look at the page, then scroll, pause, scroll, for --scrolls times (default 5). --scrolls accepts 1-200; a value outside that range is a hard error, not clamped, and it is rejected before a browser is launched. --first-wait (default 10s) is the pause before the first scroll; --wait (default 5s) is the pause between scrolls; both are jittered +/-30% because a precisely repeated interval is itself a fingerprint, and both are floored at 250ms so a very small value is not jittered down to zero — --first-wait 0 --wait 0 still pauses ~250ms per wait, not 0, so with --scrolls 200 that floor alone adds up to roughly 50s. It only scrolls — no clicking, no engagement — so it works on any site.
Progress is measured on the element that was actually scrolled: the viewport-filling overflow container when the site has one (LinkedIn's feed lives in <main id="workspace">, where window.scrollY never moves at all), otherwise the document scroller. Scrolling stops early once that element stops advancing across 2 consecutive scrolls, reported as (reached bottom) in the printed summary, so a short page does not grind against the bottom. An infinite feed keeps going for the full --scrolls.
--persona falls back to your configured default account exactly like the other wonda wab commands, and an invalid persona name is rejected up front instead of silently creating an empty logged-out profile. With a url, browse navigates in its OWN isolated tab (like every other WAB write) and reports the page actually browsed (after any redirect). Omitting the url dispatches to the persona's shared default tab instead — the driver's initial page, the same tab id wonda wab show/hide reference — scrolling whatever that tab already shows. wonda wab login does NOT leave anything there: it opens and navigates its own separate <platform>-login tab, so a persona that was just logged in still has an untouched default tab (often still about:blank). wonda wab show/hide don't navigate the default tab either — they only toggle the window's on/off-screen visibility — so neither one "opens a page" there. What actually leaves a page on default is something that navigates it directly, e.g. wonda wab start --open <platform|url> (also what the MCP wab_open tool calls). This is NOT a way to continue a page from a previous wonda wab browse <url> run: that run navigated in its own isolated tab, which the no-url form never sees. If the default tab has no page open (about:blank) — the common case right after logging a new persona in, since login's tab is separate — the command fails and asks for a url rather than reporting a successful scroll of nothing.
browse prints a plain-text summary, not JSON, so the global --json / --fields / --jq flags do not apply to it.
Anonymous PNG capture (screenshot). wonda wab screenshot has a compatibility-preserving persona mode and a new anonymous URL mode.
wonda wab screenshot [persona] still captures the persona's already-open tab without surfacing the window. Its existing flags retain their meaning: --tab selects the tab, --full-page captures the scrollable page, and --output writes a file. With --json and no output file it still returns {path, base64, mimeType} so MCP and existing automation receive the inline PNG unchanged. --persona <name> works in place of the positional (the same flag every page command takes), and --target <target> crops to one element using the page-command target grammar (@e12 from wab snapshot, label=..., or a selector); --target conflicts with --full-page.
An absolute http:// or https:// argument selects anonymous mode. It launches a fresh ephemeral Chromium with no persona, cookies, or persistent state:
Anonymous defaults are --viewport 1280x720, --scale 1, --wait-until networkidle, --delay 0, and --animations allow. --wait-until accepts load, domcontentloaded, or networkidle. Wonda automatically waits for document.fonts.ready; --wait-for <selector> waits for a visible element, and --delay <duration> adds a final settling delay. --inject-js <file> runs after navigation in an async IIFE, so top-level await works.
For deterministic motion, use --animations disabled to finish finite animations and cancel infinite ones at capture, or --freeze-at 800ms to seek Web Animations to an exact timeline offset and pause them. The two options conflict. --selector '.hero-stage' captures the first matching element. --clip x,y,width,height captures an exact document-coordinate rectangle in CSS pixels. The rectangle may extend beyond the viewport but must fit within the rendered page. --selector, --clip, and --full-page are mutually exclusive capture modes.
Pass several absolute URLs to capture all of them. Or pass relative path, query, or quoted hash arguments after one absolute base URL; when any relative target is present, the first URL is only the resolution base and is not captured separately. Add --viewports for a viewport matrix. Chromium launches once for the whole batch:
For a single capture without --output-template, --output is the PNG path. For a batch, manifest, or any template-driven capture, it is an output directory. --output-template supports {index}, {route}, and {viewport}, and expanded paths must be unique. {route} is sanitized; overlong expanded path components are shortened with a stable hash suffix. --viewport and --viewports conflict. Without output flags, one URL writes screenshot-<timestamp>.png in the current directory and a batch writes under screenshots-<timestamp>/. A template without --output is relative to the current directory, or to the manifest directory when it comes from the manifest.
A strict JSON manifest describes the same matrix:
Run it with wonda wab screenshot --manifest visual-qa.json --output tmp/qa. Use urls instead of url for independent absolute targets; routes resolves relative to url. Manifest keys mirror the anonymous scalar flags in camel case, including freezeAt, selector, clip, and fullPage. Explicit CLI flags override manifest values.
Anonymous --json returns {ok, captures: [...]} without embedding PNG bytes. Every capture reports url, finalUrl, viewport, dpr, pageDimensions, consoleErrors, failedRequests, and fontsLoaded. A successful capture also reports path, which means the PNG was written there. A failed capture omits path, carries error, and keeps any diagnostics collected before the failure. The browser continues the remaining matrix after an item fails and exits nonzero when any capture failed. Use this output for visual QA diagnostics. Non-JSON output prints only successfully generated PNG paths.
Anonymous video capture (record). wonda wab record <url> records a URL to webm in an ephemeral Chromium (fresh fingerprint each call, no persona, no cookies). Use it for cookie-banner-gated pages (Notion public shares, pdf.js renders, any site where bare Playwright trips a bot check) and marketing demo capture.
The --inject-js file is wrapped in an async IIFE so top-level await works. It runs AFTER domcontentloaded + networkidle + 400 ms paint settle, BEFORE the duration timer starts. Any await inside counts against the recording window. Use it for dark-theme injection, cookie-banner removal, scroll animations, anything that needs to happen in page context.
Node.js requirement: wonda needs Node >= v20 on PATH. Brew users get it via the node dependency; npm users have it by definition; install.sh users may need brew install node (or any Node distribution). If Node is missing, wonda wab install fetches a private copy into ~/.wonda/node/.
Cookie cloud backup. On by default (opt out per machine with wonda wab backup disable). The WAB driver pushes the synced cookie JSON for each bound platform to the wondercat backend after every wab → disk sync and graceful shutdown; auto-push no-ops when no api_key is configured. Encrypted at rest server-side (AES-256-GCM) when SOCIAL_COOKIES_KEY is set, else plaintext jsonb; the wire payload is always plaintext because the server holds the key. Cookie values are never printed by list/status/port commands.
Recovery is guarded. wonda wab backup pull <account> and wonda wab cookies port <platform> <persona> [account] refuse to overwrite a non-empty or newer local cookie file unless --force is passed. Forced writes create a hidden .before-pull-* backup first. Use --dry-run to inspect planned writes. When the backend exposes multiple device rows for the same platform/persona/account, use wonda wab cookies port ... --from-device <id|label> so the source row is explicit.
Current backend compatibility: legacy servers still expose one last-write-wins row per (account, platform, persona, account_label). Newer servers may include device_id, device_label, source, status, generation, and provenance; the CLI displays those fields when present and shows legacy rows as device legacy.
Source lives at cli/wondercat/wab/. The driver is launch.mjs and per-platform action scripts under actions/<platform>/.
WAB reads fail early when the selected browser profile has no live platform session cookie. For LinkedIn, X, Reddit, and Instagram, the error includes wonda wab login <persona> <platform> instead of surfacing an unexplained platform 401/403. This preflight is read-only; writes retain their existing error handling and native login itself is unaffected.
Per-command transport (--via). linkedin, x, and reddit commands take:
--via cookies|wab:cookiesreads the flat per-account JSON store (fast, no Chromium);wabroutes through the account's persona Chromium (cookies + TLS fingerprint inherit from a real browser session). An unsupported value errors loudly rather than silently downgrading.--via public: paid public-data API where a command explicitly supports it. For LinkedIn this avoids logged-in cookies and WAB profile reads, and uses the public scrape task route forwonda linkedin profileandwonda linkedin enrich.--account <name>: which on-disk identity to use (cookie filename / persona). Persona resolution is implicit: the first--via wabuse auto-creates a persona named after the account and (on a TTY) chains straight into login.
Defaults differ for reads vs writes. Read commands (profile, posts, search, timeline, etc.) default to cookies (direct API), because that path is fast and detection-safe. Write / engagement commands (post, comment, like, follow, connect, message, mute, repost, delete) default to wab, because the cookie-API path triggers anti-abuse heuristics on LinkedIn / X / Reddit at any meaningful volume. Pass --via cookies to a write command if you explicitly want the legacy API path (where the command supports it).
Commands that require --via wab. A few commands have no cookie path and only run through the Wonda Automation Browser: wonda linkedin comment, wonda linkedin reply-comment, wonda linkedin mute, wonda linkedin follow, wonda linkedin edit-post, wonda linkedin edit-comment, wonda linkedin delete-comment, wonda linkedin post --media, wonda x delete, wonda x reply --attach, wonda x dm send, wonda x dm accept, and wonda x dm start. On these, the default already resolves to wab (one stderr line noting it); passing --via cookies explicitly errors. Reddit's writes (vote, comment, subscribe, save, unsave, delete, and subreddit submit) are likewise wab-only.
Where it runs (--engine). --via picks the transport (browser vs. cookies); --engine picks the location, and the two are orthogonal. Values: local (this machine's WAB), cloud (the account's cloud twin, reached through the twin-action API with the control-session warm-up hidden behind a blocking wait), or auto (the default). auto resolves to local when the identity lives on this machine and cloud when it only exists as a cloud twin (a persona with no local footprint, or one cached as home=cloud from wonda twin provision). So wonda linkedin posts <profile> --account <twin> --engine cloud --via wab returns recent posts through the cloud twin's browser with the same structured result as the local command. --via works the same on both engines: reads default to cookies, writes default to wab, and an explicit supported override is preserved. Wired for LinkedIn (posts, connect, like/unlike, comment, reply-comment, edit-comment, send-message, follow, mute, delete-post, edit-post), X (like/unlike, bookmark, retweet/unretweet, follow/unfollow, delete), Reddit (vote, subscribe, save/unsave, delete), and Instagram (comment); other verbs run local only. --engine is accepted at the platform level (so both linkedin --engine cloud posts and linkedin posts --engine cloud work) but honored only on the wired verbs; passing it to another read or an unwired write is a clear error, not a silent no-op. wonda twin run-action is deprecated in favor of <platform> <verb> --engine cloud (it still works so running agents are not broken). On cloud, like supports plain likes and reactions (--reaction, routed to the react action); only comment reactions (--comment) stay local for now. auto resolves to local when a persona has no local footprint and no cloud twin (so first-use auto-create still works), and to cloud when a cloud twin exists for it. An explicit --engine cloud always runs on the cloud twin and never falls back to a live local relay.
Per-account credentials. Cookies live in per-account JSON files on disk:
~/.wonda/x-cookies/<account>.json~/.wonda/reddit-cookies/<account>.json~/.wonda/linkedin-cookies/<account>.json(auto-migrated from the legacy single-file format)
Each file is a session-owned local cache, not a portable credential. A platform account may have separate sessions on multiple computers and a cloud Twin. The file records the owning device/persona WAB session, and every injection, overwrite, cookie-only read, refresh, and backup restore checks that identity first. The CLI does not upload these cookies to or download them from the legacy shared hosted-token store. --force never bypasses a session mismatch.
Pass --account <name> to auth set to keep multiple logins side-by-side on the current device. The binding is recorded against the resolved account persona in account-bindings.json, even when --persona is omitted, and if the matching persona's Chromium is running, the rotated cookies get pushed into that live context. Never use auth set to copy cookies from another device or a Twin. Native wab login is safer. The driver also syncs cookies back to disk every 10 minutes (and on graceful shutdown), so rotated cookies (ct0 cycles, token_v2 server-side refresh, etc.) flow back to the cookies path without manual re-paste. Direct X cookie requests also absorb response cookie rotation into the same store and send the full stored cookie jar, preserving device-trust cookies across reads.
The cloud Twin is its own stable, always-on session. Ordinary provisioning does not seed LinkedIn from local cookie files; use wonda twin login <persona> --platform linkedin to log in inside the Twin. Cloud backups are recovery artifacts for their source session only. Inspect local and cloud ownership with wonda wab cookies status [persona]; it never prints cookie values. Recent LinkedIn posts have a direct structured cloud path: wonda linkedin posts <profile> --persona <persona> --engine cloud --via wab. Other reads that are not yet engine-wired can run inside the Twin with wonda twin run-now <persona> --command "<platform read command>"; retrieve the captured result with wonda twin output <twinRunId>, using the id returned by run-now. Neither path downloads Twin cookies locally.
Safely refresh LinkedIn's flat cookie file. Run wonda linkedin auth refresh --account <account> --persona <persona> before a cookie-only read batch. Fresh disk cookies are a fast local no-op. Stale or near-expiry cookies cause one WAB start (a no-op when already running), one WAB-to-disk sync, and a local/WAB-side LinkedIn session check. The refresh path never invokes the raw wonda linkedin auth check probe and never retries credentials. If the WAB session is dead, it exits nonzero with native re-login required; stop the batch and recover manually with wonda wab login <persona> linkedin.
wonda linkedin auth refresh --json returns {fresh, refreshed, ageSeconds, expiresAt, sessionAlive}. fresh describes the final disk state; refreshed is true only when WAB-to-disk sync ran; ageSeconds is the disk cookie age; expiresAt is the known expiry or null; and sessionAlive is the local/WAB-side login result. sessionAlive is null on a fresh local no-op because no WAB check was needed. --fresh-within changes both the maximum accepted disk age and the minimum remaining recorded li_at lifetime (15 minutes by default). Treat a nonzero exit as authoritative even with JSON output.
Cookie-backed LinkedIn reads accept opt-in --freshen alongside --via cookies. It runs the same local preflight only when the disk cookies are stale, then attempts the requested read once. The default remains unchanged: without --freshen, --via cookies stays fast and browser-free.
Safely refresh X's flat cookie file. wonda x auth refresh --account <account> --persona <persona> uses the same local-first lifecycle with X-specific freshness rules. Since auth_token has no locally readable expiry claim, freshness is the cookie store age only. A store newer than --fresh-within (15 minutes by default) is a browser-free no-op. A stale store starts the selected WAB if needed, syncs X cookies to disk, and checks the WAB-side session without invoking the raw x auth check probe. A dead session exits nonzero with the exact native login command.
wonda x auth refresh --json returns {fresh, refreshed, ageSeconds, sessionAlive}. Cookie-backed X reads accept opt-in --freshen; it runs the same refresh only when the selected disk store is stale, pins the read to the refreshed account, and never applies to WAB reads, writes, or auth commands.
Use wonda x auth status --account <name> for a pure-local view of cookie origin, latest capture provenance, generation, ownership, cookie names, and missing device-trust-cookie risk. It never contacts X and never prints cookie values. Use wonda x browser-bootstrap --account <name> to explicitly push the selected stored jar into running WAB personas bound to that account.
Action rate limits
Every platform command (linkedin, x, reddit, instagram), reads AND writes, runs through a per-profile rate-limit guard so a burst doesn't trip a platform's shadow-ban / anti-abuse heuristics. Accounting is per (platform, account) in a rolling 24h window, logged per profile under ~/.wonda/wab/personas/<persona>/ (so a cloud twin's caps persist across runs).
-
Reads are paced, never blocked: spacing is jittered to keep a profile under
read_per_min(75/min default), holding across separate invocations. -
Writes are checked against per-bucket daily caps. The LinkedIn defaults (other platforms track + count toward the total/day but have no per-type write cap by default):
salesnav search (Sales Navigator) is exempt: it carries no Commercial Use Limit, so it paces as an uncapped read instead of counting toward the search cap or the total/day.
Caps are SOFT by default: an over-safe / over-max action prints a shadow-ban-risk warning to stderr and proceeds. Pass --hard (or set mode: hard in config) to make over-cap writes abort (exit 1) instead.
wonda actions is a JSON data query (not a dashboard) for reading a profile's rolling-24h usage vs caps on demand; the caps/pacing/warnings run silently in the live hook regardless.
When an API key is configured, the local ledgers (actions log, WAB audit/error logs, cookie provenance) also sync to your Wonda account-health record automatically in the background on every command: best-effort, batched, and idempotent (a stable client event id per record means retries never double count), so offline use keeps working and sync catches up later. wonda actions sync forces a full flush and prints the server's insert/dedup counts; without an API key it is a silent no-op. Only event metadata travels, never cookie values or failure bundles. WONDA_TELEMETRY_DISABLED=1 turns the background sync off.
Override / disable / hard-mode via ~/.wonda/config.json under action_limits (caps are clamped to safety floors/ceilings so an override can loosen but not silently disable the guard):
Config keys
wonda config get|set|list keys:
api-key: your wondercat API key.base-url: API base (defaults to prod, set tohttps://staging.api.wondercat.aifor staging).default-account: account used when a platform command doesn't pass--account.wab-backup-enabled:true/falsefor cookie cloud backup (same aswonda wab backup enable/disable). On by default; only an explicitfalsedisables it.
Transport is NOT a config key. Each command picks it per kind (reads default to cookies, writes / engagement default to wab), identically on every platform. Override it per command with --via cookies|wab (where the platform supports it).
How to think about content creation
You are a marketing director with access to a full production toolkit. Before touching any tool, think:
- What product category? (beauty, food, tech, fashion, fitness, etc.)
- What format performs for this category? (UGC memes for everyday products, cinematic for luxury, before/after for transformations, testimonial for services)
- What's the hook? (relatable scenario, surprising twist, aspirational lifestyle, social proof)
- What specific scene? (not "product on table" but "person discovering the product in a funny situation")
Decision flow
When asked to create content, follow this order:
Step 1: Gather context
Step 2: Check content skills
Content skills are step-by-step guides for common content types. Each skill tells you exactly which models, prompts, and editing operations to use — and in what order. ALWAYS check skills before building from scratch.
Skills are server-hosted, per-account, and editable (the same model as wonda brand save) — you don't download a folder of .md files you own. Wonda ships a canonical set of default skills, served read-only as the fallback. Pull them live for a task; fork a default into your own copy when you want to tweak it; Wonda keeps full version history. skill list shows your effective skills (defaults overlaid with your own edits), and flags any fork whose default has since changed.
Default skill catalog (live source: wonda skill list, which also shows your own forks/edits and flags drift):
video
image
social-research
strategy
utility
<!-- SKILLS_TABLE_END -->Editing skills (optional). When a default doesn't quite fit, fork and edit it instead of working around it. Editing a default forks it into your account automatically:
If a skill matches → wonda skill get <slug>, read it, adapt to context, execute each step.
If no skill matches → build from scratch (Step 3).
Step 2.5: Decide whether finishing should be local
Not every media task should go back through Wonda editing. Use this routing rule:
- Use
wondafor AI generation, AI transcription/alignment, scraping, publishing, hosted transitions, and workflows that need media IDs or remote jobs. - Use local
ffmpegfor deterministic transforms on files you already have or can download: trim, crop/scale/pad, concat (merging multiple clips), replace audio, extract audio/frame, reverse, normalize for delivery, burn captions, split scenes, cut silence, and build analysis artifacts. Always merge clips locally — server-side merge can hang for 30+ minutes once any input exceeds ~7MB.
When a task starts from a Wonda media ID but the actual edit is deterministic, move it to local files first:
Before any local ffmpeg work:
Font rule for local caption/text work:
- Prefer an explicit font file path over a family name.
- Never assume a font exists. Check first with
fc-match,fc-list,/System/Library/Fonts,/Library/Fonts,~/Library/Fonts, or/usr/share/fonts. - If the task is mainly local finishing/captions/formatting/splitting/artifact extraction, check the
ffmpegskill before inventing commands. wonda edit videoruns a local ffmpeg for every editor op:trim,crop,volume,speed,reverseVideo,extractFrame,extractAudio,editAudio,imageCrop,imageToVideo,merge,overlay,splitScreen,splitScenes,skipSilence. The render runs on your machine via ffmpeg: no server-sideeditor_joband no credit hold for the render itself (inputs are downloaded and the result uploaded around it).textOverlayandanimatedCaptionsalso run locally, via the bundled hyperframes (Chromium) renderer. ffmpeg must be on PATH (wonda doctorverifies). The public API/video/edit,/image/edit,/audio/editare no longer used for these and return 410 Gone.- Always merge clips locally. Server-side merge can hang for 30+ minutes once any input exceeds ~7MB, and
wonda edit video --operation mergenow runs in local ffmpeg by default for the same reason. - Never mix per-clip audio then concat. Concat the video tracks first, then layer the full voiceover or music track once over the joined timeline. Per-clip audio bakes create cut-line collisions and silent gaps.
Default local export target unless the user asked otherwise:
Always pass -y as the first flag so the command auto-overwrites the output. ffmpeg prompts interactively when the output path exists and agent shells hang on that prompt until timeout.
Step 2.6: Pick the right local tool
Editing maps to one of four tools. Pick the first row that matches.
Run wonda doctor once on a new machine to confirm ffmpeg, node, and hyperframes are all available. Pass --warm-chrome to pre-fetch hyperframes' bundled Chromium (~150 MB) so the first clipping render doesn't pause to download it. Pass --wab to also audit the on-disk WAB browser runtime (installed driver vs this build's pin, driver tree manifest); offline, read-only, and advisory like --relay.
Examples:
Primitive trim and merge (wonda edit, local ffmpeg):
Motion graphics intro (wonda compose, hyperframes):
Kinetic captions on a finished clip (transitions service):
Raw ffmpeg for an op no primitive covers (e.g. concat with audio fade out):
Multi-step pipeline (compose intro → wonda merge with main → transitions captions):
Step 3: Build from scratch (chain endpoints)
When no skill matches, chain individual CLI commands. Each step produces an output that feeds into the next.
Single asset:
Audio (speech, transcription, dialogue):
Audio AI operations (direct-inference, NOT editor ops):
Add animated captions to a video:
The animatedCaptions operation handles everything in one step — it extracts audio, transcribes for word-level timing, and renders animated word-by-word captions onto the video.
The video's original audio is preserved. Do NOT replace the audio with TTS — Sora already generated the speech.
Transitions (effects pipelines on a single video):
Use exactly one of --preset or --clips. Requires a full (logged-in) account. Always read wonda transitions llms first when composing a clips timeline. It documents the detect/segment/effect dependencies, which ops need masks, and the full clip-spec shape (layer types, tracks, effects, transforms).
Preset variables (variables block). Each preset declares the template variables it accepts under variables in wonda transitions presets. Each entry has name, description, and required. Required variables MUST be supplied or the job is rejected with a 400 — no more silent skipping. Pass them with --var name=value (repeatable) or, for the common prompt case, the --prompt shortcut:
The prompt variable is a detection text query describing which subject to mask, fed to SAM3 to produce per-frame segmentation masks. Not a content-generation prompt.
Building a custom --clips timeline that needs detection masks? Add a clip with layer_type: "video" and a mask: {layer_type: "mask", analysis_steps: [{name: segment, params: {prompt: "..."}}]}. SAM3 handles both detection and segmentation in one step from the prompt, so no separate detect step is needed.
Pre-warming masks before render (recommended)
For presets with mask:<label> variables, run wonda transitions ensure-masks first so the render starts with masks already prepared. The first call for a (media, label) pair takes 1-3 minutes; subsequent calls are near-instant.
ensure-masks flags:
--media MEDIA_ID— required, the video the masks are for--label NAME— repeatable, one label per call (--label person --label phone)--labels NAME,NAME— comma-separated alternative (--labels person,phone)--wait— block until every label is prepared--timeout DUR— cap wait time when--waitis set (default 10m)
Multi-prompt syntax: mask:woman+phone in --var is split into separate masks (woman, phone) and unioned per-frame. Pass each sub-label separately to ensure-masks so all of them are pre-warmed.
When to skip ensure-masks:
- Non-mask presets (no
mask:<label>variables) — nothing to prepare - A previous render already used these (media, labels) — already prepared
When ensure-masks matters most:
- First render of a new media with mask-based presets
- Iterating params on a render — pre-warm once, then run as many times as you want without re-preparing
Multi-scene presets (requiresMultiScene: true). Some presets use scene-aware logic and expect a video with multiple cuts/scenes. Check requiresMultiScene in wonda transitions presets. If true, feeding a single continuous shot will produce only one scene and the effect may look underwhelming. Combine clips first or use a video with natural cuts.
Tweaking preset params. Every preset is clip-shape. Pull a single preset with wonda transitions preset <name> --json, read its clips: (single-track) or tracks: (multi-track) field, edit any clip param, and submit as --clips. For multi-track presets, flatten by giving each clip a track index drawn from the track it came from. If the preset declares sceneTransitions:, pass that array through unchanged on the request.
Auto-repair safety net (--auto-repair, --face-bbox). For --clips renders the worker runs a deterministic repair pass on the submitted JSON before rendering, default on. Repairs: width-fit font clamp, descender clamp against canvas bottom, stack-spacing snap (ROW1_py from cap-height formula), keyframe-bound clamp to [0, source_duration], same-y-row caption overlap trim, mask full-duration extension, stroke-width zeroing, letter-spacing target snap per font, mask-cutout duration extension, negative-start clamp, and (with --face-bbox) face-overlap caption shift. Pass --auto-repair=false for strict validation; out-of-spec values then surface as render errors.
--face-bbox only shifts body captions. Decorative text you want behind the speaker still routes through an explicit mask_cutout {prompt: "person"} clip.
Output URL paths differ by job type:
- Inference jobs (generate, audio):
.outputs[0].media.urland.outputs[0].media.mediaId - Editor jobs (edit):
.outputs[0].urland.outputs[0].mediaId
Model waterfall
Image
Default: gpt-image-2. OpenAI's flagship — strongest prompt adherence, best text-in-image, high-fidelity edits via reference images. Handles 1-4 reference images. Quality tiers: auto (default), low, medium, high — pass via --params '{"quality":"high"}'. Caps at 1536px output.
For img2img editing specifically (change, add/remove, restyle, bg-remove, crop, text overlay, vectorize), use wonda skill get image-edit — it has the full edit-specific decision tree.
Pick something else only when one of these applies:
- User explicitly requests another model
- More than 4 reference images →
nano-banana-2(gpt-image-2 caps at 4 refs; nano-banana-2 accepts up to 14). For 1-4 refs, stay ongpt-image-2. - Need vector output →
runware-vectorize - Need background removal →
birefnet-bg-removal - Cheapest possible / fastest drafts →
z-image - Need >1536px / true 4K output →
nano-banana-pro(1K/2K/4K) ornano-banana-2(1K/2K/4K). gpt-image-2 caps at 1536px. - gpt-image-2 unavailable / OpenAI down →
nano-banana-2orseedream-4-5orgrok-imagine-pro
Video
Default: seedance-2 (duration 5/10/15s, default 5s, quality: high). Escalation:
- Quality complaint or different style →
sora2orsora2pro - Max single-clip duration is 15s for Seedance 2, 20s for Sora → for longer content, stitch multiple clips via merge
- Veo (
veo3_1,veo3_1-fast) is available but NOT in the default waterfall. Only pick Veo when the user explicitly asks for Veo by name. - Gemini Omni (
gemini-omni-video) is available but NOT in the default waterfall. Only pick it when the user asks for Gemini by name, or specifically needs multi-image reference T2V/I2V (up to 7 reference images) or 4K output.
Image-to-video routing (MANDATORY when attaching a reference image):
- Person/face visible in the reference image → MUST use
kling_3_pro(preserves identity better for faces) - No person in reference image → use
seedance-2 - Text-to-video (no reference image): Seedance 2 generates people fine. This rule ONLY applies when you
--attachan image.
Kling model family:
kling_3_pro— Text-to-video and image-to-video, supports start/end images, custom elements (@Element1, @Element2), 3-15s duration, 16:9/9:16/1:1kling_2_6_pro— General purpose, 5-10s, 16:9/9:16/1:1, text-to-video and image-to-videokling_2_6_motion_control— Motion transfer: requires both a reference image AND a reference video, recreates the video's motion with the image's appearancekling2_5-pro— Budget Kling option, 5-10s, supports first/last frame images
Kling prompt rules (important): Kling's prompt field caps at 2,500 characters and Kling responds poorly to Sora-style structured briefs (SCENE: / SUBJECT: / MOTION: / BANNED LOOK: section headers). In that format Kling latches onto atmosphere nouns and silently drops the central subject (verified empirically: the same 2,842-char Sora-style prompt that rendered correctly on Sora 2 Pro and Seedance 2 produced no phone at all on Kling — even when trimmed to 2,250 chars). When escalating Seedance → Kling, or targeting Kling directly, rewrite the prompt as short natural-language prose (~1,000–1,500 chars) and lead with the hero subject in the opening sentence rather than burying it inside a SUBJECT: block. Do NOT pass a Sora-formatted prompt through to Kling unchanged.
Other video models:
grok-imagine-video— xAI video generation, 5-15s, supports 7 aspect ratios including 4:3 and 3:2gemini-omni-video: Google Gemini Omni. Text-to-video and image-to-video with up to 7 reference images (slotsreference_image_1throughreference_image_7). Durations 4/6/8/10s, aspect ratios 9:16 and 16:9, resolutions 720p / 1080p / 4K. Pricing: $0.15 base + $0.075/s at 720p/1080p, $0.75 base + $0.075/s at 4K. No native audio (pair with a separate audio model if speech is needed).topaz-video-upscale— Upscale video resolution (1-4x factor, supports fps conversion)sync-lipsync-v2-pro— Legacy lipsync for user-supplied video + audio pairs. Inferior to native-audio generation and almost never the right choice for new content. See the "Lip sync" section for rules.
Seedance family (DEFAULT video model, watermarks automatically removed):
seedance-2— Base Seedance 2.0 (T2V/I2V, 5-15s, high=standard/basic=fast)seedance-2-omni— Multi-reference generation (images, audio refs)seedance-2-video-edit— Edit existing video via text prompt
Video durations: Accepted --duration values vary by model. Check with wonda capabilities or wonda models info <slug>.
Audio
- Music:
suno-music(set--params '{"instrumental":true}'for no vocals) - Text-to-speech:
elevenlabs-tts— only for explicit narrator/voice-over asks over silent footage. Do NOT use to "make a UGC character talk" — Sora / Sora 2 Pro / Veo 3.1 / Kling 3 / Seedance 2 generate native synced speech in any language, which looks and sounds far better. Always set voiceId in params. Default female voice:--params '{"voiceId":"21m00Tcm4TlvDq8ikWAM"}'(Rachel). - Transcription:
elevenlabs-stt - Multi-speaker dialogue:
elevenlabs-dialogue - Enhance audio (clean up noisy speech):
replicate-resemble-enhanceviawonda audio enhance— denoise + dereverberate. Use when a voice recording sounds muffled, echoey, or has background noise. NOT a general "sounds better" button; if the source is already clean this can soften it. - Extract voice (isolate vocals / split stems):
replicate-demucsviawonda audio extract-voice— splits into voice and instrumental tracks. Use to pull a speaker or singer off a track, or to isolate the music behind a vocal.
Native synced speech (preferred over TTS + lipsync): Sora, Sora 2 Pro, Veo 3.1, Kling 3, and Seedance 2 all generate dialogue in any language directly inside the video, with mouth movements baked in. Put the line (and language) in the video model's --prompt. Never chain elevenlabs-tts → sync-lipsync-v2-pro to fake speech over a silent generation.
Characters
Characters are reusable saved combos (image + optional voice audio) you can mention in prompts with @name. The server auto-injects the image, optional face video, and audio into the right slots for the selected model. Works on Kling 3 Pro (start_image + element_1 + voice_audio) and Seedance 2 Omni (ref_image_1 + ref_video_1 + ref_audio_1). Name rules: must start with a letter, 1–31 chars, alphanumeric + _/-.
Provider gotchas (Seedance 2 Omni): when a character is mentioned, the API routes Seedance to MuAPI automatically. Replicate enforces a 15s ref_audio_1 cap and rejects famous-celebrity refs with E005 — input flagged as sensitive. MuAPI is the reliable path for character-driven jobs. Even on MuAPI, top-tier celebrity refs (think Sydney Sweeney, Leonardo DiCaprio) are blocked with "Face detected in uploaded image. Please use an image without real people." Non-celebrity faces and lesser-known public figures pass cleanly. If you see that error on a real-person ref, use Kling 3 Pro instead (its character pipeline runs voice cloning server-side, so the raw face audio never touches a moderation classifier).
From a Kling clip — extract a frame + voice from a generation you like:
From scratch — generate a portrait and a TTS sample, then bind them:
List / inspect / update / delete: wonda character list, wonda character get <name>, wonda character update <name> --audio $NEW, wonda character delete <name>. Only one character with audio can be referenced per generation.
Prompt writing rules
Follow this waterfall top-to-bottom. Use the FIRST matching rule and stop.
-
PASSTHROUGH — If the user says "use my exact prompt" / "verbatim" / "no enhancements" → copy their words exactly. Zero modifications.
-
IMAGE-TO-VIDEO — When a source image feeds into a video model, describe MOTION ONLY. The model can see the image. Do NOT describe the image content.
- Good:
"gentle breathing motion, camera slowly pushes in, atmospheric lighting shifts" - Bad:
"Two cats on a lavender background breathing softly"(describes the image)
- Good:
-
EMPTY PROMPT (from scratch) — Use the user's exact request as the prompt. Do NOT add style descriptors, lighting, composition, or mood.
- User says "create an image of a cat with sunglasses" → prompt:
"create an image of a cat with sunglasses" - Do NOT enhance to
"A playful orange tabby wearing oversized reflective sunglasses, studio lighting, shallow depth of field"
- User says "create an image of a cat with sunglasses" → prompt:
-
NON-EMPTY PROMPT (adapting a template) — Keep the structure and style, only swap content to match the user's request. Keep prompts literal and constraint-heavy.
Aspect ratio rules
Three cases, no exceptions:
- User specifies a ratio → use it:
--aspect-ratio 16:9 - User doesn't mention ratio → explicitly set
--aspect-ratio 9:16for social content (UGC, TikTok, Reels, Stories). Portrait is the default for any social/marketing video. - Editing existing media → use
--aspect-ratio autoto preserve source dimensions
UGC and social content is ALWAYS portrait (9:16). If someone asks for a TikTok, Reel, Story, or UGC video, always use --aspect-ratio 9:16. Landscape is only for YouTube, presentations, or when explicitly requested.
Square (1:1) is supported by all Kling models and some image models — use for Instagram feed posts when requested.
Common chaining patterns
These patterns show how to compose multi-step pipelines by chaining CLI commands. Each step's output feeds into the next.
No need to download and re-upload between steps. Every generation and edit produces a media ID in its output. Pass that ID directly to the next command via
--mediaor--audio-media. Use--jq '.outputs[0].media.mediaId'for inference jobs and--jq '.outputs[0].mediaId'for editor jobs. Only use-o <file>on the FINAL step to download the finished output.
Animate an image to video
Replace audio on a video (TTS voiceover or music)
Only use this when you need to REPLACE the video's audio. Sora, Sora 2 Pro, Veo 3.1, Kling 3, and Seedance 2 all generate native synced speech in any language — don't replace it with TTS unless the user explicitly asks for a different voiceover. Never reach for this step to "add speech" to a UGC/talking-head clip; put the dialogue in the video model's prompt instead.
Add static text overlay
Static overlays (meme text, "chat did i cook", etc.) use smaller font sizes than captions. They're ambient, not meant to dominate the frame.
Featured textOverlay + animatedCaptions presets. wonda edit {video,image,audio} accepts --preset <name> (scoped to --operation). --params fields override preset values on key collisions.
textOverlay (static, top-centered):
TikTok White Highlight— black text on a slightly rounded white box.TikTok Black Highlight— white text on a slightly rounded black box.TikTok Red Highlight— white text on a slightly rounded red (#E14135) box.
animatedCaptions (STT-driven, bottom-centered):
TikTok White Captions— black text, white highlight on the active word.TikTok Black Captions— white text, black highlight on the active word.TikTok Red Captions— white text, red (#E14135) highlight on the active word.
textOverlay renders locally via the bundled hyperframes (Chromium) renderer. There is no server-side image textOverlay anymore.
Font sizing guide:
- Static overlays:
sizePercent: 66,fontSizeScale: 0.5,strokeWidth: 4.5 - Animated captions:
sizePercent: 80,fontSizeScale: 0.8,strokeWidth: 2.5,highlightColor: rgb(252, 61, 61) - Font:
TikTok Sans SemiCondensedfor both
Add animated captions (word-by-word with timing)
The animatedCaptions operation extracts audio, transcribes, and renders animated word-by-word captions — all in one step.
For quick static captions (no timing, just text on screen), use textOverlay with --prompt-text:
Add background music
Editor output chaining
When chaining multiple editor operations (e.g., editAudio → animatedCaptions → textOverlay), extract the media ID from each editor job output and pass it to the next step. Note the jq path differs from inference jobs:
Merge multiple clips
Always merge locally with ffmpeg. Server-side merge (wonda edit video --operation merge) can hang for 30+ minutes once any input exceeds ~7MB.
Download every Wonda media ID, then concat. Stream-copy is fast but requires matching codec/profile/resolution; fall back to re-encode if it errors:
File order in concat.txt = playback order. See the ffmpeg skill for the full concat reference.
Split scenes / keep a specific scene
Two modes, pick by intent:
Use omit mode for "remove frozen first frame" (common with Sora videos). Use split mode to get all scenes as separate clips.
Image editing
Any image edit — img2img, background removal, crop, text overlay, vectorize — has its own skill with the full decision tree, aspect-ratio rules, and model waterfall for edits:
One gotcha worth keeping here: image and video background removal use different models (birefnet-bg-removal vs bria-video-background-removal). Never swap them.
Lip sync (last-resort fallback — prefer native-audio video models)
Sora, Sora 2 Pro, Veo 3.1, Kling 3, and Seedance 2 all generate speech in any language with correctly synced mouth movements as part of the video itself. That path produces dramatically better results than sync-lipsync-v2-pro: better lip physics, better lighting, better costs, and no second inference round-trip. For any talking UGC, ad, or spokesperson video, put the dialogue directly in the video model's prompt — do not chain TTS + lipsync.
Only reach for sync-lipsync-v2-pro when the user EXPLICITLY supplies both a pre-existing video and a pre-existing audio clip and asks you to align the mouth to that audio. If a user asks for lipsync as the default method of making a character speak, push back: the native-audio video models are the better tool and work in any language.
Video upscale
Clipping (longform → vertical shorts)
wonda clipping takes a long video (podcast, interview, talking-head)
and produces short vertical clips. Selection is LLM-driven and supports
a natural-language --brief so you can ask for specific moments instead
of generic virality.
V1 renders 9:16 with face-tracked reframe (LR-ASD active-speaker
detection + One-Euro stabilizer, default) and the existing
animatedCaptions op + a top-third hook overlay per clip. Pass
--reframe blur-fill to keep the full landscape source inside a
vertical canvas with a blurred background instead.
Translated captions: pass --caption-language <code> (ISO-639-1,
e.g. en) to render captions in another language while keeping the
original audio. Each clip's transcript is translated per sentence and the
spoken timing is preserved, so captions stay in sync. Omit the flag to
caption in the spoken language. On --restyle, the language is inherited
from the parent job unless overridden, so you can cheaply spin out an
English-captioned version of an already-clipped job (reuses the
transcript, no re-transcription).
Async: POST /api/v1/clipping returns a clippingJobId; the CLI polls
GET /api/v1/clipping/jobs/{id} under --wait. Pass --output <dir>
and the CLI downloads each rendered clip + a plan.json.
Auth: included in the paid plan.
Source: --url accepts YouTube and direct mp4 URLs.
YouTube links work; a long video can take several minutes to ingest before
transcription starts. If a YouTube ingest fails, download the file locally
and upload it first, then clip with --media:
--no-transcode skips the server-side normalize so longform sources are
usable in seconds instead of minutes (a 1 h upload otherwise transcodes for
~19 min before clipping can start). Clipping handles incompatible formats
(AV1, rotated phone video) per selected clip automatically. Omit the flag
for uploads you plan to publish or edit directly without clipping.
Job-status shape (returned by GET /api/v1/clipping/jobs/{id}):
Editor operations reference
cropvsimageCrop: videocropis ratio/percent based (aspectRatioorcropPercent+cropAxis); it does NOT take pixel coordinates and rejectscropPixelX/Y/Width/Heightwith an error. For an exact pixel rectangle, useimageCrop. Runwonda operations info <operation>for the full param list, defaults, and ranges of any op.
Valid textOverlay fonts: Inter, Montserrat, Bebas Neue, Oswald, TikTok Sans, TikTok Sans Condensed, TikTok Sans SemiCondensed, TikTok Sans SemiExpanded, TikTok Sans Expanded, TikTok Sans ExtraExpanded, Nohemi, Poppins, Raleway, Anton, Comic Cat, Gavency Valid positions: top-left, top-center, top-right, center-left, center, center-right, bottom-left, bottom-center, bottom-right
Marketing & distribution
X/Twitter
Supports reads, writes, and social graph.
⚠️ Anti-fraud caution: don't probe freshly-pasted cookies. When you've just received cookies (yours or a user's), the FIRST request on them should be the operation the user actually wants, not
wonda x auth check, notwonda x home, not anything that fires a probe. Burst activity on a new IP / device / process is the textbook signal X (and Reddit / LinkedIn / IG) flag as credential theft, and the cookies get shadow-banned or hard-killed. If you must verify, usewonda x auth check --account <name> --via wab(that routes through the account's existing logged-in browser session: same IP, same fingerprint, same browsing history) instead of firing a raw API request from a fresh process.
All paginated commands support: -n <count>, --cursor, --all, --max-pages, --delay <ms>.
wonda x user includes additive emails and links fields when the member has published explicit contact details in their public bio or X exposes expanded URL entities for that profile. These fields are read-only and equivalent to manually reading the profile. No inferred emails, enrichment data, storage, or DOM scraping is involved.
Tweet modes: The tweet command has two transports:
--via cookies(internal API): X's internal GraphQL (CreateTweetfor ≤280 chars,CreateNoteTweetfor long-form Premium). Fast (<1s), supports--attachfor media. Occasionally fails with error 226 when X rotates query IDs or feature flags. When that happens, recapture viatwitter-tone-research/_artifacts/scripts/capture-ct-bw.mjsand bump the three knobs inxclient/.--via wab(default for writes): Routes through the account's WAB Chromium (auto-spawned on first--via wabuse), opens x.com compose, types with human-style jitter, clicks Post. Supports--attach(image/gif/video, up to 4); files are driven through the hidden compose input via Playwright'ssetInputFiles, no native picker dialog opens; the script waits for X's upload pipeline to finalize (up to 5 min for video) before submitting. Zero fingerprinting risk. Slower (~10s text, ~30-90s with video) but fully drift-proof: no queryIds, feature flags, or request shape to maintain. The stealth browser + Chromium install once viawonda wab install(~315 MB, one-time, idempotent). Cookies live in~/.wonda/x-cookies/<account>.json, bound to the account's persona viaaccount-bindings.json.wonda x reply --attachis wab-only (no cookie path).
Inbound engagement
Received engagement means inbound activity on the twin's own content: replies, mentions, comments, and DMs. This is the reverse direction of outbound posting, liking, and following.
It is not keyword monitor, not sent-invite/outreach acceptance tracking, and not the outbound action ledger. Use platform search commands for keyword monitoring. Use wonda actions for outbound action accounting.
wonda inbound fans out direct read clients only: X mentions/replies, LinkedIn notifications, and Reddit classic inbox. Output is always normalized JSON:
By default it persists a high-water mark per persona/platform/account at ~/.wonda/wab/personas/<persona>/inbound-state.json and only emits items newer than the last run. Corrupt or missing state is treated as a first run. --all bypasses state reads and writes and returns the full fetched set.
X DMs: wonda x dm inbox, wonda x dm requests, and wonda x dm read read through the cookie-backed xclient path by default and can use --via wab for browser-session transport parity. requests returns the untrusted/message-request inbox so recipient-side acceptance can be handled explicitly. wonda x dm send, wonda x dm start, and wonda x dm accept are WAB-only DOM writes: sends open the real X messages UI, type into [data-testid="dmComposerTextInput"], and click the composer send button; accept opens the request thread through Chat and clicks the visible Accept/Allow control. start verifies that the selected suggestion's screen name equals the requested handle before sending, and WAB DM writes verify the active X browser account when --account or persona bindings provide one. If x dm send or x dm start fails with recipient_cannot_receive_dm, X refused the DM before send. Check that the recipient can receive DMs from the sender: mutual follow/connect may be required, and the recipient may need to accept any pending message request before retrying. If X shows an encrypted XChat passcode gate, store the passcode with wonda x dm passcode set --account <name>; it is encrypted locally using the same local-secret pattern as cookie-adjacent storage and is never printed or included in failure bundles. Deeper XChat/Juicebox message decrypt/encrypt support remains unsupported until X exposes a thread that needs it.
Supports search, profiles, companies, messaging, and engagement.
⚠️ Same anti-fraud caution as X: don't probe freshly-pasted cookies. First request on new cookies = the actual operation, never a check. LinkedIn's anti-fraud is the most aggressive of all the platforms (force-logout, password reset, account flag). If you must verify, use
wonda linkedin auth check --account <name> --via wabto route through the account's existing WAB session.
Paginated commands support: -n <count>, --start, --all, --max-pages, --delay <ms>.
wonda linkedin profile includes additive emails, websites, phone, and twitter fields when the member exposes them in LinkedIn's Contact info. The contact-info read uses the same logged-in Voyager cookie and CSRF path as profile reads, is best-effort when hidden or rate-limited, and is equivalent to manually opening the Contact info overlay. It does not infer, enrich, persist, or batch-harvest contacts.
LinkedIn member identity resolution: wonda linkedin resolve <id>... accepts bare ACoAA... fsd profile ids plus urn:li:fsd_profile: and urn:li:fs_miniProfile: wrappers. It performs one cookie-only Voyager read per unique cache miss and stores positive results for 30 days under the selected account. --no-cache skips persistent cache reads and writes but still deduplicates within the invocation. Numeric urn:li:member:<id> values are rejected because LinkedIn's authenticated web redirect preserves the numeric id and no stable Voyager mapping endpoint is available. conversations --resolve and messages --resolve use the same cache and only fill existing blank vanityName fields; without --resolve, their output is unchanged.
wonda linkedin profile also returns education (a list of {school, degree, fieldOfStudy, startYear, endYear}, endYear 0 meaning ongoing), yearsOfExperience (career span in years, computed from the earliest dated experience; 0 if none), and the full experiences work history (each {title, companyName, companyUrn, companyUniversalName, startYear, startMonth, endYear, location}, endYear 0 meaning current) plus currentCompany in the same shape for the present role. These read over the same logged-in Voyager cookie/CSRF path as the rest of the profile and are best-effort when hidden or rate-limited.
LinkedIn profile visit warm-up: wonda linkedin visit <vanity-or-url> --account <name> is WAB-only and intentionally separate from wonda linkedin profile. profile is a stealth data read and does not fire viewed-your-profile. visit opens the real /in/<vanity>/ page in the persona browser, dwells for roughly 6-20s by default, makes a light scroll pass unless --no-scroll is set, and returns { ok, status, profile, dwellMs, scrolled }. Private or anonymous LinkedIn profile-viewing settings can suppress the target notification.
LinkedIn account creation: wonda linkedin signup provisions a brand-new LinkedIn account: it mints a throwaway mailbox (or uses --email), drives the join flow (email plus password, name, emailed verification code, profile) in a headful WAB window, fetches the verification PIN from the inbox, then binds the persona and syncs cookies so LinkedIn reads and writes route through the new account. Gated by the linkedinAccountCreationEnabled flag (server-evaluated preflight: GET /linkedin/signup/enabled). Bind a mobile or residential proxy to the persona first (wonda wab config set <persona> proxy_url socks5://...): LinkedIn shadowbans accounts born on datacenter IPs even harder than Reddit. LinkedIn frequently inserts a captcha or "security verification" puzzle the flow cannot pass automatically; when a field cannot be located the flow pauses and leaves the window on that screen, so finish by hand then re-run with --resume <step> (account|name|code|profile|persist). On success it prints {name, email, password, persona, account} plus a ready-to-paste op item create block for the "LinkedIn logins" vault.
Connection request modes: The connect command has two transports:
--via cookies(API): Voyager REST API with fingerprint mitigations (profile visit, drawer warm-up, connect). Fast (~3s), supports notes viacustomMessage.--via wab: Routes through the account's persona Chromium (auto-spawned) for full stealth via DOM dispatch. Zero fingerprinting risk. Slower (~10s) but fully safe. Use when you need extra protection. The stealth browser + Chromium install once viawonda wab install(~315 MB, idempotent). The persona reuses its persistent profile under~/.wonda/wab/personas/<persona>/profile. Cookies live in~/.wonda/linkedin-cookies/<account>.json, bound to the persona viaaccount-bindings.json; rotating viawonda linkedin auth set --account <name>pushes the new cookies into the live Chromium if it's running.
Connection request JSON failures: A failed local WAB connect exits nonzero and writes a structured JSON error to stderr. error.code stays send-rejected for a genuine one-off rejection. A weekly invitation limit or silent send drop uses send-throttled with machine reason weekly_invitation_limit or silent_drop; HTTP 429 keeps blocked:rate-limit with reason http_429. Connect failures include throttled, and visible UI failures include the original toast in detail. With JSON output, a local --hard action-budget stop emits error: "hard-rate-limit", throttled: true, and reason: "local_hard_limit" on stdout while keeping the existing stderr sentence. Cloud-wrapped failures preserve the existing outer HTTP status, error.code, and error.reason; the raw connect token and reason are additive error.actionCode and error.actionReason fields.
LinkedIn InMail: wonda linkedin inmail <vanity-or-url> --message <body> [--subject <subject>] [--spend-credit] [--dry-run] drives LinkedIn's native InMail composer in WAB only. It is not connect and not send-message; first-degree recipients are refused with a pointer to wonda linkedin send-message. Open Profile / free InMail can send without credit confirmation. If LinkedIn indicates one InMail credit will be consumed, the command stops with sent:false and exits 0 unless --spend-credit is provided (--yes-consume-credit still works as a deprecated alias). --dry-run fills the composer and stops before Send.
LinkedIn Recruiter Lite (Wonda 1.61.0+): wonda linkedin recruiter provides six commands through the selected LinkedIn account/persona and requires its Recruiter Lite seat. Cookie search, list, bounded native open and pagination have passed normal CLI verification on an existing saved query. Profile comparison on one intended candidate retained seven positions, two education records and 42 skills across cookie and WAB transports; the final cookie profile returned exit 0 and partial:false. Matching the native parameter encoding and Recruiter app context resolved the observed request failures using the existing cookies. Native WAB save/list/open were observed on an intended real save with alerts off. Cookie save uses the inspected native client contract with duplicate/capacity checks, single-attempt creation and saved-search ID verification, but has fixture coverage only; no extra record was created for live testing. Exhaustive filter/operator parity is not established, and native open rejects filter overrides.
--via cookies means the authenticated internal Recruiter API; --via wab means actual browser navigation, rendered DOM and trusted input. Neither performs an authentication probe, automatic session refresh, retry or transport fallback. The initial cookie-search HTTP 500 was fixed by correcting native parameter serialization; a subsequent bounded search passed. Matching the native Recruiter app context also restored the requested profile skills. Missing requested sections still produce partial results. Do not work around an error by silently changing transports.
Search accepts repeatable --filter TYPE=VALUE, --exclude-filter TYPE=VALUE, --range-filter TYPE=MIN:MAX, --toggle TYPE[=true|false], --filters-json <JSON-or-local-file>, and a captured --search-url, including native project/saved-search GET URLs. Cookie native URL replay retains the project/owner context and currently rejects filter overrides. Structured clauses support operator: "require", scopes and raw observed values. WAB rejects unmapped raw fields before navigation. --interactive --via wab pauses for the operator to configure the real search; it does not save a native search. Catalog coverage does not yet imply every dependency or operator has live parity.
--count is results per logical page, not a total bound. The defaults are 10 results × 10 pages; --all permits up to 100 pages unless --max-pages is supplied. Use an explicit page bound for a small intended operation. Recruiter's physical UI page currently contains 25 results even when fewer are returned. --delay sets a minimum interval between every cookie request, including preparatory collection/project/typeahead reads, and before WAB interactions. It never lowers the existing cookie throttle. --max-pages also caps physical WAB pages, so a far-away --start can fail before it is reached. --details full adds profile reads, with --max-profiles defaulting to 250. --override-max-profiles permits a requested maximum above 250; the requested --max-profiles still caps visits, and platform errors and request pacing remain enforced. These values are Wonda controls, not LinkedIn-approved allowances.
Use only captured opaque candidate references or actual /talent/profile/... links. Public /in/... URLs are emitted only when exposed, with provenance. Candidate output retains current/past positions, multiline descriptions, education, skills, available optional sections and per-section states such as complete, partial, unavailable and unobserved. A missing field is not an empty complete section. On page/read failure, useful results remain on stdout with partial/errors and a nonzero exit. --csv generates Wonda CSV, retaining structured history in quoted JSON cells; LinkedIn native export is not required. CSV partial includes both candidate and result failures; extraction retains candidate-specific status, while result_partial and JSON result_errors preserve the overall outcome on each row. record_type=candidate identifies candidate rows; an empty failed result emits one record_type=result row with no candidate identity. Formula-like CSV text receives a quoted leading tab for spreadsheet import; that tab remains in the CSV data, while JSON preserves the original text. --no-cache bypasses profile-cache reuse. Partial profiles, explicit credentials and unknown/expired seat context bypass reusable cache entries. The seat/contract namespace comes from the existing local Recruiter session cookie without an authentication request; a changed context is not cached. Standalone WAB cache reuse has passed through the normal CLI.
Native project flags are --recruiter-project (exact name for WAB, ID or exact name for cookies) and --new-recruiter-project; global --project still means Wonda project context. Cookie project lookup prefers an exact name, including numeric names; a complete captured project URN selects an ID unambiguously. Alerts default off. WAB save/open were observed on an intended real record. An unverified save outcome must not be retried automatically. viewedMarkerChanged:null means the native GET marker effect is unknown. Never manufacture disposable records or capacity failures in a real account to verify them.
Local MCP and twin payloads share all six command contracts, explicit transport selection and bounded arguments. Tool search defaults are 25 × 1, capped at 375 requested candidates; all:true expands to at most 15 pages within that candidate cap, and an explicit maxPages remains authoritative. Full-profile inputs/enrichment cap at 10. Payloads whose deliberate waits exceed the 280-second synchronous budget (with 40 seconds reserved for work) are rejected before dispatch; use smaller batches or the local CLI for longer operations. The requested delay is never shortened. Tools accept JSON filter content, not local file paths. They require Wonda 1.61.0 or newer. CLI --engine local|cloud|auto uses the existing engine resolver and twin dispatch. Local cookie overrides cannot be sent to cloud; --interactive and --override-max-profiles require local execution. Runner/local MCP commands pin the local engine to prevent redispatch. Partial hosted results retain stdout and a nonzero exit.
LinkedIn InMail credit balance: wonda linkedin inmail-credits is read-only and returns the seat's remaining InMail credit balance as remaining (plus the grant type and id) from the Sales Navigator credit-grant endpoint over the cookies transport. It never sends anything and never consumes a credit; use it to plan before wonda linkedin inmail --spend-credit. It requires an active Sales Navigator seat on the account; flagship-Premium-only accounts cannot read their balance this way.
LinkedIn public profile enrichment: wonda linkedin enrich <profile-url-or-vanity...> defaults to --via cookies, uses the local logged-in profile read path one profile at a time, reuses _artifacts/linkedin-cache, adds jitter between misses, aborts on 429-style rate-limit signals, and rejects more than 25 inputs when run directly. The linkedin_enrich tool (local MCP and the twin action route) caps a batch targets call at 10 instead: that path shares tighter caller-side deadlines (the Go API client's 300s request-cancel, other synchronous HTTP callers) that a 25-profile run could exceed. Both --via cookies and --via wab include each profile's education (school, degree, fieldOfStudy, startYear, endYear) and derived yearsOfExperience in the returned profile, but the education fetch is best-effort: a transient failure of that optional request (anything other than a rate limit) still returns a successful enrichment with no education field, and that result is cached as-is. A cache hit is served as-is, and entries written by commands that share this cache without fetching education (enrich-engagers, search-posts author profiles), by enrich before this change, or by a prior best-effort miss lack education, so pass --force-refresh when education matters. enrich-engagers does not request education, so its per-engager profile read count is unchanged. Use --via public for paid public enrichment through POST /api/v1/scrape/linkedin-profiles, polling GET /api/v1/scrape/linkedin-profiles/{taskId} until completion unless --no-wait is set. Cancel a running paid scrape with wonda scrape cancel <taskId>. Public mode supports --force-refresh, --timeout 10m, --idempotency-key, and --output json|table. The single-profile wonda linkedin profile --via public path uses the same paid route and also supports --force-refresh, --timeout, and --idempotency-key.
--via wabruns the same batch through the persona's Wonda Automation Browser session: every profile read executes inside the logged-in browser (authentic fingerprint, no flat-store staleness). Same 25-input cap, pacing, cache, and education/yearsOfExperience fields as--via cookies. Prefer it when a batch follows a Sales Navigator export or when cookie reads have been failing.
LinkedIn sent invitation tracking: wonda linkedin sent-invitations is read-only and defaults to --via cookies. It lists live pending outgoing invites and, by default, reconciles them with the local WAB connect audit to label audited targets as accepted, pending, withdrawn, or unknown; pass --reconcile=false for the raw paginated pending list or --no-connection-check to skip per-target connection-status reads. --via wab routes the same reads through the account persona. LinkedIn's sent-list Voyager endpoints currently churn, so the command tries pageable Voyager variants before falling back to the read-only sent-invitations manager page. Results always report complete, fetched, knownTotal, and source. complete is true only when the inventory is proven complete; fetched counts distinct pending rows returned; knownTotal is LinkedIn's advertised total or null when absent or inconsistent; and source is voyager or html. HTML fallback results also report fetchCap, the rendered HTML row cap; it is omitted when HTML was not used. A partial inventory returns useful rows with complete: false. An audited target absent from a partial inventory becomes unknown when connection checks are disabled or fail, never an inferred withdrawal. acceptRate excludes unknown rows. --via public and --via api are not supported.
LinkedIn connection status batches: wonda linkedin connection-status <inputs...> returns exactly one result per input in the original order. Every row includes the exact original input. Lookup failures are labeled with status: "error" and error instead of being dropped. One input keeps the object shape; multiple inputs keep the array shape. A single lookup failure adds the labeled object on stdout while preserving the existing stderr error and nonzero exit.
Engager enrichment: wonda linkedin enrich-engagers --activity-id <id> scrapes reactors (and optionally commenters via --comments), then fetches each engager's profile + current employer + company page, and emits a single joined JSON document keyed by vanity with profile and currentEmployer (industry, headcount, HQ, description, employee count) blocks per engager. Use --company-detail=false to skip company page lookups and keep only inlined employer identity (name, urn, universalName). Use --max-profiles N to cap the batch (default 250, hard ceiling 250 unless --override-max-profiles is set) and --out file.json to write to disk. --profile-source cookies|public controls only per-profile detail enrichment; engager collection still uses logged-in LinkedIn access.
For ICP qualification of post engagers, run wonda skill get linkedin-icp-qualify.
A first-class platform with three transports selected by --via, the same legitimacy gradient as the others:
--via api— official Graph API via your connected OAuth account (--connection). ToS-safe, used for publishing.--via cookies— private mobile API via the local cookie--account. Used for reads (saved posts, comments).--via wab— browser DOM via the account's Wonda Automation Browser persona. Used for thecommentwrite (drives the reel's inline comment composer with a real-browser fingerprint, the same stealth path X reply / LinkedIn comment use).
Transports are per-operation capabilities: saved and comments are cookies-only (no Graph endpoint for them), post/carousel are api-only, and comment is wab-only. The two identities are distinct: --account/--sessionid = the local cookie identity; --connection = the OAuth instagram_account UUID. For --via wab the persona is auto-derived from --account (or pass --persona directly); the WAB injects the bound account's sessionid (+ ds_user_id) into the Chromium cookie jar at spawn.
⚠️ Same anti-fraud caution as the others: don't probe freshly-pasted cookies. The first request on a new
sessionidshould be the operation you wanted. Instagram flags burst activity from a new IP/process on a freshly-handed session.
--account selects the cookie file under ~/.wonda/instagram-cookies/<account>.json. For saved, carousels contribute every child's media URL (videos win over images for the per-item URL); pagination uses the max_id cursor (--cursor, --all, --max-pages, --delay <ms>). comments takes a /p/<code>/ or /reel/<code>/ URL (or a bare shortcode), decodes it to the numeric media id locally, then pages the same max_id cursor; the result carries each comment's id, text, authorHandle, authorName, createdAt, likeCount, replyCount plus the parent media's total commentCount. For posting, wonda instagram post --via api and wonda publish instagram share the same Graph-API path. comment (write) takes a /reel/<code>/ or /p/<code>/ URL (or bare shortcode) plus the text, and is wab-only: it auto-spawns the persona's WAB if needed, types into the inline composer, submits, and writes a comment audit row to ~/.wonda/wab/audit.jsonl (failures fire wab_action_failed telemetry and drop a failure bundle).
Reddit's transport is fixed per command kind, so --via is mostly not yours to choose here:
- Reads (search, subreddit, feed, user, user-posts, user-comments, post, trending, home, saved) run direct via a Chrome-fingerprinted Go HTTP client (fast, ~700ms p50). Cookies only.
--via wabis not available for reads and errors. - Writes (vote, comment, subscribe, save, unsave, delete, and subreddit
submit) dispatch through the account's Wonda Automation Browser so the shreddit GraphQL mutations carry a real-browser signal. WAB only.--via cookieserrors on these. - Submit to a profile self-post (
u_<handle>/u/<handle>) or a link post goes via the tls-client (cookies) only.--via wabis not available for those (no DOM submit URL), so--dry-run(DOM-only) does not apply to them either.
--account selects the cookie file under ~/.wonda/reddit-cookies/ (and, for writes, the account's auto-derived persona). You don't pass a persona here.
⚠️ Anti-fraud caution on freshly-pasted cookies.
wonda reddit auth checkis safe (it only decodes the JWT exp locally), but the FIRST read or write you fire on new cookies hits Reddit's API from your IP / process. If those cookies were last used elsewhere (different machine, different country), Reddit's anti-fraud trips the session-theft heuristic and may force-logout the cookies. Pattern: paste cookies, go straight to the operation the user wanted. Never do a "let me just check this works" round-trip first.
wonda reddit user includes additive emails and links fields when the user has published explicit contact details in their public About text. These fields are read-only and equivalent to manually reading the profile. No inferred emails, enrichment data, storage, or DOM scraping is involved.
Add --dry-run on a subreddit comment or submit to type into the composer but not click Post (useful for review). It is DOM-only, so it does not apply to profile self-posts or link posts. For subreddit submits, pass --flair "<label>" to select post flair in the composer. Flair labels match case-insensitively with normalized whitespace, first by exact label and then by a unique contains fallback. The picker expands View all flairs, scrolls the requested row clear of the footer, and verifies the committed flair ID or exact label before posting. If a subreddit requires flair and it is omitted, or if the requested flair does not exist, the command fails with the available flair labels. --flair is not supported with profile self-posts or --url link posts. Native image/video submits accept --media with an optional --text body and verify it remains in the composer before submitting. Media-only and text-only submits keep their existing behavior; --url still cannot combine with text or media.
wonda reddit signup provisions a brand-new Reddit account: it mints a throwaway mailbox (or uses --email), drives the 5-stage register form (email, emailed code, username plus password, age, interests) in a headful WAB window, fetches the verification code from the inbox subject line, then binds the persona and syncs cookies so reddit reads and writes route through the new account. Gated by the redditAccountCreationEnabled flag (server-evaluated preflight: GET /reddit/signup/enabled). Bind a mobile or residential proxy to the persona first (wonda wab config set <persona> proxy_url socks5://...): Reddit shadowbans accounts born on datacenter IPs. If a field cannot be located the flow pauses and leaves the window on that screen; finish by hand, then re-run with --resume <step>. On success it prints {username, email, password, persona} plus a ready-to-paste op item create block for the "Reddit logins" vault.
Paginated commands support: -n <count>, --after <cursor>, --all, --max-pages, --delay <ms>.
Feed-engage author mode (wonda {linkedin,x,reddit,instagram} feed-engage): scroll the home feed like a human and engage only the posts that scroll past from your target authors. On reddit that engagement is an upvote, on the others a like. It is WAB-only (it drives a live browser scroll + click through the account's persona, so there is no cookie path) and opportunistic: it never opens profiles or searches, it just rides the feed and acts when a target's post appears. Pass the targets with --authors "alice,bob" (comma-separated handles/vanities/usernames, leading @ and u/ are stripped) or --authors-file <path> (one author per line, local use only). --duration caps the wall-clock browse time (e.g. 5m, 90s, default 2m) and --max-engage caps the number of successful likes/upvotes (default 8); whichever limit hits first stops the run. Pacing is human (eased scrolling, randomized dwell, jittered cursor motion), so let it run for the full duration rather than expecting an instant result.
LinkedIn author mode can also opt into comment reactions with --engage-comments: by default it reacts only to comments from --authors, --engage-comments-from any includes all visible comments, and --max-comment-engage caps comment reactions per engaged post (default 3).
LinkedIn engage-commenters (wonda linkedin engage-commenters): read commenters on one LinkedIn post, then run the selected per-commenter actions in order: like, reply, connect. Pass the post with --post <activity-id|urn|url>. --actions accepts like,reply,connect and defaults to like; --reply-text is required when reply is selected; --connect-note is optional; --max-commenters defaults to 25; --duration defaults to 3m; --dry-run resolves targets and opens controls/composers without submitting writes. Access tier: paid/WAB. Transport: commenter reads use the established LinkedIn DOM comments path; every write uses WAB DOM primitives.
Feed-engage monitor mode (wonda {linkedin,x,reddit} feed-engage --keywords ...): run one bounded keyword/intent monitor pass, score unseen candidates with /text/generate, and optionally generate persona replies with /text/generate. --keywords enables monitor mode and is mutually exclusive with --authors and --authors-file. Reddit supports --subreddits SaaS,marketing to scope search. Without --reply, monitor mode is read-only and prints candidates plus relevance verdicts. With --reply --dry-run, it generates reply text but posts nothing and leaves the ledger unchanged. With --reply, it posts through DOM/WAB writers only: Reddit comment composer, X reply composer, and LinkedIn comment composer. Reads use cookies/tls-client where available: Reddit search and X search use the platform clients; LinkedIn content search uses the existing WAB DOM search-posts flow because LinkedIn content search is not reliable through the cookie API. Safety controls: --max-scan (default 50), --relevance-threshold (default 0.7), --max-reply (default 3), --per-day-cap (default 8), and a persistent dedupe ledger under ~/.wonda/reply-ledger/. Generated public replies are post-processed to remove em dashes. Monitor mode is not supported for Instagram.
Cross-platform DM coverage
Reddit chat / DMs
Direct messaging through the WAB. Chat uses the logged-in Reddit browser session
and is available as first-class twin actions with --via wab. There is no
separate chat token.
Important: Rate limit DM sends to 15-20/day with varied text to avoid detection. The send command drives the browser composer and verifies the rendered message.
Cloud digital twins (wonda twin)
Manage cloud-hosted social personas that run behind mobile proxies. Sessions are server-side; schedules drive recurring tasks (saved-content sync, engagement, agent runs) on a cron.
Interactive (watchable) vs headless: know which surface you are on. A cloud twin runs the SAME antidetect Chromium two ways, and only some of them stream a live screen you (or the user) can watch and click into:
- Watchable, interactive surfaces (a live cloud browser streamed to a
viewerUrlyou open in the local WAB or any browser):wonda twin login,wonda twin view,wonda twin signup, andwonda twin attach. Use these whenever a human must SEE and DRIVE the cloud browser: sign in, re-authenticate, create a brand-new account, or solve anything a human has to touch.login/view/signupSTART a fresh run and take the twin's profile lease;attachis different — it hooks onto an ALREADY-running run by its<runId>(fromwonda twin runs) and NEVER takes the lease, so an operator can watch (or take control of) ANY live run on demand — an autopilot run, a schedule firing, arun-nowcommand — not only a session they just started. A CAPTCHA (or any step-up challenge) can only be solved on one of these watchable surfaces. You cannot type credentials or solve a captcha for the user; hand them theviewerUrland let them finish, then re-checkwonda twin login-status/health/liveness. - Headless surfaces (no screen, no viewer, unattended):
wonda twin run-now, schedules,wonda twin login-status,wonda twin seed-from-cookies, and every platform action run with--engine cloud. These run invisibly behind the proxy; nobody is watching. If one hits a captcha or a login wall it cannot be solved in place, so the run just FAILS (fetch itswonda twin artifact <runId>screenshot to see why), OR — while it is still live — youwonda twin attach <runId> --control --opento jump into it and drive it yourself, then let it continue. - To take control of a live run, use
wonda twin attach <runId> --control. There is nowonda twin controlverb; drive-capable control of an arbitrary live run happens throughattach --control(or, for a view you started,wonda twin view).attachwithout--controlis a read-only watcher. - Watching / screenshotting a streamed viewer (agents): use the WAB, never a bare Playwright or headless-Chromium script.
login/view/signupopen the viewer as a tab INSIDE the local WAB; that tab IS the observation surface, sowonda wab show <persona>surfaces the window andwonda wab screenshot <persona>captures a PNG without surfacing it. For a read-onlyattachcapture, omit both--controland--open, then runwonda wab screenshot <viewerUrl> --output run.png --wait-until domcontentloaded --wait-for '#screen[width][height]'. The attach viewer addswidthandheightattributes to#screenonly after its first streamed image has loaded into the canvas, so this waits for real frame content instead of merely waiting for the viewer shell or idle network. The broker fans frames out to multiple read-only viewers, so this screenshot can run alongside other viewers. Only control is single-winner: if several controller-role viewers connect, the first keeps control and later ones are demoted to view-only. First confirm the run is actually streaming (wonda twin liveness <persona>→live:true) before building any observer at all.
wonda twin login-automated is NOT a headless happy path. The route never finishes a login on its own; it ALWAYS mints a streamed-login fallback and returns { status: "needs_human", viewerUrl }. Treat it as "open this viewerUrl and sign in by hand", the same as twin login.
Two gotchas that cause false "it's running / it's signed in" reports:
--engine cloudparks a warm control session that holds a ~15-min busy lease. After a cloud action the twin stays "busy" for the idle window, so a followingtwin login/view/signup(or another--engine cloudaction) can 409 as busy until the lease decays. This is expected warm-session behavior, not a bug; wait it out orwonda twin stop <persona>.- The "You are now signed in" panel in the viewer is sticky. It confirms the login WAS detected in that session; it does NOT mean a live stream is open now. For "is a live interactive stream open THIS SECOND" use
wonda twin liveness <persona>(flips false ~30s after the tab closes); for run state usewonda twin runs/wonda twin health. Never infer "still running / still watching" from the sticky panel.
The MCP twin surface tags read tools as Always allow and write tools as Needs approval. Use run_campaign and schedule_loop for one-approval autopilot loops; use the per-verb platform tools for supervised actions.
Claude web and Cowork users can connect the cloud twin as a custom remote MCP connector from https://wonda.sh/docs/connect-claude.
Local WAB relay (wonda relay)
Run registered twin actions on this machine's local WAB while callers still use the cloud broker and the same /twin/sessions/{persona}/actions/{platform}/{action} API.
relay health and relay restart require Wonda CLI v1.57.0 or newer. Before calling either command, check wonda --version. On an older or unknown version, do not attempt the unavailable command: use wonda doctor --relay instead of relay health. Do not substitute relay install for relay restart, because reinstalling without every configured --persona can replace the relay's persona set. Report that a safe CLI restart requires v1.57.0+ and ask the user to upgrade when that release is available.
The relay runs on your real device and IP. Platform cookies stay local. It only serves actions while the machine and wonda relay run are up. If no live relay is present, the broker routes the same twin action call to the cloud twin instead.
Safety is intrinsic to running anything on a twin. Account safety is not a special feature of the sequencer: it is a property of "run a command on a twin." The SAME per-identity safety gate guards an ad-hoc agent (wonda linkedin connect --account <name>), the autopilot, a twin schedule, and a user-authored sequence, through ONE enforcement path against ONE shared per-identity counter. So a heavy ad-hoc day on a persona tightens the headroom on that persona automatically (one set of ceilings per identity, no double-charge). The caps are a per-twin MODE (warmup / conservative_steady (default) / moderate_max / unlimited; set via wonda twin limits set) selecting the per-action daily ceilings, the per-action rolling-7d weekly ceilings, AND a global daily aggregate cap; unlimited turns every cap off, and a custom override is UNCLAMPED. The gate classifies each command from its argv: WRITE / social-action commands (connect, send-message, like, comment, visit, ...) are capped; reads (connection-status, conversations, profile, search), generation (image/video/text), and unknown commands pass freely ungated. This is why the SENSE verbs exist: an agent asks can-act / actions / health BEFORE firing, then branches on the typed result instead of attempting a write and catching a deny.
General sequences (MCP and public API)
A sequence is a durable Wonda workflow, not an outreach campaign. It has exactly four composable step kinds:
run: one real Wonda argv. The generated CLI manifest validates command paths, nested Sales Navigator/X DM/Reddit chat verbs, long flags and positional arity.wait: exact or ranged spacing such as2h ± 15mor3m-9m, sampled deterministically.if: ordinary conditions over{{vars.*}}, earlier{{steps.<id>.*}}results and current-item bindings. AND has normal precedence over OR.each: sequential bounded fan-out over any array, with{{item}}and an optional alias. Default 500 items, explicit maximum 5,000.
Use the MCP tools sequence_validate, sequence_create, sequence_run, sequence_runs, sequence_update, sequence_cancel_run and sequence_schedule. Validate before create. Every required {{vars.*}} value is checked before a manual or scheduled run is created. Scheduling is optional and accepts either a validated 5-field cron plus IANA timezone or a one-shot instant, deterministic jitter, stored vars, and explicit allowOverlap (default false).
Relay execution is manifest-driven, not registry-limited: every generated LinkedIn, Sales Navigator, X and Reddit command can run on the paired machine, including local-only commands such as X DMs. transport: auto falls back from a Premium cloud preference to relay when a definition requires one of those local-only commands. Each run snapshots its definition, resolved transport, organization billing scope and trigger, so edits or context changes never rewrite work already in flight.
Uniform error taxonomy (branch on code, never on the message). Every twin/outreach surface returns the existing { error: { code, message } } envelope, and on a quota deny it also carries deferUntil (the ISO time the quota resets) + reason (the granular gate signal). Agents branch on the typed code:
wonda twin can-act returns the SAME code + reason for a deny that a later write would, so an agent that reads code: "needs_auth" from can-act and then sees needs_auth on a write branches on the one code.
Calendar (wonda calendar)
One read-only feed over two things that otherwise have to be checked separately: upcoming twin schedule fires, and past runs and platform actions. Resolved in a real timezone, so "Tuesday" means the user's Tuesday.
A range returns per-day COUNTS only; calendar day returns the individual entries. That split is deliberate: a busy account has tens of thousands of rows in a month and a month grid only ever shows counts.
Two axes that must never be added together: counts.run is twin runs, actions is rows in the source-neutral action ledger. One run performs many platform actions, and an action run on the user's own machine through the relay performs no cloud run at all. Under-reporting is stated rather than implied: truncated.scheduleFires means a schedule's fire enumeration hit its ceiling, so those counts are a floor.
Workflow & discovery
Brand extraction (brand extract)
Extract a website's design system (colors, typography, radii, shadows, spacing, fonts, logo, hero decor, CSS pattern backgrounds, dashed/dotted border treatments, :root custom properties, headline emphasis pattern, film-grain/noise overlay) into a DESIGN.md + tokens.json + assets/. Runs locally via the bundled stealth browser + Chromium driver (the same wonda wab install as wonda wab record and the authenticated session flows).
Requires a one-time wonda wab install to download the stealth browser + Chromium (~300 MB, shared across wonda wab record, the authenticated session flows, and brand extract).
This is the in-house replacement for the previous npx-based brand-extraction CLI used in the slide-generation / creative-static-ads / premium-static-ads skills.
Flags:
--save: uploadassets/via the media presign flow and POST{tokens, mediaIds}to/api/v1/brand/save. Requires auth.--make-active: implies--save. Sets the new brand as active.--output <dir>: override the local output dir. Default is./output/<domain>/. Mutually exclusive with--no-output.--no-output: don't write to disk (in-memory extract for piping). Mutually exclusive with--output.--name "Brand Name": override the brand name when persisting. Defaults to the domain stem capitalized.--screenshot: also savepage.pngalongside DESIGN.md.--viewport WxH: viewport size for the headless browser. Default1920x1080.
Outputs (when --no-output is not set, always to <output-dir>/<domain>/):
DESIGN.md: Markdown summary of tokens, typography, hero decor, logo, CSS patterns, dashed borders, and root CSS variables. Read this in the slide / static-ad skills before composing HTML.tokens.json: raw structured JSON of the extraction.page.png: only when--screenshotis passed.assets/: raw hero decor files plusassets/fonts/for any non-Google@font-faceURLs. Always written when not--no-output.
Prints written file paths to stdout. With --save, also prints the API response (brandId, sourceDomain, warnings). Non-zero exit on failure (network error, navigation timeout, browser crash, save failure).
Video analysis
Analyze a video to extract a composite frame grid (visual) and audio transcript (text). Useful for understanding video content before creating variations. Requires a full account (not anonymous) and costs credits based on video duration (ElevenLabs STT pricing).
If the video was just uploaded and is still normalizing, the CLI auto-retries until the media is ready.
Error handling: 402 = insufficient credits, 409 = media still processing (CLI auto-retries).
Manage throwaway email accounts and read mailbox messages. These commands require the emailServerApiEnabled flag.
For signup flows, capture --since before triggering the verification email, or snapshot the current max message id with wonda email inbox list <email> --jq '[.[].id] | max // 0' and pass it to --since-id.
Jobs
Discovery
Editing audio & images
For any image edit (crop, text overlay, img2img, background removal, vectorize) pull the dedicated skill: wonda skill get image-edit.
Alignment (timestamp extraction)
Quality tiers
Troubleshooting
Timing expectations
- Image: 30s - 2min
- Video (Sora): 2 - 5min
- Video (Sora Pro): 5 - 10min
- Video (Veo 3.1): 1 - 3min
- Video (Kling): 3 - 8min
- Video (Grok): 2 - 5min
- Music (Suno): 1 - 3min
- TTS: 10 - 30s
- Editor operations: 30s - 2min
- Lip sync: 1 - 3min
- Video upscale: 2 - 5min
Error recovery
- Unknown model:
wonda models list - No API key:
wonda auth loginor setWONDA_API_KEYenv var - Job failed:
wonda jobs get inference <id>for error details - Bad params:
wonda models info <slug>for valid params - Timeout:
wonda jobs wait inference <id> --timeout 20m - Insufficient credits (402):
wonda topupto add credits LinkedIn WAB attachments: uselinkedin send-message <target> [text] --attach <local-path>with one supported local file (bmp/gif/jpeg/jpg/png/doc/docx/pdf/mp4/m4a, max 20,000,000 bytes). Add--dry-runto stage the upload without sending. Local paths are rejected by cloud mode.

