Bright Data — Scrape
Get clean content (markdown, HTML, JSON, screenshot) from one or more URLs via the Bright Data CLI. This skill owns the "fetch raw or lightly-structured content" job. For platform-specific structured data (Amazon, LinkedIn, TikTok, etc.), stop and use data-feeds instead — you'll get clean JSON without selector logic.
Setup gate (run first)
Before any scrape, verify the CLI is installed and authenticated:
If either check fails, halt and route the user to skills/bright-data-best-practices/references/cli-setup.md. Do not attempt the legacy curl fallback silently — ask the user first.
Pick your path
Action
Core commands:
Full flag reference: references/flags.md [blocked].
Verification gate (run before claiming success)
- Non-empty output:
test -s "$out_path"— or, for stdout, at least 200 bytes of content. - Not a block page — grep the output for any of these signatures (case-insensitive):
Access DeniedJust a momentAttention RequiredChecking your browsercaptchacf-browser-verificationcloudflare(with < 2KB total body)
- Expected markers present for the task: e.g., a product page should contain a price pattern (
\$\d); an article should contain at least one<h1>or#heading. - On failure, escalation ladder:
- Retry with a different
--country(e.g.,--country deif the origin site is US) - Escalate to
bdata browserfor full JS rendering (hand off tobrightdata-cliskill)
- Retry with a different
Do not report success until all checks above pass.
Red flags
- Claiming success without inspecting the output.
- Silencing errors with
2>/dev/null— you'll miss auth failures and rate-limit errors. - Running
bdata scrapeon Amazon/LinkedIn/TikTok/Instagram/YouTube/Reddit URLs — these are supported bydata-feedsand return structured data directly. Scraping loses the structure. - Scraping the same URL repeatedly in the same task — cache the first result.
- Looping
bdata scrapesequentially for large lists instead of usingxargs -P 4(or similar) with a parallelism cap. - Using
curlagainstapi.brightdata.comdirectly — legacy path; only when the CLI isn't available.
References
references/flags.md[blocked] — every flag with when-to-use notes.references/patterns.md[blocked] — shell-loop batching,xargsparallelism, pagination recipe, retry/backoff, block-page recovery chain, legacycurlfallback.references/examples.md[blocked] — (1) single page → markdown, (2) batch a list of URLs with parallelism cap, (3) paginated listing, (4) block-page recovery.



