Search

作者 brightdatae825f02fbcd7无许可证264 个星标收录于 2026年10月8日更新于 2026年10月8日仓库昨天更新

Search the web via the Bright Data CLI — `bdata search` for Google/Bing/Yandex SERP, `bdata discover` for intent-ranked semantic results. Use when the user wants SERP results, needs URLs to feed into scraping, or wants semantic web discovery with optional page content. Hands off to `scrape` once target URLs are chosen, and to `data-feeds` when the user wants structured data from a known platform. Requires the Bright Data CLI; proactively guides install + login if missing.

AI 生成的概览

通过 Bright Data CLI 指导网页搜索与语义发现,以 JSON 返回搜索结果或按意图排序的结果。

功能
该技能说明两个 Bright Data CLI 命令:bdata search 用于获取 Google、Bing 或 Yandex 的关键词搜索结果,bdata discover 用于按意图排序的语义结果并可附带页面内容。内容涵盖环境检查、参数选择、分页、本地化、新闻与图片等垂直搜索,以及返回 JSON 的校验。它还规定了在选定 URL 或平台后转交给抓取或结构化数据源的流程。
适用场景
当用户需要搜索引擎结果、需要可用于抓取的 URL,或需要带可选页面内容的语义网页发现时使用。它也用于在关键词搜索与基于意图的发现之间做选择。
运行要求
需要安装并已通过 bdata login 认证的 Bright Data CLI(bdata),以及运行搜索所需的网络访问。该技能不包含脚本,只有参考文档,并建议使用 jq 校验 JSON 输出。

Bright Data — Search

Find things on the web. Two commands live in this skill:

  • bdata search — classic keyword SERP (Google/Bing/Yandex). Best when you want "what ranks for keyword X."
  • bdata discover — AI intent-ranked discovery with optional page content. Best when you want "pages about topic Y that match intent Z."

For structured data from a known platform (Amazon, LinkedIn, TikTok, …), stop and use data-feeds instead.

Setup gate (run first)

bash
if ! command -v bdata >/dev/null 2>&1; then    echo "bdata CLI not installed — see bright-data-best-practices/references/cli-setup.md"elif ! bdata zones >/dev/null 2>&1; then    echo "bdata not authenticated — run: bdata login  (or: bdata login --device for SSH)"fi

Halt and route to skills/bright-data-best-practices/references/cli-setup.md if either check fails.

Pick your path

SituationAction
Single keyword query, just SERPbdata search "<query>" --engine google --json --pretty
Paginated SERP (more results)loop --page 0, --page 1, … (0-indexed)
Multiple queriesshell loop over a queries file
Intent-ranked / semantic (not keyword)bdata discover "<query>" --intent "<intent>" --num-results 20
Want page bodies along with results, one passbdata discover ... --include-content
News / images / shopping SERPbdata search "<query>" --type news (or images, shopping)
Want Amazon/LinkedIn/TikTok/… structured datastop — hand off to data-feeds
Have URLs, want contenthand off to scrape

Action

Core commands:

bash
# Google SERP, structured JSONbdata search "site:example.com privacy policy" --engine google --json --pretty
# Localized Bing (German results, German language)bdata search "datenschutz" --engine bing --country de --language de --json
# Second page of results (0-indexed)bdata search "machine learning papers" --page 1 --json
# Mobile SERP (rankings differ from desktop)bdata search "best coffee shops" --device mobile --json
# News verticalbdata search "openai" --type news --json --pretty
# Intent-ranked discoverybdata discover "enterprise LLM platforms" \    --intent "vendor pages with pricing" \    --num-results 15 --json
# Discovery with page content in markdownbdata discover "webhook best practices" \    --include-content --num-results 10 -o results.json
# Date-filtered discoverybdata discover "react server components" \    --start-date 2025-01-01 --end-date 2025-12-31 --num-results 20

Full flag reference: references/flags.md [blocked].

search vs discover — pick the right one

You wantUse
"What Google ranks for this exact keyword"search
"Pages that match this meaning/intent"discover
"News / images / shopping vertical SERP"search --type <vertical>
"Results + page bodies in one call"discover --include-content
"Dedup / semantic ranking across queries"discover

Verification gate

  1. JSON parses cleanly: jq . <output> returns 0.
  2. Result array non-empty — if empty, the query is legitimately zero-result; relax the query and re-run. Don't claim success on empty results without telling the user.
  3. Required fields present:
    • search: results live at .organic[]; each has title + link
    • discover: results live at .results[]; each has title + link; if --include-content, also content
  4. For discover --include-content: no block-page signatures in the content field (same list as scrape, case-insensitive):
    • Access Denied
    • Just a moment
    • Attention Required
    • Checking your browser
    • captcha
    • cf-browser-verification
    • cloudflare (with < 2KB total body)
  5. Geo sanity: if the user expected country-specific results, inspect TLDs / languages of top results. If mis-localized, re-run with explicit --country and --language.

Red flags

  • Using search to fetch content from Amazon, LinkedIn, TikTok, etc. when data-feeds returns clean structured data in one call.
  • Scraping every SERP result blindly — filter first (domain allowlist, keyword in title, relevance heuristic).
  • Confusing search (keyword) with discover (semantic). They answer different questions.
  • Running multiple queries without deduping URLs across result sets before scraping.
  • Assuming SERP order is universal — it's personalized by geo + device. Always set --country and --device explicitly for reproducibility.
  • Using --page as a result count — it's a page index, not a limit. Each page returns ~10 results.
  • Assuming SERP results are at .results[] — for bdata search they live at .organic[]. (Discover uses .results[].)
  • Hardcoding --num-results 100 on discover without realizing the pipeline polls until that many are found; can be slow.

References

  • references/flags.md [blocked] — full flags for search and discover with when-to-use notes.
  • references/patterns.md [blocked] — multi-query dedup, SERP → filter → scrape pipeline, search vs discover decision, legacy curl fallback, shared verification checklist.
  • references/examples.md [blocked] — (1) single Google query, (2) localized Bing, (3) batch queries + dedup into URL list, (4) discover --include-content end-to-end.

来源与署名

来源:brightdata/skills位于skills/search提交e825f02

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架