Data Scraper Agent

by affaan-mef648e01899bNo license275K starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 3 days ago

任意のパブリックソース(ジョブボード、価格、ニュース、GitHub、スポーツなど)用の完全自動化されたAI搭載データ収集エージェントを構築します。スケジュールでスクレイプし、無料LLM(Gemini Flash)でデータを豊かにし、Notion/Sheets/Supabaseに結果を保存し、ユーザーフィードバックから学習します。GitHub Actions上で100%無料で実行。ユーザーがパブリックデータを自動的に監視、収集、または追跡したい場合に使用します。

Instructions onlyAI & Agents
AI-generated overview

Builds automated AI data-collection agents that scrape public sources, enrich results with a free LLM, and store them on a schedule.

What it does
This skill guides the construction of a production-style data collection agent in three layers: collect, enrich, and store. It describes scraping public sites or APIs with Playwright or BeautifulSoup, enriching the collected text with Gemini Flash for scoring, summarising, or classification, and saving results to Notion, Sheets, or Supabase. It also covers scheduling runs with GitHub Actions and adding a feedback loop that learns from user decisions.
When to use it
Use it when someone wants to scrape or monitor a public website or API, track jobs, prices, news, repositories, sports scores, events, or lists, or automate data collection without paying for hosting. It also fits requests for an agent that improves over time based on user decisions.
Requirements
Instructions only; no scripts are shipped. It assumes Python, a scraping library such as Playwright or BeautifulSoup, a Gemini Flash API key, a storage target such as Notion, Google Sheets, or Supabase, and GitHub Actions for scheduling.

データスクレイパーエージェント

任意のパブリックデータソース用の本番環境対応、AI搭載データ収集エージェントを構築。 スケジュールで実行され、無料LLMで結果を豊かにし、データベースに保存し、時間とともに改善されます。

スタック:Python · Gemini Flash(無料) · GitHub Actions(無料) · Notion / Sheets / Supabase

アクティベーション時期

  • ユーザーが任意のパブリックWebサイトまたはAPIをスクレイプまたは監視したい場合
  • ユーザーが「チェックするボットを構築」「Xを監視」「データを収集」と言う
  • ユーザーがジョブ、価格、ニュース、リポ、スポーツスコア、イベント、リストを追跡したい場合
  • ユーザーがホスティング用に支払わずにデータ収集を自動化する方法を尋ねる
  • ユーザーが決定に基づいて時間とともにより スマートになるエージェントを望む

コアコンセプト

3つのレイヤー

すべてのデータスクレイパーエージェントには3つのレイヤーがあります:

COLLECT → ENRICH → STORE  │           │        │Scraper    AI (LLM)  Databaseruns on    scores/   Notion /schedule   summarises Sheets /           & classifies Supabase

無料スタック

LayerToolWhy
COLLECTPlaywright/BeautifulSoup無料のオープンソーススクレイピング
ENRICHGemini Flash無料で高速LLM
STORESupabase / Sheets無料データベースとスプレッドシート
SCHEDULEGitHub Actions無料クロンジョブ

ワークフロー

  1. ソースを定義 - どこからスクレイプするか、何を抽出するか
  2. スクレイパーを構築 - BeautifulSoup または Playwright ベースのコレクタ
  3. LLMを構成 - Gemini Flash でテキストをスコア付け/要約/分類
  4. ストレージを設定 - Notion、Sheets、Supabase のいずれか
  5. GitHub Actions を設定 - 毎日/毎週実行するスケジュール
  6. フィードバックループを追加 - ユーザーの判断から学習

例

  • ジョブボード監視:新しい公開

Source and attribution

Source:affaan-m/eccindocs/ja-JP/skills/data-scraper-agentat commitef648e0

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal