Firecrawl Knowledge Ingest

作者 firecrawl94cc91229d6cISC185 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫6 週前更新

Ingest public or authenticated knowledge bases and docs portals with Firecrawl browser. Use for JS-heavy docs, login-gated portals, paginated help centers, support knowledge bases, or structured JSON/markdown extraction from documentation sites.

AI 產生的概覽

使用 Firecrawl 擷取文件入口網站與知識庫,包括大量使用 JavaScript 及需登入的頁面,輸出結構化 Markdown 或 JSON。

功能
此技能引導代理使用 Firecrawl 瀏覽器從文件入口網站或知識庫蒐集內容。它會檢視導覽結構,追蹤側邊欄連結、分頁與「載入更多」控制項,將文章內容擷取為 Markdown,並擷取標題、章節、最後更新日期、作者與標籤等中繼資料。它產出結構化交付內容,包含摘要、章節與文章數量、失敗或受限頁面、來源網址、重跑輸入,以及使用 source、url、extractedAt、totalArticles 與 sections 欄位的 JSON。
適用情境
當文件入口網站或說明中心需要瀏覽器導覽、身分驗證、分頁或 JavaScript 渲染時使用。適用於大量使用 JavaScript 的文件、需登入的入口網站、分頁說明中心、支援知識庫,以及從文件網站進行結構化 JSON 或 Markdown 擷取。
執行需求
需要以 FIRECRAWL_API_KEY 提供的 Firecrawl API 金鑰,用於託管的 Firecrawl 請求,並需要連線至目標入口網站的網路存取。此技能不附帶指令碼,僅為指示文件。

Firecrawl Knowledge Ingest

Use this when a docs portal needs browser navigation, auth, pagination, or JS rendering.

Onboarding Interview

Infer the portal URL, output format, auth needs, and page limit from context. If the portal is clear, proceed immediately.

Ask at most 1-3 concise questions only if blocked, such as the portal URL, whether authentication is required, or the desired output format.

Firecrawl Collection Plan

Use Firecrawl browser to:

  • open the portal and inspect navigation
  • identify sections, categories, sidebar links, and article URLs
  • follow sidebar navigation, next links, pagination, load-more controls, or search
  • scrape article content as markdown
  • extract metadata such as title, section, last updated date, author, and tags

Try Firecrawl map as a supplement for public URLs, but use browser navigation for auth-gated or JS-heavy content.

Final Deliverable

markdown
# Knowledge Ingest: [Portal]
## Summary[Pages extracted, sections covered, limitations]
## Output[JSON/markdown/merged file path or content]
## Sections[Section names and article counts]
## Failed Or Restricted Pages[Any access/loading issues]
## Sources[URLs extracted]
## Rerun Inputsworkflow: firecrawl-knowledge-ingesturl: [portal url]format: [json/markdown/merged]max_pages: [number]

JSON Shape

Use source, url, extractedAt, totalArticles, and sections[] with article title, url, section, content, and metadata.

Quality Bar

  • Preserve code examples, tables, and formatting.
  • Strip nav chrome, headers, and footers.
  • Track extraction progress and page failures.
  • Respect authentication boundaries.

來源與署名

來源:firecrawl/firecrawl-workflows位於skills/firecrawl-knowledge-ingest提交94cc912

授權條款: ISC

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架