Firecrawl Knowledge Ingest

作者 firecrawl94cc91229d6cISC185 个星标收录于 2026年10月8日更新于 2026年10月8日仓库6周前更新

Ingest public or authenticated knowledge bases and docs portals with Firecrawl browser. Use for JS-heavy docs, login-gated portals, paginated help centers, support knowledge bases, or structured JSON/markdown extraction from documentation sites.

AI 生成的概览

使用 Firecrawl 抓取文档门户和知识库,包括重度依赖 JavaScript 和需登录的页面,输出结构化 Markdown 或 JSON。

功能
该技能指导智能体使用 Firecrawl 浏览器从文档门户或知识库采集内容。它会检查导航结构,跟踪侧边栏链接、分页和“加载更多”控件,将文章内容抓取为 Markdown,并提取标题、章节、最后更新日期、作者和标签等元数据。它产出结构化交付物,包含摘要、章节与文章数量、失败或受限页面、来源 URL、重跑输入,以及使用 source、url、extractedAt、totalArticles 和 sections 字段的 JSON。
适用场景
当文档门户或帮助中心需要浏览器导航、身份验证、分页或 JavaScript 渲染时使用。适用于重度依赖 JavaScript 的文档、需登录的门户、分页帮助中心、支持知识库,以及从文档站点进行结构化 JSON 或 Markdown 提取。
运行要求
需要以 FIRECRAWL_API_KEY 提供的 Firecrawl API 密钥,用于托管的 Firecrawl 请求,并需要访问目标门户的网络连接。该技能不附带脚本,仅为说明文档。

Firecrawl Knowledge Ingest

Use this when a docs portal needs browser navigation, auth, pagination, or JS rendering.

Onboarding Interview

Infer the portal URL, output format, auth needs, and page limit from context. If the portal is clear, proceed immediately.

Ask at most 1-3 concise questions only if blocked, such as the portal URL, whether authentication is required, or the desired output format.

Firecrawl Collection Plan

Use Firecrawl browser to:

  • open the portal and inspect navigation
  • identify sections, categories, sidebar links, and article URLs
  • follow sidebar navigation, next links, pagination, load-more controls, or search
  • scrape article content as markdown
  • extract metadata such as title, section, last updated date, author, and tags

Try Firecrawl map as a supplement for public URLs, but use browser navigation for auth-gated or JS-heavy content.

Final Deliverable

markdown
# Knowledge Ingest: [Portal]
## Summary[Pages extracted, sections covered, limitations]
## Output[JSON/markdown/merged file path or content]
## Sections[Section names and article counts]
## Failed Or Restricted Pages[Any access/loading issues]
## Sources[URLs extracted]
## Rerun Inputsworkflow: firecrawl-knowledge-ingesturl: [portal url]format: [json/markdown/merged]max_pages: [number]

JSON Shape

Use source, url, extractedAt, totalArticles, and sections[] with article title, url, section, content, and metadata.

Quality Bar

  • Preserve code examples, tables, and formatting.
  • Strip nav chrome, headers, and footers.
  • Track extraction progress and page failures.
  • Respect authentication boundaries.

来源与署名

来源:firecrawl/firecrawl-workflows位于skills/firecrawl-knowledge-ingest提交94cc912

许可证: ISC

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架