Firecrawl Knowledge Base

作者 firecrawl94cc91229d6cISC185 个星标收录于 2026年10月8日更新于 2026年10月8日仓库6周前更新

Build a knowledge base from web content with Firecrawl. Use for local reference docs, RAG-ready chunks, fine-tuning datasets, documentation mirrors, topic corpora, or LLM-ready markdown organized from web sources.

AI 生成的概览

使用 Firecrawl 抓取、映射和搜索网页来源,构建有条理、可供大模型使用的知识库。

功能
该技能通过规划 Firecrawl 采集流程(使用 map、search 和 scrape 转为 markdown),把网址或主题转化为有条理、可供大模型使用的内容。它支持多种输出模式:参考型 markdown 加 index.md 和 sources.json;RAG 型 markdown 加分块文件和 manifest.json;训练型数据加 training-data.jsonl 和 training-metadata.json;或带目录的完整文档镜像。它还规定了最终交付报告的结构,涵盖摘要、输出结构、覆盖范围、使用说明、来源和重跑输入,并设定质量要求,例如保留代码示例和表格、尽量去除样板式导航。
适用场景
当需要从网页来源整理本地参考文档、RAG 分块、微调数据集、文档镜像或主题语料时使用。它适合把网页内容收集并整理为可供大模型使用的 markdown 的请求。
运行要求
需要 Firecrawl API 密钥(FIRECRAWL_API_KEY)以调用托管的 Firecrawl 请求,并需要访问 Firecrawl 和目标站点的网络连接。该技能不附带脚本,仅为说明文档。

Firecrawl Knowledge Base

Use this to turn URLs or topics into organized LLM-ready content.

Onboarding Interview

Infer the source, goal, depth, and output location from context. If the source and goal are clear, proceed immediately.

Ask at most 1-3 concise questions only if blocked, such as the source URL/topic, whether the output is reference/RAG/training/docs, or training format if training is requested.

Firecrawl Collection Plan

Use Firecrawl map for documentation sites, search for topic-based corpora, scrape pages into markdown, and preserve code examples and tables.

For files, follow the Firecrawl download-style convention:

text
.firecrawl/  <hostname>/    <path>/      index.md

Parallel Work

If appropriate, use sub-agents or equivalent parallel task runners:

  • one docs section per researcher
  • official docs, tutorials, community discussions, and references by source type
  • source scraping vs chunk generation vs manifest generation

Output Modes

  • Reference: markdown files, index.md, and sources.json.
  • RAG: markdown files plus chunk files and manifest.json.
  • Training: scraped source files plus training-data.jsonl and training-metadata.json.
  • Docs mirror: complete markdown mirror with a table of contents.

Final Deliverable

markdown
# Knowledge Base: [Source]
## Summary[What was collected and why]
## Output Structure[Files/directories created]
## Coverage[Sections, source types, counts]
## Usage Notes[How to use in RAG, docs, training, or agent context]
## Sources[URLs collected]
## Rerun Inputsworkflow: firecrawl-knowledge-basesource: [url/topic]goal: [reference/rag/train/docs]depth: [quick/thorough/exhaustive]output_dir: [.firecrawl/]

Quality Bar

  • Preserve code examples and formatting.
  • Remove boilerplate navigation where possible.
  • Include source URLs in frontmatter or metadata.

来源与署名

来源:firecrawl/firecrawl-workflows位于skills/firecrawl-knowledge-base提交94cc912

许可证: ISC

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架