Firecrawl Knowledge Base

by firecrawl94cc91229d6cISC185 starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 6 weeks ago

Build a knowledge base from web content with Firecrawl. Use for local reference docs, RAG-ready chunks, fine-tuning datasets, documentation mirrors, topic corpora, or LLM-ready markdown organized from web sources.

AI-generated overview

Builds an organized, LLM-ready knowledge base from web sources using Firecrawl scraping, mapping and search.

What it does
This skill turns URLs or topics into organized, LLM-ready content by planning a Firecrawl collection using map, search and scrape into markdown. It produces one of several output modes: reference markdown with index.md and sources.json, RAG markdown with chunk files and manifest.json, training data with training-data.jsonl and training-metadata.json, or a full documentation mirror with a table of contents. It also defines a final deliverable report covering summary, output structure, coverage, usage notes, sources and rerun inputs, and sets quality rules such as preserving code examples and tables and removing boilerplate navigation.
When to use it
Use it when you need to assemble local reference docs, RAG-ready chunks, fine-tuning datasets, documentation mirrors or topic corpora from web sources. It fits requests to collect and organize web content into LLM-ready markdown.
Requirements
Requires a Firecrawl API key (FIRECRAWL_API_KEY) for hosted Firecrawl requests, plus network access to Firecrawl and the target sites. It ships no scripts; it is instructions only.

Firecrawl Knowledge Base

Use this to turn URLs or topics into organized LLM-ready content.

Onboarding Interview

Infer the source, goal, depth, and output location from context. If the source and goal are clear, proceed immediately.

Ask at most 1-3 concise questions only if blocked, such as the source URL/topic, whether the output is reference/RAG/training/docs, or training format if training is requested.

Firecrawl Collection Plan

Use Firecrawl map for documentation sites, search for topic-based corpora, scrape pages into markdown, and preserve code examples and tables.

For files, follow the Firecrawl download-style convention:

text
.firecrawl/  <hostname>/    <path>/      index.md

Parallel Work

If appropriate, use sub-agents or equivalent parallel task runners:

  • one docs section per researcher
  • official docs, tutorials, community discussions, and references by source type
  • source scraping vs chunk generation vs manifest generation

Output Modes

  • Reference: markdown files, index.md, and sources.json.
  • RAG: markdown files plus chunk files and manifest.json.
  • Training: scraped source files plus training-data.jsonl and training-metadata.json.
  • Docs mirror: complete markdown mirror with a table of contents.

Final Deliverable

markdown
# Knowledge Base: [Source]
## Summary[What was collected and why]
## Output Structure[Files/directories created]
## Coverage[Sections, source types, counts]
## Usage Notes[How to use in RAG, docs, training, or agent context]
## Sources[URLs collected]
## Rerun Inputsworkflow: firecrawl-knowledge-basesource: [url/topic]goal: [reference/rag/train/docs]depth: [quick/thorough/exhaustive]output_dir: [.firecrawl/]

Quality Bar

  • Preserve code examples and formatting.
  • Remove boilerplate navigation where possible.
  • Include source URLs in frontmatter or metadata.

Source and attribution

Source:firecrawl/firecrawl-workflowsinskills/firecrawl-knowledge-baseat commit94cc912

License: ISC

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal

Firecrawl Knowledge Base Agent Skill | SourceWeft