Firecrawl Knowledge Ingest

by firecrawl94cc91229d6cISC185 starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 6 weeks ago

Ingest public or authenticated knowledge bases and docs portals with Firecrawl browser. Use for JS-heavy docs, login-gated portals, paginated help centers, support knowledge bases, or structured JSON/markdown extraction from documentation sites.

Instructions onlyResearch & Analysis
AI-generated overview

Scrapes documentation portals and knowledge bases with Firecrawl, including JS-heavy and login-gated pages, into structured markdown or JSON.

What it does
This skill guides an agent through collecting content from a documentation portal or knowledge base using the Firecrawl browser. It inspects navigation, follows sidebar links, pagination and load-more controls, scrapes article content as markdown, and extracts metadata such as title, section, last updated date, author and tags. It produces a structured deliverable with a summary, section and article counts, failed or restricted pages, source URLs, rerun inputs, and JSON using source, url, extractedAt, totalArticles and sections fields.
When to use it
Use it when a docs portal or help center needs browser navigation, authentication, pagination or JavaScript rendering. It suits JS-heavy documentation, login-gated portals, paginated help centers, support knowledge bases, and structured JSON or markdown extraction from documentation sites.
Requirements
Requires a Firecrawl API key supplied as FIRECRAWL_API_KEY for hosted Firecrawl requests, plus network access to the target portal. It ships no scripts; it is instructions only.

Firecrawl Knowledge Ingest

Use this when a docs portal needs browser navigation, auth, pagination, or JS rendering.

Onboarding Interview

Infer the portal URL, output format, auth needs, and page limit from context. If the portal is clear, proceed immediately.

Ask at most 1-3 concise questions only if blocked, such as the portal URL, whether authentication is required, or the desired output format.

Firecrawl Collection Plan

Use Firecrawl browser to:

  • open the portal and inspect navigation
  • identify sections, categories, sidebar links, and article URLs
  • follow sidebar navigation, next links, pagination, load-more controls, or search
  • scrape article content as markdown
  • extract metadata such as title, section, last updated date, author, and tags

Try Firecrawl map as a supplement for public URLs, but use browser navigation for auth-gated or JS-heavy content.

Final Deliverable

markdown
# Knowledge Ingest: [Portal]
## Summary[Pages extracted, sections covered, limitations]
## Output[JSON/markdown/merged file path or content]
## Sections[Section names and article counts]
## Failed Or Restricted Pages[Any access/loading issues]
## Sources[URLs extracted]
## Rerun Inputsworkflow: firecrawl-knowledge-ingesturl: [portal url]format: [json/markdown/merged]max_pages: [number]

JSON Shape

Use source, url, extractedAt, totalArticles, and sections[] with article title, url, section, content, and metadata.

Quality Bar

  • Preserve code examples, tables, and formatting.
  • Strip nav chrome, headers, and footers.
  • Track extraction progress and page failures.
  • Respect authentication boundaries.

Source and attribution

Source:firecrawl/firecrawl-workflowsinskills/firecrawl-knowledge-ingestat commit94cc912

License: ISC

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal

Firecrawl Knowledge Ingest Agent Skill | SourceWeft