Web Content Fetcher
Given a URL, return its main content as clean Markdown — headings, links, images, lists, code blocks all preserved.
Extraction Strategy
Always try one method per URL — don't cascade blindly. Pick the right one upfront.
Scrapling script
<SKILL_DIR> is the directory where this SKILL.md lives. Resolve it before calling the script.
The script has two modes built in:
- Default (fast): HTTP fetch, ~1-3s, works for most sites
--stealth: Headless browser, ~5-15s, for JS-rendered or anti-scraping sites
When run without --stealth, the script automatically falls back to stealth if the fast result has too little content. So you rarely need to specify --stealth manually — the only reason to force it is when you already know the site needs it (see routing table), which saves the initial fast attempt.
Domain Routing
Use this table to pick the right mode on the first call:
Script Options
Install Dependencies
First use only — the script checks and tells you if anything is missing:
If on system-managed Python (macOS/Linux), add --break-system-packages or use a venv.
Failure Rules
- Same URL fails once → give up, tell the user "unable to extract content from this URL"
- Do not retry — each failed call wastes context tokens
