Web Fetch

alleneubank/claude-code/.claude/skills/web-fetch

作者 alleneubank2921eb8a685a無授權條款52 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫3 個月前更新

Fetches web content as clean markdown by preferring markdown-native responses and falling back to selector-based HTML extraction. Use for documentation, articles, and reference pages at http/https URLs.

已封存包含腳本Research & Analysis
AI 產生的概覽

抓取網頁並轉換成乾淨的 Markdown,優先採用原生 Markdown 回應,其次使用 CSS 選擇器或內附的 Bun 指令碼。

功能
此技能會從 http/https 網址取得內容,並以乾淨的 Markdown 回傳。它會先檢查回應是否為原生 Markdown,否則使用 html2markdown 搭配包含與排除的 CSS 選擇器擷取 HTML,並提供已知文件網站的選擇器表。當選擇器效果不佳時,會執行內附的 Bun 指令碼作為備援解析器。
適用情境
適合將文件、文章與參考頁面抓取成 Markdown,以便閱讀或後續處理。適用於需要乾淨文字版本而非原始 HTML 的頁面。不適用於只回傳載入中佔位內容的 JavaScript 渲染頁面。
執行需求
需要 curl、html2markdown 與 Bun,以及連線至目標網址的網路。內附的 fetch.ts 指令碼需以 bun install 安裝 Bun 相依套件。此技能附有可執行指令碼。

Web Content Fetching

Fetch web content in this order:

  1. Prefer markdown-native endpoints (content-type: text/markdown)
  2. Use selector-based HTML extraction for known sites
  3. Use the bundled Bun fallback script when selectors fail

Prerequisites

Verify required tools before extracting:

bash
command -v curl >/dev/null || echo "curl is required"command -v html2markdown >/dev/null || echo "html2markdown is required for HTML extraction"command -v bun >/dev/null || echo "bun is required for fetch.ts fallback"

Install Bun dependencies for the bundled script:

bash
cd ~/.claude/skills/web-fetch && bun install

Default Workflow

Use this as the default flow for any URL:

bash
URL="<url>"CONTENT_TYPE="$(curl -sIL "$URL" | awk -F': ' 'tolower($1)=="content-type"{print tolower($2)}' | tr -d '\r' | tail -1)"
if echo "$CONTENT_TYPE" | grep -q "markdown"; then  curl -sL "$URL"else  curl -sL "$URL" \    | html2markdown \        --include-selector "article,main,[role=main]" \        --exclude-selector "nav,header,footer,script,style"fi

Known Site Selectors

SiteInclude SelectorExclude Selector
platform.claude.com#content-container-
docs.anthropic.com#content-container-
developer.mozilla.orgarticle-
github.com (docs)articlenav,.sidebar
Genericarticle,main,[role=main]nav,header,footer,script,style

Example:

bash
curl -sL "<url>" \  | html2markdown \      --include-selector "#content-container" \      --exclude-selector "nav,header,footer"

Finding the Right Selector

When a site isn't in the patterns list:

bash
# Check what content containers existcurl -s "<url>" | grep -o '<article[^>]*>\|<main[^>]*>\|id="[^"]*content[^"]*"' | head -10
# Test a selectorcurl -sL "<url>" | html2markdown --include-selector "<selector>" | head -30
# Check line countcurl -sL "<url>" | html2markdown --include-selector "<selector>" | wc -l

Universal Fallback Script

When selectors produce poor output, run the bundled parser:

bash
bun ~/.claude/skills/web-fetch/fetch.ts "<url>"

If already in the skill directory:

bash
bun fetch.ts "<url>"

Options Reference

bash
--include-selector "CSS"  # Keep only matching elements--exclude-selector "CSS"  # Remove matching elements--domain "https://..."    # Convert relative links to absolute

Troubleshooting

Empty output with selectors: The page might be markdown-native. Check headers first:

bash
curl -sIL "<url>" | grep -i '^content-type:'

Wrong content selected: The site may have multiple article/main regions:

bash
curl -s "<url>" | grep -o '<article[^>]*>'

html2markdown not found: Install it, then retry selector-based extraction.

bun or script deps missing: Run cd ~/.claude/skills/web-fetch && bun install.

Missing code blocks: Check if the site uses non-standard code formatting.

Client-rendered content: If HTML only has "Loading..." placeholders, the content is JS-rendered. Neither curl nor the Bun script can extract it; use browser-based tools.

來源與署名

來源:alleneubank/claude-code位於.claude/skills/web-fetch提交2921eb8

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架