Defuddle

joeseesun/defuddle-skill/skills/defuddle

作者 joeseesunb1639f3627d776b20d93a8098d2c896d34bdc672無授權條款112 個星標收錄於 2026年10月9日更新於 2026年10月9日儲存庫7 個月前更新

Extract clean article content from web pages or local HTML files. Removes clutter (ads, sidebars, nav) and returns readable content with metadata.

AI 產生的概覽

從網頁或本機 HTML 檔案擷取乾淨的文章正文與中繼資料,並儲存為 Markdown。

功能
Defuddle 會從網頁或本機 HTML 檔案擷取主要文章內容,移除廣告、側邊欄與導覽列。它會回傳乾淨的 Markdown,以及標題、作者、來源、發布日期與字數等中繼資料。接著流程會把結果儲存為帶 frontmatter 的 Markdown 檔案,並向使用者確認檔案路徑。
適用情境
當使用者想擷取或清理網頁內容、從 URL 取得文章正文、去除 HTML 中的雜亂元素,或把網頁轉換成乾淨的 Markdown 時使用。它適合部落格、新聞與文件等文章類頁面。
執行需求
需要 Node.js 與 npm,並全域安裝 defuddle CLI,同時以 jsdom 作為同級相依套件。處理 URL 時需要網路存取。此技能未附帶指令碼,說明中僅描述 CLI 指令。

Defuddle - Web Content Extraction

Extract main article content from web pages, removing ads, sidebars, navigation, and other clutter. Output clean Markdown with metadata.

Prerequisites

Before first use, check if defuddle is installed:

bash
command -v defuddle >/dev/null 2>&1 || npm install -g defuddle jsdom

Default Workflow

When user provides a URL, follow this workflow:

Step 1: Extract content as Markdown + JSON metadata

Always use both -m and -j flags to get markdown content with full metadata:

bash
defuddle parse "<url>" -m -j

Step 2: Present a summary to the user

Show the user:

  • Title: from JSON title field
  • Author: from JSON author field
  • Source: domain
  • Word count: from JSON wordCount field
  • A brief preview (first 2-3 sentences)

Step 3: Ask where to save

If this is the first time using defuddle in this conversation, ask the user:

"Save to which directory? (e.g. ~/Documents, ~/Desktop, or a custom path)"

Remember the user's chosen directory for subsequent uses in the same conversation.

Step 4: Save as Markdown file

Write the file with frontmatter + full content:

markdown
---title: {title}author: {author}source: {url}date: {published or "Unknown"}clipped: {today's date YYYY-MM-DD}wordCount: {wordCount}---
# {title}
{markdown content}

File naming: Use the article title as filename, sanitized for filesystem:

  • Replace special characters with spaces
  • Trim whitespace
  • Example: The Shape of the Essay Field.md

Step 5: Confirm to user

Tell the user the file path where it was saved.

CLI Reference

bash
defuddle parse <source> [options]

Arguments:

  • <source> — URL (https://...) or local HTML file path

Options:

FlagDescription
-m, --markdownConvert content to Markdown
-j, --jsonOutput as JSON with full metadata
-o, --output <file>Write to file instead of stdout
-p, --property <name>Extract single property (title, description, domain, author, published, wordCount, content)
--debugVerbose logging

JSON Response Fields

When using -j, the response includes:

  • title — Article title
  • author — Author name
  • published — Publication date
  • description — Meta description
  • content — Extracted Markdown (when -m used)
  • domain — Source domain
  • favicon — Favicon URL
  • image — Featured image URL
  • site — Site name
  • wordCount — Word count
  • parseTime — Processing time in ms

Notes

  • Requires Node.js and npm
  • jsdom is required as a peer dependency
  • Works best with article-style pages (blogs, news, documentation)
  • Not designed for SPAs or JavaScript-heavy pages (e.g. WeChat articles need browser rendering)

來源與署名

來源:joeseesun/defuddle-skill位於skills/defuddle提交b1639f3

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架