Defuddle

joeseesun/defuddle-skill/skills/defuddle

作者 joeseesunb1639f3627d776b20d93a8098d2c896d34bdc672无许可证112 个星标收录于 2026年10月9日更新于 2026年10月9日仓库7个月前更新

Extract clean article content from web pages or local HTML files. Removes clutter (ads, sidebars, nav) and returns readable content with metadata.

AI 生成的概览

从网页或本地 HTML 文件中提取干净的文章正文和元数据,并保存为 Markdown。

功能
Defuddle 从网页或本地 HTML 文件中提取主要文章内容,去除广告、侧边栏和导航。它返回干净的 Markdown 以及标题、作者、来源、发布日期和字数等元数据。随后该流程会把结果保存为带 frontmatter 的 Markdown 文件,并向用户确认文件路径。
适用场景
当用户想要提取或清理网页内容、从 URL 获取文章正文、去除 HTML 中的杂乱元素,或把网页转换为干净的 Markdown 时使用。它适合博客、新闻和文档等文章类页面。
运行要求
需要 Node.js 和 npm,并全局安装 defuddle CLI,同时以 jsdom 作为同级依赖。处理 URL 时需要网络访问。该技能不附带脚本,说明中只描述 CLI 命令。

Defuddle - Web Content Extraction

Extract main article content from web pages, removing ads, sidebars, navigation, and other clutter. Output clean Markdown with metadata.

Prerequisites

Before first use, check if defuddle is installed:

bash
command -v defuddle >/dev/null 2>&1 || npm install -g defuddle jsdom

Default Workflow

When user provides a URL, follow this workflow:

Step 1: Extract content as Markdown + JSON metadata

Always use both -m and -j flags to get markdown content with full metadata:

bash
defuddle parse "<url>" -m -j

Step 2: Present a summary to the user

Show the user:

  • Title: from JSON title field
  • Author: from JSON author field
  • Source: domain
  • Word count: from JSON wordCount field
  • A brief preview (first 2-3 sentences)

Step 3: Ask where to save

If this is the first time using defuddle in this conversation, ask the user:

"Save to which directory? (e.g. ~/Documents, ~/Desktop, or a custom path)"

Remember the user's chosen directory for subsequent uses in the same conversation.

Step 4: Save as Markdown file

Write the file with frontmatter + full content:

markdown
---title: {title}author: {author}source: {url}date: {published or "Unknown"}clipped: {today's date YYYY-MM-DD}wordCount: {wordCount}---
# {title}
{markdown content}

File naming: Use the article title as filename, sanitized for filesystem:

  • Replace special characters with spaces
  • Trim whitespace
  • Example: The Shape of the Essay Field.md

Step 5: Confirm to user

Tell the user the file path where it was saved.

CLI Reference

bash
defuddle parse <source> [options]

Arguments:

  • <source> — URL (https://...) or local HTML file path

Options:

FlagDescription
-m, --markdownConvert content to Markdown
-j, --jsonOutput as JSON with full metadata
-o, --output <file>Write to file instead of stdout
-p, --property <name>Extract single property (title, description, domain, author, published, wordCount, content)
--debugVerbose logging

JSON Response Fields

When using -j, the response includes:

  • title — Article title
  • author — Author name
  • published — Publication date
  • description — Meta description
  • content — Extracted Markdown (when -m used)
  • domain — Source domain
  • favicon — Favicon URL
  • image — Featured image URL
  • site — Site name
  • wordCount — Word count
  • parseTime — Processing time in ms

Notes

  • Requires Node.js and npm
  • jsdom is required as a peer dependency
  • Works best with article-style pages (blogs, news, documentation)
  • Not designed for SPAs or JavaScript-heavy pages (e.g. WeChat articles need browser rendering)

来源与署名

来源:joeseesun/defuddle-skill位于skills/defuddle提交b1639f3

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架