竞品内容 Pattern 雷达 · AI 答案里的重复说法

cn.savantcatv1.0.1Updated Oct 6, 2026

把 AI 答案/竞品内容当语料,找出反复出现的说法、被反复引用的信源与答案同质化程度,含对照组自检

VerifiedSTDIODesktop onlyWeb Search & ScrapingData & Analytics

Overview

AI-generated overview

Analyzes AI answers or competitor content as a corpus to surface repeated phrases, recurring cited sources, and answer homogenization.

What it does
Treats AI answers or competitor content as a corpus and finds phrases that recur across many answers, with coverage rates and a null-hypothesis threshold. It also reports which brands or entities occupy the answer space, which domains the AI repeatedly cites, and how similar answers to the same question are. Three tools are exposed: analyze_patterns for the main analysis, self_check to verify the measurement baseline, and explain_method for definitions and limits.
When to use it
Useful when you want to see what default wording an AI has settled on for a topic, who gets mentioned in those answers, and which sources are being cited. Suited to content, brand, or GEO/AI-visibility work on Chinese short-answer corpora of at least twelve samples.
Requirements
Runs as a local stdio process with Python and no third-party dependencies, or can be reached through the provider's hosted remote endpoint. Input is a JSON file of records with question and answer fields, or a plain array of such records. No accounts, API keys, or environment variables are declared.
Before you install
Read-only: it does not modify input files. It performs character-level n-gram analysis only, does not merge synonyms, does not judge sentiment or content quality, and does not predict traffic or ROI. It refuses to score corpora with fewer than twelve samples, and thresholds need recalibration for English or long-form text. Cited URLs are extracted and counted but not checked for reachability.

Installation

In SourceWeft

  1. Open 竞品内容 Pattern 雷达 · AI 答案里的重复说法 in the dashboard and add it to a workspace.
  2. Enable the server for the chats that should use its tools.

Desktop only via STDIO. STDIO servers start a local process, so they need the SourceWeft desktop host.

Other MCP clients

Follow the launch instructions in the repository.

README

竞品内容 Pattern 雷达 · pattern-radar

把「AI 答案 / 竞品内容」当语料,找出反复出现的说法——谁被反复提到、AI 反复引用哪些信源、答案空间是否已经固化。

纯标准库、零依赖、只读。clone 下来就能跑。

已托管为公网 MCP 服务,不用装任何东西就能用:https://savantcat.cn/mcp-radar

[License] [Deps] [CI]

[M8ven Verified]


python radar.py --selftest                        # 必跑:对照组自检python radar.py --input <记录.json> --brands 我家,竞品A,竞品B --out 报告.md

它回答的问题

  1. AI 对这个问题已经形成什么固定说法?(重复 Pattern + 覆盖率)
  2. 谁在这类答案里占位?(主体覆盖率)
  3. AI 实际上把谁当信源?(被反复引用的域名)
  4. 答案空间固化到什么程度?(同题相似度)

口径(重要,别误用)

  • Pattern = 覆盖度:出现在越多回答里,越像「AI 对这个问题的默认说法」。覆盖率 = 覆盖回答数 / 总回答数。
  • 阈值不是拍脑袋的:先取语料的高覆盖 n-gram 作种子,向左右逐字扩展成完整短语(所以给出的是「百度智能云」而不是碎片「度智能云」)。
  • 空阈值用零假设模拟得出:把每篇答案的字符内部打乱(保留长度与字频、摧毁短语),重算最大覆盖,取 p95;真实阈值 = max(空阈值+1, 3)。这是自校准的,不是经验常数。
  • 占位符域名会被剔除:答案正文里的「你的域名」这类模板文字不是信源,报告中标 ⚠️ 且不进结论。

对照组标定结果(v1.0.0 实测)

--selftest 造两份答案已知的语料:

对照组期望实测
A 健康组(注入 10/12 条重复短语 + 同一信源 8 次)检出✅ 检出,覆盖 100%;信源识别正确
B 病态组(每条只用互不重叠的字,真无重复)零误报✅ 检出 0 个
负对照(A 语料字符打乱)恒为 0✅ 0 个(链路不造伪 Pattern)
样本闸门(5 条样本)必须拒✅ 拒答并说明原因

对照组本身也被验证:自检会先断言「B 真的没有跨答案重复」,不满足就报语料构造错误而不是改工具。第一版自检就是这么抓出我的语料写错了(12 条答案共用同一模板串)。

边界(明说,不含糊)

  • 只做字符级 n-gram,不做语义聚类。同义不同形(「AI 客服」vs「智能客服」)不会合并。
  • 不判情绪、不判内容好坏、不预测流量或 ROI。覆盖率高 ≠ 评价好。
  • 样本 < 12 条直接拒绝出分,不硬凑。低于这个量算出来的是噪声,比没有更误导。
  • 换语料形态(英文、长文)需重新标定:当前阈值是在中文短答案(每篇数百字)上标定的。
  • 只读:不修改任何输入文件(可用 md5 前后比对自证)。
  • 引用 URL 只做提取与域名统计,不做可达性验证(打不开可能是反爬,不等于幻觉)。

输入格式

两种都吃:

  • geo-monitor 的记录:{"samples": [{"question","platform","run","answer"}, ...]}
  • 纯数组:[{"question","answer","platform"?,"run"?}, ...]

空答案会被丢弃并计数,不当成数据。

报告骨架

一、AI 答案里的重复 Pattern(含零假设阈值)二、竞品占位(谁被提到 / 覆盖率 / 是否出现在靠前位置)三、AI 反复引用的信源(信源池占位,占位符已剔除)四、答案空间同质化程度(同题两两相似度)五、最该先修的一件事(只给一条)附录:口径与边界

第五节刻意只给一条建议——需要的是下一步动作,不是十条清单。

怎么接入

用远程端点(推荐,零安装) —— 任何支持 MCP 的客户端填 URL 即可:

json
{"mcpServers": {"pattern-radar": {"url": "https://savantcat.cn/mcp-radar"}}}

三个工具:analyze_patterns(主分析)· self_check(先验尺子可用)· explain_method(口径与边界)。

本地 stdio(要改代码或离线用):

python server.py                    # stdiopython server.py --http --port 8768 # streamable-http

相关

License

MIT — 见 LICENSE。

作者

合尘猫 · 一个人 + 一个 AI 分身的 AI 落地实践 —— 知识库 · AI 客服合规 · 内容自动化。

版本

  • v1.0.0(2026-10-04)首版。阈值口径:零假设 p95 自校准 + 逐字扩展短语 + 碎片剔除 + 占位符识别。

Source: README.md at commit da037db

Tools

0
Tool metadata has not been indexed yet.

Version history

1
  1. v1.0.1LatestOct 6, 2026