MCP Security Guard

io.github.petrovicistefanv0.7.0更新於 Oct 6, 2026

Audit MCP servers for tool poisoning, rug pulls and supply-chain risk (OWASP MCP Top 10).

概覽

AI 產生的概覽

稽核已安裝的 MCP 伺服器,檢查工具投毒、隱藏文字、rug pull、高風險能力與供應鏈風險,並提供可選的執行階段掛鉤。

功能
探索 Claude Code、Claude Desktop、外掛及其他本機用戶端中設定的 MCP 伺服器,並稽核其工具定義、說明、提示與資源,檢查投毒、隱藏字元、工具遮蔽與名稱衝突。它會以雜湊釘住工具定義,以便發現後續變更,為每個伺服器評分並歸類執行、寫入、刪除、外送等能力。可選掛鉤會在 MCP 呼叫時提示憑證或注入指令,並記錄不含內容的稽核日誌。發現項目會標註 OWASP MCP Top 10 編號,可匯出為 markdown、JSON、SARIF 或 HTML。
適用情境
當你安裝了多個 MCP 伺服器,想了解它們能做什麼、描述是否遭竄改時使用。它適合在安裝前審查某個伺服器、釘住定義以便日後發現變更,以及在 CI 中執行已核准伺服器政策。
執行需求
以 npm 套件透過 stdio 在本機執行,通常使用 npx,免費功能不需帳號或 API 金鑰。進行工具層級稽核需要啟動伺服器並明確確認;可選的供應鏈檢查需要連線至 npm、PyPI 與 OSV。付費方案使用 MCP_SECURITY_API_KEY。
安裝前請注意
部分工具會啟動已設定的伺服器,其中 adversarial_test 會以注入酬載呼叫其工具,因此只應用於你自己擁有的伺服器。啟用 write 後 apply_fixes 會在備份後修改 .mcp.json 與 .claude/settings.json。供應鏈檢查會把套件名稱與版本傳送給第三方,可選的付費方案會透過 MCP_SECURITY_API_KEY 傳送資料。靜態檢查無法證明伺服器安全。

安裝

在 SourceWeft 中

  1. 開啟 儀表板中的 MCP Security Guard,將其新增到工作區。
  2. 為需要使用其工具的對話啟用該服務。

Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。

其他 MCP 客戶端

參照 儲存庫 中的啟動說明。

README

mcp-security-guard

[mcp-security-guard logo]

A Claude Code plugin that audits the MCP servers you have installed. Your code is covered by other tools. This one checks the servers that inject text into Claude's context.

CheckWhat it catches
Tool poisoningInstruction overrides, "don't tell the user", <IMPORTANT> tags, directives to read secrets (~/.ssh, .env), conversation harvesting, exfiltration via URLs, parameters and Markdown images, HTML comments, encoded payloads
Full-schema poisoningThe same checks on parameter names, descriptions, defaults, enums, required, plus non-schema text in type
Hidden textZero-width, bidi-control and Unicode-tag characters, ANSI terminal escapes, homoglyph tool names
Tool shadowingA server whose descriptions reference another server's tools, and tool-name collisions between servers
Rug pullsTool definitions or launch commands that changed after you pinned them (SHA-256 per tool). Re-checked automatically at every session start
Capabilities & scorePer-tool classification (execute, delete, write, egress), unauthenticated remote write access, 0–100 score and A–F grade per server, recommended permission rules
Supply chainOSV vulnerabilities and malicious versions, typosquats, missing or brand-new packages, install scripts, publisher changes (opt-in network check)
RuntimeHooks on every MCP call: ask before credentials are sent, warn on injected instructions or credentials in outputs, content-free audit log
Policy.mcp-security.json approved/blocked servers and hosts, enforced in audits, CI and at session start
ConfigurationPlaintext secrets in env/headers/args/URLs, plain-HTTP remotes, unpinned npx/uvx packages, privileged or unpinned Docker images, pipe-to-shell launches, duplicate names across scopes

It discovers servers from every place Claude Code and Claude Desktop load them: user, local and project scope, servers shipped inside installed plugins and plugins synced from your claude.ai account (named <plugin>:<server>), claude_desktop_config.json and Claude Desktop extensions, the organisation-managed managed-mcp.json, other clients on the machine (Cursor, VS Code, Windsurf, user and project configs), and the claude.ai connectors you have used (names only: their configuration lives in your account).

It scans everything a server puts into Claude's context, not only tools: server instructions, prompts, resources and resource templates go through the same poisoning checks and are pinned for rug-pull detection.

Everything runs locally. Nothing is sent anywhere.

OWASP MCP Top 10 coverage

Every finding is tagged with its OWASP MCP Top 10 id, in reports and in SARIF.

IDRiskCovered by
MCP01Token Mismanagement & Secret Exposure✅ Plaintext secrets in env, headers, args and URLs. At runtime, asks before a credential is sent to an MCP server and warns when one comes back.
MCP02Privilege Escalation via Scope Creep✅ Capability inventory (execute, delete, write, egress), ready-to-paste permissions.ask rules, privileged or broadly mounted containers
MCP03Tool Poisoning✅ 26/27 published techniques detected, including full-schema poisoning, shadowing and name collisions, plus rug-pull pinning
MCP04Supply Chain Attacks✅ Unpinned packages, images and git sources; OSV vulnerabilities and malicious versions; typosquats; new packages; install scripts; publisher changes
MCP05Command Injection & Execution✅ Flags tools that can execute commands; adversarial_test finds injectable parameters in servers you own
MCP06Prompt Injection via Contextual Payloads✅ In tool metadata and, at runtime, in tool outputs (English patterns)
MCP07Insufficient AuthN/AuthZ✅ Plain-HTTP remotes; remote servers exposing write or exec tools without authentication; OAuth detection
MCP08Lack of Audit and Telemetry✅ Local, content-free audit log of every MCP call, with query_audit_log
MCP09Shadow MCP Servers✅ Discovery across user, project, local, plugin and Claude Desktop configs; approved-server policy enforced in audits, CI and at session start
MCP10Context Injection & Over-Sharing✅ Conversation and system-prompt harvesting, Markdown-image exfiltration, credentials in outputs, network-egress inventory

What a local tool cannot do (planned for a hosted Team plan): org-wide discovery and audit aggregation, and OAuth scope review.

Measured: detects 26/27 attacks from a corpus of publicly documented techniques, with 0 false positives on 13 hard benign samples and on 17 real servers (83 tools). See bench/RESULTS.md.

Install

/plugin marketplace add petrovicistefan/mcp-security-guard/plugin install mcp-security-guard@mcp-security-guard

Then run /mcp-audit, or ask Claude "are my MCP servers safe?".

Other MCP clients (Cursor, VS Code, Windsurf, Claude Desktop, Cline, ...)

The same server is published on npm and runs with no install step:

json
{ "mcpServers": { "mcp-security-guard": { "command": "npx", "args": ["-y", "mcp-security-guard"] } } }

The command line tool is the same package: npx mcp-security-guard audit --project-only. It is also listed in the official MCP Registry as io.github.petrovicistefan/mcp-security-guard.

Tools

ToolLaunches servers?
list_mcp_serversNo
audit_mcp_configNo
audit_server_toolsYes, after explicit confirm_launch: true. Sends only initialize and list requests (tools, prompts, resources); never calls a tool, renders a prompt or reads a resource
pin_toolsYes (same as above). Writes ~/.claude/mcp-security/pins.json
analyze_tool_definitionsNo. Offline analysis of a tools/list payload, for MCP server authors
check_supply_chainNo. Sends package names and versions to npm, PyPI and OSV after confirm_network: true
apply_fixesOnly for the permissions fix. Dry run by default; with write: true it edits the project's .mcp.json / .claude/settings.json after a backup to ~/.claude/mcp-security/backups/
security_dashboardOnly with scan: "full" and confirm_launch: true. Interactive dashboard (MCP App)
generate_policyNo. Returns a .mcp-security.json approving the current servers
query_audit_logNo. Summarises the runtime audit log
adversarial_testYes, and calls tools with injection payloads. Only for servers you own; needs i_own_this_server and confirm_launch; skips destructive tools

Session-start check

A SessionStart hook re-verifies only the servers you have pinned (pinning is your consent to launch them) and stays silent unless something changed. Control it with MCP_SECURITY_SESSION_CHECK:

  • full (default): compare launch configs and re-list tools
  • config: compare launch configs only, launch nothing
  • off: disable the check

CI / GitHub Action

Fail pull requests that add risky MCP servers to .mcp.json, and show the findings in GitHub code scanning:

yaml
name: MCP securityon: [pull_request]permissions:  contents: read  security-events: writejobs:  audit:    runs-on: ubuntu-latest    steps:      - uses: actions/checkout@v4      - uses: petrovicistefan/mcp-security-guard@main        id: mcp        with:          fail-on: high          # critical | high | medium | low | info | none      - uses: github/codeql-action/upload-sarif@v3        if: always()        with:          sarif_file: ${{ steps.mcp.outputs.sarif-file }}

The same checks run locally without Claude:

node plugin/dist/cli.mjs audit --project-only --format sarif --output mcp.sarifnode plugin/dist/cli.mjs analyze-tools tools.json --name my-server   # for MCP server authors: a saved tools/list result

Check a server before installing it (launches it, sends only initialize and tools/list):

node plugin/dist/cli.mjs scan some-server.mcp.json --confirm-launch

Test your own server for command injection and path traversal (it calls the tools; run a test instance, ideally in a container):

node plugin/dist/cli.mjs adversarial my-server.mcp.json --server my-server --i-own-this-server --confirm-launch

Fix what the audit found (dry run first, then --write):

node plugin/dist/cli.mjs fix --pin-versions            # npx pkg → [email protected], uvx pkg → pkg==x.y.z (looks up npm/PyPI)node plugin/dist/cli.mjs fix --env-refs --write        # literal secrets in .mcp.json → ${VAR} referencesnode plugin/dist/cli.mjs fix --permissions --confirm-launch --write   # permissions.ask rules for risky tools

Start a team policy from the servers configured today:

node plugin/dist/cli.mjs policy-init && git add .mcp-security.json

Any command takes --format markdown|json|sarif|html. The HTML report is a single self-contained file you can open in a browser or attach to a ticket.

Exit codes: 0 clean, 1 findings at or above --fail-on, 2 usage error.

Limitations

Remote servers that require OAuth (most hosted MCP servers) cannot be scanned at the tool level: the scanner cannot reuse Claude Code's tokens. Their configuration is still audited.

Interactive dashboard (MCP App)

security_dashboard is an MCP App: hosts that support MCP Apps (Claude Desktop, claude.ai, VS Code Copilot…) render it inline. It shows every server with its score and grade, findings filterable by severity and server, the OWASP MCP Top 10 breakdown, recommended permission rules, and Full scan and Pin buttons. Selecting a server tells Claude what you are looking at, so follow-up questions have context. Claude Code in a terminal gets the text summary instead.

To use it in Claude Desktop, add the server to claude_desktop_config.json and ask Claude to "open the MCP security dashboard":

json
{ "mcpServers": { "mcp-security-guard": { "command": "node", "args": ["/path/to/mcp-security-guard/plugin/dist/index.mjs"] } } }

The UI is a single self-contained HTML file. Server-supplied text reaches the page only as text (never as HTML), and the host's sandbox applies. Develop it with a local host that drives the real server: npm run dashboard:dev -- /path/to/project.

Runtime hooks

HookWhat it doesSetting
PreToolUse on mcp__*Asks for confirmation when a call's arguments contain a credentialMCP_SECURITY_SECRET_GUARD=ask (default), deny or off
PostToolUse on mcp__*Warns Claude and you when an output contains injected instructions, hidden characters, exfiltration markup or a credentialalways on
Audit log~/.claude/mcp-security/audit.jsonl: server, tool, time, input hash and sizes. Never arguments or outputs. Rotates at 10 MB.MCP_SECURITY_AUDIT_LOG=off

The hooks add about 40 ms per MCP call.

Trust model

  • Read-only, apart from the pin file, the audit log, and policy-init (which writes a file you asked for).
  • Network only when you opt in: check_supply_chain / --supply-chain send package names and versions to npm, PyPI and OSV. The adversarial_test tool is the only one that calls tools.
  • Evidence from scanned servers is sanitised (invisible characters revealed, length capped) and labelled as untrusted data.
  • Secrets are masked in all output.
  • Static checks reduce risk. They do not prove a server safe: malicious behaviour in tool responses or server code is out of scope.

Development

npm installnpm run build      # bundles to plugin/dist/ (committed, so the plugin runs without npm install)npm testnpm run bench      # false-positive gate on real servers (network; Docker images optional, see bench/RESULTS.md)

test/fixtures/poisoned-server.mjs is a deliberately malicious server used by the end-to-end test.

Pro & Team (early access)

Everything above is free and stays free: it runs locally and needs no account. Paid plans add what needs a server: a daily threat feed of known malicious MCP servers and packages, alerts when a server you use ships changed tool descriptions, history, and team policies and dashboards. They are opt-in through MCP_SECURITY_API_KEY; see PRIVACY.md for exactly what is sent.

Interested? Join the early access list. Early sign-ups get launch pricing, including a limited lifetime license.

Security, privacy, license

About the author

I'm Stefan Petrovici: passionate about IT, a husband and a father. I built mcp-security-guard on my own. I'm looking for a job.

I build web applications end to end, frontend, backend, APIs and deployment, and I'm happy to work on anything else that needs building. This repository shows how I work: tests, CI, careful documentation and attention to security.

I also build WordPress and WooCommerce plugins, available at pluginsforstores.com.

If your team is hiring, write to me at [email protected].

來源:README.md,提交 dbc07be

工具

0
工具後設資料尚未被收錄。

版本歷史

1
  1. v0.7.0最新Oct 6, 2026