LLMVerify

io.github.subodhkcv1.8.0更新於 Oct 9, 2026

Local-first MCP server for LLM output verification — risk signals, injection and PII checks.

已驗證STDIO僅桌面AI & MLSecurity & Monitoring

概覽

AI 產生的概覽

本機 MCP 伺服器,用來檢查 LLM 輸出的幻覺風險、提示注入與 PII,並可對敏感資料進行遮蔽。

功能
LLMVerify 透過 stdio 向相容 MCP 的代理程式提供一個本機驗證引擎。它提供六個工具:verify_llm_content、assess_hallucination_risk、check_prompt_injection、check_pii、redact_pii 與 get_llmverify_capabilities。檢查採用確定性的樣式比對,涵蓋幻覺與一致性訊號、注入與越獄樣式,以及電子郵件、電話號碼、SSN、信用卡與常見 API 金鑰等標準 PII 格式。結果會附上明確的 limitations 或 notChecked 清單。
適用情境
當助理需要處理不受信任的使用者輸入或模型輸出,並希望在內容送到使用者或日誌之前加一層本機防護時使用。適合初步篩檢注入嘗試、PII 偵測與遮蔽,不適合做事實查核或核准決策。
執行需求
透過 npx 從 npm 套件 llmverify 在本機執行;MCP 命令需要 Node.js 20 或更新版本。不需要帳號或 API 金鑰,免費版不會發出網路請求。選用環境變數可設定稽核、基準、日誌與狀態目錄,以及輸入輸出大小上限與單一工具逾時。
安裝前請注意
偵測採用樣式比對:幻覺訊號無法證明某個說法為假,經過混淆或編碼的 PII 以及新型注入可能被漏掉。遮蔽與驗證結果不應取代人工審查。選用密鑰 LLMVERIFY_AUDIT_HASH_KEY 可在稽核記錄中啟用帶金鑰的內容雜湊;設定 LLMVERIFY_AUDIT_NO_CONTENT_HASH 可停用內容雜湊。另有本機 HTTP 伺服器沒有身分驗證,將其暴露到 localhost 之外有風險。

安裝

在 SourceWeft 中

  1. 開啟 儀表板中的 LLMVerify,將其新增到工作區。
  2. 為需要使用其工具的對話啟用該服務。

Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。

其他 MCP 客戶端

參照 儲存庫 中的啟動說明。

README

llmverify

You shipped an AI feature. Your LLM hallucinated a citation, leaked a customer's email, and followed a prompt-injection buried in user input — on the same day. llmverify is the safety layer that sits between your LLM and your users.

Local-first verification, PII redaction, prompt-injection defense, and runtime monitoring for any LLM. One npm install. Zero telemetry. No API keys on the free tier.

[npm version] [CI] [License: MIT]

Last Updated: August 21, 2026 Version: 1.7.0 Node: >= 18.0.0 License: MIT


Links


The problem

You build with GPT-4, Claude, Gemini, or any LLM. The model:

  • Hallucinates facts and citations that do not exist.
  • Leaks PII — emails, phone numbers, SSNs, API keys in responses.
  • Follows prompt injections — users trick it into ignoring your instructions.
  • Returns broken JSON that crashes your parser.
  • Drifts in quality over time, and nobody notices until a user complains.

You need a guardrail between the model and your users. That is llmverify.


Install

bash
npm install llmverify

Everything runs locally. The free tier makes zero network requests and needs no API key. Free tier limit: 500 verification calls per day (tracked locally, never sent anywhere).


What you get

FunctionOne-linerWhat it does
verify(content)await verify(aiResponse)Runs hallucination, consistency, safety, and CSM6 checks; returns a risk level and findings
isInputSafe(input)isInputSafe(userMessage)Blocks prompt injection, jailbreaks, and malicious input before it reaches the model
redactPII(text)redactPII(aiResponse)Masks emails, phones, SSNs, credit cards, and API keys
containsPII(text)containsPII(text)Returns true if PII is present
detectAndRepairJson(...)detectAndRepairJson(prompt, response)Detects and repairs broken JSON output
monitorLLM(client)monitorLLM(openaiClient)Wraps any LLM client; tracks latency, token drift, and behavioral changes
sentinel.quick(...)await sentinel.quick(client, model)Runs regression tests against your model before users see changes
classify(...)classify(prompt, response)Intent detection, hallucination signals, and instruction compliance
auditLog(event)auditLog({ ... })Appends a local, hash-only audit entry for SOC 2 / HIPAA / GDPR evidence
run, prodVerify, ciVerifyawait prodVerify(content)Preset pipelines for dev, prod, strict, fast, and CI use

Quick start (30 seconds)

javascript
const { verify, isInputSafe, redactPII } = require('llmverify');
// 1. Block prompt injection before it reaches the model.if (!isInputSafe(userMessage)) {  return { error: 'Invalid input detected' };}
// 2. Verify the model's output.const aiResponse = await yourLLM.generate(userMessage);const result = await verify(aiResponse);
if (result.risk.level === 'critical') {  return { error: 'Response failed safety check' };}
// 3. Strip PII before the response reaches a user or a log.const { redacted } = redactPII(aiResponse);console.log(redacted);

Three lines of safety between your LLM and your users. No config file required. No API key required.


How it works

llmverify runs deterministic, pattern-based engines locally — no model calls, no network on the free tier. Same input plus same rules equals same result. Every result carries an explicit limitations array stating what was and was not checked, so you never mistake a clean score for a guarantee.

Framework alignment (baseline mapping only — not certification):

  • OWASP LLM Top 10
  • NIST AI RMF
  • EU AI Act
  • ISO 42001
  • CSM6 (HAIEC's 38-rule control set)

CLI

bash
# Verify a string from the terminal.npx llmverify verify "The capital of France is London."
# Start a local HTTP API for IDE / tool integration (localhost only by default).npx llmverify-serve --port=9009
# Expose to the network only on a trusted network. There is no auth on the API.npx llmverify-serve --host=0.0.0.0 --port=9009

The server binds to 127.0.0.1 by default, restricts CORS to localhost origins, and rate-limits clients (100 requests / 60s). It requires express (an optional dependency that installs by default).


MCP server

llmverify ships a built-in Model Context Protocol server — the same engine, exposed to MCP-compatible agents and IDEs over stdio:

bash
npx llmverify mcp
jsonc
// MCP client config{  "mcpServers": {    "llmverify": {      "command": "npx",      "args": ["-y", "llmverify", "mcp"]    }  }}

Six tools: verify_llm_content, assess_hallucination_risk, check_prompt_injection, check_pii, redact_pii, get_llmverify_capabilities. Stdio-only, zero outbound network, bounded inputs/outputs, PII-filtered responses, honest notChecked/audit semantics.

Requires Node.js ≥ 20 (the MCP SDK's floor; the rest of the package supports ≥ 18). The MCP SDK and zod are regular dependencies — the mcp command lazy-loads them so other commands pay no startup cost.

See docs/MCP.md for the full tool reference and security model.


Limitations

llmverify is a triage tool, not a truth oracle. Be honest with yourself about what it can and cannot do:

  • It cannot definitively prove hallucinations. Hallucination signals are pattern-based. "The capital of France is London" scores low because the text looks internally consistent. Ground-truth verification requires a source document you provide.
  • It does not replace human review. Use it to triage, not to approve.
  • PII detection is regex-based. It catches standard formats (emails, US phones, SSNs, credit cards, common API keys). It misses obfuscated, image-embedded, or encoded PII. Accuracy is roughly 90% for standard formats, lower for variations.
  • Prompt-injection detection is pattern-based. Novel or obfuscated injections can evade it.
  • Free tier is 100% local. ML-enhanced features require a paid tier and an explicit API key; the free tier never makes network requests and never sends data anywhere.

If a claim matters, verify it yourself. llmverify narrows the risk surface; it does not eliminate it.


Documentation


Part of HAIEC

llmverify is part of the HAIEC (Human AI Evidence Company) AI governance platform. Use it alongside the AI Security Scanner, the CI/CD pipeline integration, and Runtime Injection Testing.


Support


License

MIT — see LICENSE.


Recommendation (not legal advice): Run verify() on every model output that reaches a user, and isInputSafe() on every user input that reaches a model. Treat the risk level as a triage signal, not an approval.

來源:README.md,提交 2d130f6

工具

0
工具後設資料尚未被收錄。

版本歷史

1
  1. v1.8.0最新Oct 9, 2026