
Keyspoor
io.github.majiayu000v0.1.3更新於 Oct 4, 2026
Offline secret scanning for AI agents. Redacted findings; scans text and files in a chosen root.
概覽
供 AI 代理使用的離線密鑰掃描:掃描文字與指定根目錄下的檔案,回傳已去識別化的發現結果。
- 功能
- Keyspoor 是唯讀的 MCP 伺服器,提供 scan_text 與 scan_paths 工具,可在設定的根目錄下偵測文字與檔案中的憑證。它使用內建的 225 條規則,回傳包含規則、路徑、位置、信賴度與指紋的去識別化報告,不含原始密鑰或原始碼片段。它也提供 CLI 與 Rust 函式庫,可掃描儲存庫、暫存的 Git 變更、本機 Git 歷史與壓縮檔。
- 適用情境
- 當編碼代理需要在提交或分享前檢查程式碼、設定或貼上的文字是否外洩密鑰時使用。適合必須保持發現結果去識別化的離線本機工作流程。若要在變更被接受前強制執行,應搭配必要的 CI 檢查或 Git 鉤子,因為代理可能不會主動呼叫這些工具。
- 執行需求
- 透過 stdio 在本機執行,通常使用 Node.js 20+ 的 npx,或使用 Rust 執行檔。未宣告帳號、API 金鑰或環境變數。掃描本身離線進行,僅首次 npm/npx 安裝需要存取套件來源。必須設定根目錄,檔案路徑必須解析到該根目錄之下。
安裝
在 SourceWeft 中
- 開啟 儀表板中的 Keyspoor,將其新增到工作區。
- 為需要使用其工具的對話啟用該服務。
Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。
其他 MCP 客戶端
參照 儲存庫 中的啟動說明。
README
Keyspoor
Offline secret scanning for Rust applications, CI and AI coding agents. Keyspoor scans files, staged Git changes, local Git history and ZIP/tar/gzip archives, with redacted JSON, JSONL and SARIF results. Use it as a native CLI, a reusable Rust library or a read-only MCP server.
Website · Rust API · crates.io · npm · Releases · 简体中文
Try it in 60 seconds
With Node.js 20+, paste this into a macOS/Linux terminal. The value below is made up for this demo and is not a credential:
Expected: exit=1 and one finding. The relevant report fields are:
This is an excerpt; the actual report also includes ranges, fingerprints,
statistics and scan context. It contains neither the value nor a source snippet.
On Windows PowerShell, pipe the same synthetic string to
npx.cmd -y [email protected] scan - --format json, then check $LASTEXITCODE.
Install
Current release: 0.1.3. The Rust API is pre-1.0 and may change. The npm
package bundles native binaries for macOS (Apple Silicon/Intel), Linux GNU
(ARM64/x64) and Windows x64; it is a CLI launcher, not a JavaScript scanning SDK.
There are no install hooks or runtime binary downloads. First-time npm/npx
installation requires registry access; scanning itself is offline.
Standalone binaries are also
available. To build from source, run cargo build --release --locked and use
target/release/keyspoor.
Why Keyspoor
- Offline by design: scanning does not call credential providers, and reports omit raw secrets and source snippets.
- Reusable Rust engine: compile rules once and share an
Engineacross calls and threads; no scanner SDK dependency. - Agent interfaces: MCP
scan_text/scan_paths, persistent JSONL requests, progress and cancellation, plus explicit incomplete-scan reporting. - Repository-aware inputs: staged index contents, local Git history, bounded archive expansion, ignore files and fingerprint baselines.
- Evidence you can inspect: versioned rules, documented exclusions and reproducible quality, throughput and allocation measurements.
The matching, keyword dispatch, filtering, decoding, location mapping, Git acquisition and agent interfaces are implemented here. This project does not depend on Kingfisher or another secret scanner SDK. It uses general-purpose Rust regex, Aho–Corasick, hashing and archive libraries.
The built-in catalog contains 221 rules adapted from the MIT-licensed Gitleaks v8.30.1 catalog plus four independently written rules (three generic assignment rules and one URI-password rule), for 225 rules in the default engine. See rule provenance and exclusions and third-party notices. This is not a Gitleaks-compatible engine: rule semantics that cannot be preserved are explicitly excluded. The implementation decision records the build/adapt boundary and rejected SDK and wrapper alternatives.
CLI and repository scans
Exit codes: 0 means the selected scan completed with no reported findings; 1 means it completed with findings; 2 means an error or incomplete scan. You can check both other outcomes with synthetic input:
When a finding appears, inspect its path and position locally, remove exposed
credentials from tracked files and rotate any real credential that was exposed.
Review known intentional findings before recording a baseline. Investigate every
2 result before treating the scan as complete.
An ignored finding is not evidence that the original input contained no secret.
Baselines suppress existing fingerprints but preserve changed secret values.
An incomplete scan cannot replace a baseline. Baseline schema 2 binds the scan
mode, canonical roots, effective ignore policy (including Git/Jujutsu repository
boundaries), size budget and engine settings
(including rule content and fingerprint key identity). A different scope or
policy cannot compare against or overwrite that baseline. Create a separate
baseline when intentionally changing policy. Schema 1 baselines are rejected.
Canonical filesystem identity makes aliases such as ./ stable; reported paths
and rule path predicates remain relative to the selected source root.
Every finding is redacted and contains a rule, path, byte range, one-based line,
zero-based byte column, confidence, explanation and fingerprint. No raw secret
or source snippet is serialized. The default unkeyed fingerprint is a digest,
not encryption and not protection against guessing low-entropy secrets. Supply
--fingerprint-key-file with exactly 32 private bytes for keyed BLAKE3 identities.
Keep the same private key when comparing baselines.
Exact matches at the same secret byte span within one decoded view are merged.
rule_id identifies the primary rule: non-generic- IDs take precedence, then
higher confidence, then lexical rule ID. Merged findings include sorted
matched_rule_ids containing the primary and other matching rules; single-rule
findings omit that field. Adjacent occurrences, partial overlaps and distinct
positions inside one Base64 container stay separate. The primary rule determines
the fingerprint. This semantic update changes the engine configuration identity,
so previous baselines must be explicitly recreated rather than compared silently.
Scanning never contacts credential providers or validates whether a credential
is active. Filesystem scans respect ignore files by default; --no-ignore
includes ignored files, while Git metadata remains excluded. Oversized inputs
are reported as incomplete, not clean. Default file and archive expansion limits
are 64 MiB; --max-bytes changes the budget. Archive traversal never writes
members to disk. ZIP, tar and gzip have a shared expansion budget and at most
four nesting levels. Encrypted or damaged members remain visible as errors.
Staged mode reads the index version of each changed file, not the working
tree. It scans whole staged files to preserve credential-pair context. History
mode scans locally reachable commits, reads each unique blob once and retains
commit/path occurrences. It does not fetch LFS, submodules or remote refs, and
also unpacks supported archives in Git blobs, retaining commit/member paths. stats.files/bytes count
unique inputs; detection_passes counts actual path-sensitive engine calls.
GitHub Action
Copy the complete pull request workflow
into .github/workflows/keyspoor.yml, or follow the CI setup guide
for reports, permissions, baselines and failure handling. Its scan job can be
required in branch protection. The core steps are:
Place these steps in a job. The Action installs the exact npm package version
recorded at its source ref, scans the selected path and saves a redacted report
outside the checkout. report-path points to that file; exit-code preserves
0 (clean), 1 (findings), or 2 (error/incomplete). Findings and errors fail the
step; reports may be incomplete after an error. Use a full commit SHA to pin
the Action immutably instead of the maintained v1 ref. Supported report
formats are SARIF (default), JSON and JSONL. Installation needs npm registry
access; detection itself remains offline. The Action scans files; use the CLI
for staged/history modes and custom engine options.
Releases publish through GitHub Actions with npm and crates.io trusted publishing. Homebrew updates in the existing tap run hourly and may be delayed by GitHub scheduling. See the release guide for triggers and verification.
Library and custom rules
Add keyspoor = "0.1" and anyhow = "1" to your dependencies for this example.
Reuse an engine across calls and threads to avoid repeated compilation. It does not install a global runtime, logger or executor.
Finding.path and Finding.explanation are shared Arc<str> values. This is a
breaking Rust API change from String: use .into() when assigning an owned
string, .as_ref() to read &str, and replace the field to change its text.
JSON still contains ordinary strings; fingerprints, scan context and baseline
identity are unchanged. Generated findings share paths within one engine scan
and explanations across calls using the same engine. Deserialization does not
intern repeated strings.
Custom JSON rule files use
{"rules": [...]} and can be supplied with --rules path.json or a rules
directory. --no-builtin selects only custom rules. Each rule has id, name,
pattern, optional secret_group, keywords, min_entropy, confidence,
path, allowlist and exclude_paths. Keywords are case-insensitive OR-ed
necessary conditions; rules without keywords are always considered. Allowlist
groups have condition (or/and), target (secret/match/line),
regexes, paths, and stopwords. Groups are OR-ed; predicates within one
category are OR-ed, then populated categories are combined using condition.
Stopwords are case-insensitive substrings of the secret. Empty groups fail.
Omitted/null secret_group selects the first nonempty capture (or whole match);
explicit 0 selects the whole match. These rule-schema changes are breaking.
Unknown rule
fields and invalid regexes fail visibly. Rule patterns are never echoed in
compilation errors. Custom rule descriptions and identifiers are trusted
configuration and should not contain real secrets.
Generic API/access/auth token assignments suppress long explanatory prose (at least eight words plus sentence/clause punctuation). This heuristic does not suppress natural-language password/secret assignments, but a real token formatted as a long punctuated phrase can still be missed. Provider rules retain upstream limitations; see the rule provenance document.
Agent interfaces
Start with the tested configuration guide for Codex, Claude Code and Cursor. MCP exposes read-only tools; whether an agent calls them depends on the client, its permissions and the task. Use a required CI check or Git hook when scans must run before changes are accepted.
serve accepts one JSON object per line: {"id":1,"path":"file.txt","text":"..."}.
Responses contain the same id and a complete redacted scan report. Rules are
compiled once. Malformed JSON returns an error without echoing the payload;
an oversized request terminates this simple stream.
The MCP stdio server supports initialization, ping, tool listing and calls to
scan_text and scan_paths. File paths must resolve beneath the configured
root. The server does not perform writes or network verification. Root checks
are not an OS sandbox against another local process replacing paths concurrently.
MCP errors and tool failures are distinct, and malformed/oversized frames do not
turn into successful clean results. stdout contains protocol messages only.
One scan runs at a time while ping and cancellation remain responsive. Clients
may supply a progress token and cancel by request id; cancelled requests receive
no final response, and the session remains usable. Cancellation is checked at
file, Git blob and archive-member boundaries (also archive read chunks), not
inside a single regex operation. Tool results include at most 100 findings and
100 errors within a shared 512 KiB entry budget, plus total counts and
output_truncated; truncation is independent of scan completeness.
JSON, JSONL and SARIF are available. Ordinary filesystem/Git JSONL scans emit
finding, error and progress records during scanning, followed by a summary
with counts and completeness. Bounded worker queues avoid collecting all file
reports; output errors stop further acquisition. Each worker still buffers one
input and its findings, and Git history metadata remains resident. JSON/SARIF,
stdin, and JSONL scans using baselines collect reports; baseline validation must
finish before filtered results are emitted. SDK callers can use the scan_*_stream
APIs with ScanControl and a fallible event sink. Regular progress is emitted at
most once per 100 ms, with first and final snapshots retained. These are
event-driven updates, not a heartbeat during a single long regex operation.
The CLI flushes its first finding and each error immediately; later findings use
buffered writes, and progress/final records flush remaining output.
Verification and comparison
Benchmark methodology and reproduction separates synthetic labelled quality from unlabelled real-source throughput. Versions, commands, hardware, binary fingerprints, time, memory, parser completeness and unsupported capabilities are retained. Online verification is disabled. These measurements do not establish production accuracy or universal speed leadership.
The latest v6 allocation and performance report measured 20 alternating before/after pairs on Apple M2 Max/macOS. For one 100,000-finding synthetic JSONL workload, scanner peak RSS fell from 99.86 to 59.59 MiB and scanner CPU time fell 8.74%; regular throughput and Git workloads were broadly unchanged. These are changes within this project, not competitor speed rankings. Previously observed synthetic regression sets retained 600 true positives, 25 false positives and zero false negatives; this is not a real-world precision estimate.
See measured results, the
versioned cross-tool comparison,
comparison methodology and
feature-by-feature research. The research covers 22 external
projects; unsupported, unavailable and cloud-dependent entries are identified
rather than scored as failures. Historical artifacts call this project
secret-scan, its name before Keyspoor; their recorded commands and binary
hashes have been preserved.
Feature implementation matrix maps every item in the research catalog to implemented, partial or unimplemented status. Cloud connectors, credential validation/revocation, GPU/ML backends and language bindings are not implied by the presence of a CLI or MCP server.
License and support
Keyspoor is Apache-2.0 licensed. Adapted rule data retains its upstream MIT license and attribution in THIRD_PARTY_NOTICES. Report reproducible bugs or request features in GitHub Issues; use synthetic examples and do not include real credentials.
來源:README.md,提交 7eb9bf9
工具
0版本歷史
1- v0.1.3最新Oct 4, 2026

