
Ghostchars
io.github.ghostcharsv1.0.2更新于 Sep 29, 2026
Find and remove invisible Unicode: zero-width, tag smuggling, bidi, homoglyphs. Offline.
安装
在 SourceWeft 中
- 打开 控制台中的 Ghostchars,将其添加到工作区。
- 为需要使用其工具的对话启用该服务。
Desktop only,通过 STDIO。 STDIO 服务会启动本地进程,因此需要 SourceWeft 桌面宿主。
其他 MCP 客户端
参照 仓库 中的启动说明。
README
ghostchars
Find and remove invisible Unicode characters and machine-writing punctuation, in text and in documents. Everything runs on your machine: no network, no telemetry, no account, and zero runtime dependencies.
Needs Node 22 or newer. The whole package is a launcher and three bundles: the CLI, the document reader, and the MCP server. Only the first is loaded unless you name a document or start the server.
Commands
clean reads stdin or the files you name, writes the cleaned text to stdout, or
rewrites each file in place with -i. With --check it writes nothing and exits
1 if anything would change.
inspect lists every suspicious character with its line, column, codepoint,
Unicode name and the flag that would act on it. It changes nothing and exits 1
when it finds something.
report measures the habits that make a draft read as machine written: sentence
length uniformity, stock transitions, tell words, dash density, tricolons,
hedging, missing specifics. It is not a detector and emits no probability.
check holds a whole tree to a bar and prints every file that would change,
with exact character positions. This is the command for CI and for a pre-commit
hook. Configure surfaces and pins in ghostchars.json.
mcp runs a Model Context Protocol server on stdin and stdout, so an agent can
call clean_text, inspect_text, style_report and check_paths as tools
instead of shelling out.
help prints the commands, help <command> prints one command's flags, and
help --json prints the whole surface as one machine-readable document.
version prints the version.
Flags
The thirteen cleaning switches, their spellings and their meanings are exactly the Python reference CLI's, because it is the same engine underneath.
Default cleaning is context aware. It strips a zero-width joiner sitting between
two Latin letters and keeps the one inside a family emoji or a Persian word. It
strips free tag characters, which is how text is smuggled into a message, and
keeps a complete subdivision flag. It strips a left-to-right mark in English text
and keeps one next to Hebrew. -a turns that judgement off.
An option can only be turned on. There is no way to turn one off, because the reference CLI has none either.
The three bars
A bar is a named set of options, so a person or an agent picks one word instead of thirteen booleans.
Resolution order is fixed, so the result is always predictable:
- start from every option off,
- apply
--bar, - apply
--maxif given, - apply each explicit flag, which can only turn an option on.
So --bar text --collapse is -s --punct --collapse, and --bar max --punct is
still just max.
max folds look-alike Cyrillic and Greek letters and applies NFKC, which
rewrites letters. Do not run it over text you did not write. It is not offered as
a gate bar at all: see check below.
Documents
clean and inspect recognise a document by its extension and handle it
structurally.
--as text forces plain-text treatment of any path, which is what the Python
reference CLI does with an .html file.
Without -i, a cleaned document is written to <name>.clean<ext> beside the
original, or to the path you give with -o.
inspect on a document prints the character report over the concatenated
document text, then what hides around that text: hidden runs, white-on-white and
tiny runs, tracked changes, comments, the count of distinct rsid editing-session
ids, metadata fields, macros, media. --metadata blanks the authoring fields and
drops custom properties; --rsids removes the editing-session fingerprint.
Hidden runs, comments and tracked changes are reported and never removed. Deleting them changes the document, and that stays a human call.
Out of scope here: PDF, images, .pptx, .xlsx, .rtf, .doc, .epub and
.pages. check does not open documents at all.
check
--bar accepts exactly default and text. There is deliberately no max gate
bar: max turns on confusables and NFKC, which target scripts and symbols rather
than typography, and a gate that condemns legitimate Cyrillic, Greek and emoji
content gets deleted by the people who installed it. The default bar is
default, for the same reason: typography is a house style and has to be opted
into.
With no path arguments, check finds files with git ls-files, so it honours
the same ignore rules your repository already has and still sees files that are
new and not yet added. Outside a repository, or with --no-git, it walks the
tree instead and skips node_modules, build output, virtual environments and
every directory whose name starts with a dot, except the ones coding agents read
instructions from (.cursor, .github, .claude, .agents, .windsurf,
.devin, .gemini, .codex, .clinerules, .continue, .kiro, .roo). --git makes a git that cannot be
used a hard error, whether or not you named paths, so CI cannot silently switch
strategy between one container and another. Path arguments are always walked
rather than looked up in git, so --git <paths> asserts that git was available
and reports "discovery": "walk". Results are sorted by path, so the output
never depends on directory order. A directory the walk cannot open is an error
on stderr and exit 2, never a silent omission.
A file is skipped, silently and without failing the run, when its extension is a
known binary one, when it is larger than --max-bytes (2 MiB by default), when
one of its first 8192 bytes is NUL, or when it is not valid UTF-8. A file you
name on the command line is different: you asserted it is text, so both content
checks stop applying to it, the strict decode decides, and a decode failure is an
error that exits 2.
check <dir> finds <dir>'s own ghostchars.json even when you run it from
somewhere else: with no config beside the working directory, the search moves up
from each path you named, stopping at that path's repository root. Two paths
governed by two different config files is an error, because one run has one root
and a pin is written relative to it. --verbose prints which config was used.
ghostchars.json
Put it at the root of the tree, or point at it with --config.
A surface names a bar (or an options list, one of the two, never both), the
globs it covers, and its pins. A file that matches no surface is not checked at
all. An exclusion wins over everything, including a path you named explicitly.
An unknown key is an error that lists the known keys, so a typo in pins is
never silently ignored.
The glob subset is ** for any run of path segments, * for any run inside one
segment, and ? for one character. No braces, no character classes.
decodeEscapes resolves \uXXXX and \u{XXXX} to the characters they render
before the file is checked. Turn it on for a source file whose strings a person
reads, and leave it off everywhere else: a table of codepoints is full of escapes
that match characters rather than ship them.
Pins
A pin says a file changes by exactly N characters on purpose. It is enforced in both directions.
That last row is the one people forget. It catches a pin that went stale when somebody quietly fixed the file.
Output
The two placeholders stand for the real characters. This file is held to the
text bar itself, so it names a codepoint where it would otherwise print one.
One path:line:col U+XXXX NAME line per offending character, and the source line
printed once per distinct line number. Exit 1. When everything passes, one line:
ghostchars: 3 files pass the text bar. --verbose also names the files that
passed.
In CI
That is the whole step. Never pass --fix in CI. --fix cleans the failing
files in place and then re-checks, which turns a red build green without anyone
reading what changed. It also skips pinned files and names them, because fixing
one would invalidate its own pin.
After a local --fix, read back any file that had a dash finding. The cleaner
turns an em dash between two words into a space, a hyphen and a space, which is
correct and reads badly. Those positions are exactly what rewriteSites in the
JSON output lists.
--sarif prints the same run as a SARIF 2.1.0 log, one result per offending
character with its line, column and source line, and one result per pin failure.
GitHub code scanning shows those inline on the pull request:
The || true is there because exit 1 means findings, and the upload step is
what turns them into annotations. Keep a second plain check step if the job
must also fail. --sarif and --json are two formats for one stdout, so the
command refuses both at once.
As a pre-commit hook
integrations/pre-commit/pre-commit-config.yaml is a paste block for
.pre-commit-config.yaml. It is a repo: local hook with language: node, so
the framework installs the pinned npm package into its own environment and no
repository is cloned. It runs over the staged text files and, like the CI step,
never passes --fix.
Exit codes
The priority is strict and accumulates across inputs: 3 beats 2 beats 1 beats 0. A decode failure in one file forces 2 even when another file had findings.
Code 3 prints one line and, with GHOSTCHARS_DEBUG=1 in the environment, a stack
trace. It exists so an agent can tell "I passed the wrong flag" from "the tool
broke", which are different next actions.
EPIPE on stdout is a silent exit 0, so ghostchars inspect big.md | head
behaves like every other command.
JSON output
--json on clean, inspect and check prints exactly one JSON object to
stdout, never a stream of lines. Four rules, all of them for a machine reader:
ok and exitCode are in the payload as well as in the process status; counts
come before lists, so a reader that truncates still has the verdict; a cap that
bites always sets findingsTruncated and findingsTotal; and keys are camelCase.
All three open with the same five identity fields, so one reader can tell what it is holding before it looks at anything else.
clean and inspect continue with bar, the resolved options list, a
summary that totals across inputs, and results, one entry per input. A
clean result carries changed, its own summary, a charsDelta in
codepoints, the whole cleaned text (replaced by written: true under -i),
and rewriteSites. An inspect result carries clean, total, a summary, a
byKind rollup, then findings, each with line, col, index, cp, u,
name, kind, flag, action and, for a substitution only, replacement.
rewriteSites is the field worth reading. It lists exactly the positions
where a dash sits between two non-space characters, which is where --punct
substitutes a bare hyphen. Those are the sentences to rewrite rather than
accept. Curly quotes, ellipses and no-break spaces flatten correctly and are not
listed.
check continues differently, because a gate has no single bar: root,
discovery, config, a surfaces array carrying each surface's own bar,
options and counts, then filesScanned, filesSkipped, filesFailing,
failures with their status and hits, and a skipped list capped at 50 with
skippedTruncated and skippedTotal.
The one exception: report --json has no envelope and its keys are
snake_case. That output is byte-identical to the Python reference report,
including how it prints a whole-number float as 12.0, so a payload from either
implementation diffs cleanly against the other. One result unwraps to the bare
report object; any other number produces a {label: report} map. The condition
is on the number of RESULTS, not of file arguments, so two files where one could
not be read still print a bare object.
Install for your agent
The CLI on its own is enough: any agent that can run a shell command can run
npx -y ghostchars check --bar text. The files below make it happen without
being asked each time. They ship inside the package, so copy them from
$(npm root -g)/ghostchars/integrations/ after a global install, or from
node_modules/ghostchars/integrations/ after a local one.
integrations/README.md has the per-tool copy commands.
The hook. integrations/hooks/ is a Claude Code PostToolUse hook that
holds every file the agent writes or edits to the text bar and hands a finding
back as feedback, so the model fixes it in the same turn. It never rewrites the
file. Copy the script to .claude/hooks/ and merge the settings fragment into
.claude/settings.json.
The skill. integrations/skills/ghostchars/ is one directory holding
SKILL.md and two reference files. Copy it to .claude/skills/ghostchars/ and
to .agents/skills/ghostchars/, or put it in one and symlink the other. That
covers Claude Code, Codex, Cursor, VS Code with Copilot, and Gemini CLI from a
single copy. It is the highest-value file here: it tells the agent when to run
the tool and, more importantly, what to do with each kind of finding.
The always-on rule blocks. integrations/rules/ holds a short paste block
for AGENTS.md, CLAUDE.md and GEMINI.md, plus ready-made rule files for
Cursor, Windsurf and Copilot. A skill loads on demand; these are always in
context. Windsurf has no skill support, so its rule file is the whole integration
there.
The MCP server. integrations/mcp/ holds one config snippet per host. Adding
it gives the agent four read-only tools (clean_text, inspect_text,
style_report, check_paths) with typed inputs and outputs, which saves a shell
round trip and a JSON parse on hosts that prefer structured tools. A fifth tool,
clean_files, writes to disk and is registered only under --allow-write.
Take one of the three or all three. They do not conflict.
Differences from the Python reference CLI
The golden fixtures pin clean and inspect to the Python byte for byte across
every text fixture and every option combination, and report --json to its
exact bytes. Everything below is a deliberate difference, and the tests pin
each one.
The last row runs the other way from the rest of the table: it is the one place this tool
prints more than the reference, not less. Every codepoint the Unicode Character
Database does name is named identically here, look-alike letters included, and a
test compares all 1508 nameable look-alikes against unicodedata in both
directions, so a missing name and a wrong one both fail.
New here, with no Python counterpart: --bar, --json, --stdin-name, --as,
--limit, --fail-on, --fix, the whole check command, and the mcp server.
Two consequences of the atomic write worth knowing: a real write changes the
file's inode, so a hard link to it breaks, and a crash mid-write leaves a
.ghostchars-*.tmp file beside the original instead of a truncated original. A
file that would not change is never opened for writing, so its modification time
and inode survive untouched.
What this cannot do
- Statistical watermarks such as SynthID-Text or KGW. They live in word choice, not in characters. Only rewriting touches them.
- Stylometric detection. Vocabulary, rhythm and structure.
reportmeasures some of the same habits so you can edit them, and it is not a detector: it emits no probability and no score, and it never decides whether a text was generated. - PDF, images,
.pptx,.xlsx,.rtf,.doc,.epub,.pagesand Google Docs. EXIF, XMP and C2PA are a different tool. - Removing what hides around the text. Hidden runs, comments and tracked changes are reported, never deleted.
- Escapes outside
check. A U+2014 EM DASH written as a backslash-u escape in a source file is invisible unless aghostchars.jsonsurface turnsdecodeEscapeson, and there is no equivalent for extracting the string literals out of a language the tool does not parse. - Documents in
checkor over MCP. A gate walking a repository does not open zip archives. - Anything over the network. There is no network code in this package and no telemetry.
来源:cli/README.md,提交 faffeef
工具
0版本历史
1- v1.0.2最新Sep 29, 2026

