Ghostchars

io.github.ghostcharsv1.0.2更新于 Sep 29, 2026

Find and remove invisible Unicode: zero-width, tag smuggling, bidi, homoglyphs. Offline.

已验证STDIO仅桌面Other

安装

在 SourceWeft 中

  1. 打开 控制台中的 Ghostchars,将其添加到工作区。
  2. 为需要使用其工具的对话启用该服务。

Desktop only,通过 STDIO。 STDIO 服务会启动本地进程,因此需要 SourceWeft 桌面宿主。

其他 MCP 客户端

参照 仓库 中的启动说明。

README

ghostchars

Find and remove invisible Unicode characters and machine-writing punctuation, in text and in documents. Everything runs on your machine: no network, no telemetry, no account, and zero runtime dependencies.

npx ghostchars clean -i notes.md          # clean a file in place, atomicallycat draft.md | npx ghostchars inspect     # what is hiding: line:col, codepoint, namenpx ghostchars check --bar text           # the gate: exit 1 if the tree is dirtynpx ghostchars report draft.md            # habits that make a draft read as machine writtennpx ghostchars clean -i essay.docx        # documents too: .docx, .odt, .html

Needs Node 22 or newer. The whole package is a launcher and three bundles: the CLI, the document reader, and the MCP server. Only the first is loaded unless you name a document or start the server.

Commands

clean reads stdin or the files you name, writes the cleaned text to stdout, or rewrites each file in place with -i. With --check it writes nothing and exits 1 if anything would change.

inspect lists every suspicious character with its line, column, codepoint, Unicode name and the flag that would act on it. It changes nothing and exits 1 when it finds something.

report measures the habits that make a draft read as machine written: sentence length uniformity, stock transitions, tell words, dash density, tricolons, hedging, missing specifics. It is not a detector and emits no probability.

check holds a whole tree to a bar and prints every file that would change, with exact character positions. This is the command for CI and for a pre-commit hook. Configure surfaces and pins in ghostchars.json.

mcp runs a Model Context Protocol server on stdin and stdout, so an agent can call clean_text, inspect_text, style_report and check_paths as tools instead of shelling out.

help prints the commands, help <command> prints one command's flags, and help --json prints the whole surface as one machine-readable document. version prints the version.

Flags

The thirteen cleaning switches, their spellings and their meanings are exactly the Python reference CLI's, because it is the same engine underneath.

flagwhatwhy it is not default
(default)zero-width space, word joiner, byte-order mark, soft hyphen, invisible math operators, bidi overrides, free-floating tag characters, noncharacters, reserved default-ignorable codepoints, every other format control
-a, --aggressiveglue characters even when they are load-bearingbreaks emoji sequences, Persian, Indic, subdivision flags
-s, --spacesno-break and other exotic spaces to a plain space
-n, --newlinesCRLF, CR, NEL, LS and PS to a plain newlinechanges line endings
--bidibidi marks and isolates even next to right-to-left textbreaks mixed Hebrew and Arabic
--controlsC0 and C1 control characters, keeping tab and the line endings
--puaprivate-use charactersbreaks icon-font glyphs
--punctem and en dashes, curly quotes, ellipsis and minus to ASCIItypography, not hiding
--confusableslook-alike Cyrillic and Greek letters to Latin, only inside mixed-script wordsrewrites letters; opt in on principle
--nfc, --nfkcUnicode normalisation before anything else runsNFKC rewrites letters and is lossy
--guttersthe quote bar a terminal draws at the start of a line (U+258C to U+258F, U+2502, U+2503) and one space after it, keeping indentation and box-drawn table rowsa directory tree or a drawn box can start a line with the same bar
--trailingtrailing spaces and tabs at the end of every linemarkdown hard breaks
--collapsea run of two or more spaces inside a line to one
--maxall of the above except --nfc, which --nfkc supersedes

Default cleaning is context aware. It strips a zero-width joiner sitting between two Latin letters and keeps the one inside a family emoji or a Persian word. It strips free tag characters, which is how text is smuggled into a message, and keeps a complete subdivision flag. It strips a left-to-right mark in English text and keeps one next to Hebrew. -a turns that judgement off.

An option can only be turned on. There is no way to turn one off, because the reference CLI has none either.

The three bars

A bar is a named set of options, so a person or an agent picks one word instead of thirteen booleans.

baroptionsuse it for
defaultnoneinvisible characters only. Nothing legitimately needs them
text-s --punctanything a person will read. This is the bar for your own prose
maxthe twelve of --maxeverything, including confusables and NFKC. Lossy

Resolution order is fixed, so the result is always predictable:

  1. start from every option off,
  2. apply --bar,
  3. apply --max if given,
  4. apply each explicit flag, which can only turn an option on.

So --bar text --collapse is -s --punct --collapse, and --bar max --punct is still just max.

max folds look-alike Cyrillic and Greek letters and applies NFKC, which rewrites letters. Do not run it over text you did not write. It is not offered as a gate bar at all: see check below.

Documents

clean and inspect recognise a document by its extension and handle it structurally.

extensionhowextra flags
.docxthe package is unzipped, every text node is cleaned, the zip is rewritten--metadata, --rsids
.odtthe same, through the ODF parts--metadata
.html, .htmtext nodes only. Tags, attributes, script and style are untouched
anything elseplain text

--as text forces plain-text treatment of any path, which is what the Python reference CLI does with an .html file.

Without -i, a cleaned document is written to <name>.clean<ext> beside the original, or to the path you give with -o.

inspect on a document prints the character report over the concatenated document text, then what hides around that text: hidden runs, white-on-white and tiny runs, tracked changes, comments, the count of distinct rsid editing-session ids, metadata fields, macros, media. --metadata blanks the authoring fields and drops custom properties; --rsids removes the editing-session fingerprint.

Hidden runs, comments and tracked changes are reported and never removed. Deleting them changes the document, and that stays a human call.

Out of scope here: PDF, images, .pptx, .xlsx, .rtf, .doc, .epub and .pages. check does not open documents at all.

check

ghostchars check [--bar default|text] [option flags]                 [--config <path>] [--no-config]                 [--git|--no-git] [--max-bytes N] [--limit N]                 [--fix] [--json] [--verbose] [paths...]

--bar accepts exactly default and text. There is deliberately no max gate bar: max turns on confusables and NFKC, which target scripts and symbols rather than typography, and a gate that condemns legitimate Cyrillic, Greek and emoji content gets deleted by the people who installed it. The default bar is default, for the same reason: typography is a house style and has to be opted into.

With no path arguments, check finds files with git ls-files, so it honours the same ignore rules your repository already has and still sees files that are new and not yet added. Outside a repository, or with --no-git, it walks the tree instead and skips node_modules, build output, virtual environments and every directory whose name starts with a dot, except the ones coding agents read instructions from (.cursor, .github, .claude, .agents, .windsurf, .devin, .gemini, .codex, .clinerules, .continue, .kiro, .roo). --git makes a git that cannot be used a hard error, whether or not you named paths, so CI cannot silently switch strategy between one container and another. Path arguments are always walked rather than looked up in git, so --git <paths> asserts that git was available and reports "discovery": "walk". Results are sorted by path, so the output never depends on directory order. A directory the walk cannot open is an error on stderr and exit 2, never a silent omission.

A file is skipped, silently and without failing the run, when its extension is a known binary one, when it is larger than --max-bytes (2 MiB by default), when one of its first 8192 bytes is NUL, or when it is not valid UTF-8. A file you name on the command line is different: you asserted it is text, so both content checks stop applying to it, the strict decode decides, and a decode failure is an error that exits 2.

check <dir> finds <dir>'s own ghostchars.json even when you run it from somewhere else: with no config beside the working directory, the search moves up from each path you named, stopping at that path's repository root. Two paths governed by two different config files is an error, because one run has one root and a pin is written relative to it. --verbose prints which config was used.

ghostchars.json

Put it at the root of the tree, or point at it with --config.

json
{  "version": 1,  "surfaces": [    {      "id": "invisibles",      "bar": "default",      "include": ["**/*"],      "exclude": ["vendor/**"]    },    {      "id": "prose",      "bar": "text",      "include": ["**/*.md", "src/messages/*.ts"],      "decodeEscapes": true,      "pins": { "docs/glyphs.md": 12 }    }  ],  "maxBytes": 2097152}

A surface names a bar (or an options list, one of the two, never both), the globs it covers, and its pins. A file that matches no surface is not checked at all. An exclusion wins over everything, including a path you named explicitly. An unknown key is an error that lists the known keys, so a typo in pins is never silently ignored.

The glob subset is ** for any run of path segments, * for any run inside one segment, and ? for one character. No braces, no character classes.

decodeEscapes resolves \uXXXX and \u{XXXX} to the characters they render before the file is checked. Turn it on for a source file whose strings a person reads, and leave it off everywhere else: a table of codepoints is full of escapes that match characters rather than ship them.

Pins

A pin says a file changes by exactly N characters on purpose. It is enforced in both directions.

pinnedmeasuredresult
none0passes, silently
noneN greater than 0FAIL changed, every character listed
NNpasses
NM, and M is not NFAIL pin-mismatch: changes M characters, but exactly N are pinned
N greater than 00FAIL pin-stale: pinned as deliberately dirty but now changes nothing. Unpin it

That last row is the one people forget. It catches a pin that went stale when somebody quietly fixed the file.

Output

docs/post.md: the text bar would change this file:  docs/post.md:1:26 U+2014 EM DASH      Ghostchars is a text tool<EM DASH>it finds what a reader cannot see.  docs/post.md:2:43 U+200B ZERO WIDTH SPACE      There is a zero<ZWSP>width space in this one.
ghostchars: 1 file would change, 3 scanned

The two placeholders stand for the real characters. This file is held to the text bar itself, so it names a codepoint where it would otherwise print one.

One path:line:col U+XXXX NAME line per offending character, and the source line printed once per distinct line number. Exit 1. When everything passes, one line: ghostchars: 3 files pass the text bar. --verbose also names the files that passed.

In CI

npx -y ghostchars check --bar text

That is the whole step. Never pass --fix in CI. --fix cleans the failing files in place and then re-checks, which turns a red build green without anyone reading what changed. It also skips pinned files and names them, because fixing one would invalidate its own pin.

After a local --fix, read back any file that had a dash finding. The cleaner turns an em dash between two words into a space, a hyphen and a space, which is correct and reads badly. Those positions are exactly what rewriteSites in the JSON output lists.

--sarif prints the same run as a SARIF 2.1.0 log, one result per offending character with its line, column and source line, and one result per pin failure. GitHub code scanning shows those inline on the pull request:

- run: npx -y ghostchars check --bar text --sarif > ghostchars.sarif || true- uses: github/codeql-action/upload-sarif@v3  with:    sarif_file: ghostchars.sarif

The || true is there because exit 1 means findings, and the upload step is what turns them into annotations. Keep a second plain check step if the job must also fail. --sarif and --json are two formats for one stdout, so the command refuses both at once.

As a pre-commit hook

integrations/pre-commit/pre-commit-config.yaml is a paste block for .pre-commit-config.yaml. It is a repo: local hook with language: node, so the framework installs the pinned npm package into its own environment and no repository is cloned. It runs over the staged text files and, like the CI step, never passes --fix.

Exit codes

codemeaningfrom
0done, nothing to reportevery command
1findings, would-change, or a gate failureinspect, clean --check, check, report --fail-on
2the run could not be completed as asked: a usage error, an unreadable or non-UTF-8 file, a bad config, a broken document, a Node older than 22every command
3an internal error, which is a bug in ghostcharsevery command

The priority is strict and accumulates across inputs: 3 beats 2 beats 1 beats 0. A decode failure in one file forces 2 even when another file had findings.

Code 3 prints one line and, with GHOSTCHARS_DEBUG=1 in the environment, a stack trace. It exists so an agent can tell "I passed the wrong flag" from "the tool broke", which are different next actions.

EPIPE on stdout is a silent exit 0, so ghostchars inspect big.md | head behaves like every other command.

JSON output

--json on clean, inspect and check prints exactly one JSON object to stdout, never a stream of lines. Four rules, all of them for a machine reader: ok and exitCode are in the payload as well as in the process status; counts come before lists, so a reader that truncates still has the verdict; a cap that bites always sets findingsTruncated and findingsTotal; and keys are camelCase.

All three open with the same five identity fields, so one reader can tell what it is holding before it looks at anything else.

json
{ "tool": "ghostchars", "version": "1.0.2", "command": "clean", "ok": true, "exitCode": 0}

clean and inspect continue with bar, the resolved options list, a summary that totals across inputs, and results, one entry per input. A clean result carries changed, its own summary, a charsDelta in codepoints, the whole cleaned text (replaced by written: true under -i), and rewriteSites. An inspect result carries clean, total, a summary, a byKind rollup, then findings, each with line, col, index, cp, u, name, kind, flag, action and, for a substitution only, replacement.

rewriteSites is the field worth reading. It lists exactly the positions where a dash sits between two non-space characters, which is where --punct substitutes a bare hyphen. Those are the sentences to rewrite rather than accept. Curly quotes, ellipses and no-break spaces flatten correctly and are not listed.

check continues differently, because a gate has no single bar: root, discovery, config, a surfaces array carrying each surface's own bar, options and counts, then filesScanned, filesSkipped, filesFailing, failures with their status and hits, and a skipped list capped at 50 with skippedTruncated and skippedTotal.

The one exception: report --json has no envelope and its keys are snake_case. That output is byte-identical to the Python reference report, including how it prints a whole-number float as 12.0, so a payload from either implementation diffs cleanly against the other. One result unwraps to the bare report object; any other number produces a {label: report} map. The condition is on the number of RESULTS, not of file arguments, so two files where one could not be read still print a bare object.

Install for your agent

The CLI on its own is enough: any agent that can run a shell command can run npx -y ghostchars check --bar text. The files below make it happen without being asked each time. They ship inside the package, so copy them from $(npm root -g)/ghostchars/integrations/ after a global install, or from node_modules/ghostchars/integrations/ after a local one. integrations/README.md has the per-tool copy commands.

The hook. integrations/hooks/ is a Claude Code PostToolUse hook that holds every file the agent writes or edits to the text bar and hands a finding back as feedback, so the model fixes it in the same turn. It never rewrites the file. Copy the script to .claude/hooks/ and merge the settings fragment into .claude/settings.json.

The skill. integrations/skills/ghostchars/ is one directory holding SKILL.md and two reference files. Copy it to .claude/skills/ghostchars/ and to .agents/skills/ghostchars/, or put it in one and symlink the other. That covers Claude Code, Codex, Cursor, VS Code with Copilot, and Gemini CLI from a single copy. It is the highest-value file here: it tells the agent when to run the tool and, more importantly, what to do with each kind of finding.

The always-on rule blocks. integrations/rules/ holds a short paste block for AGENTS.md, CLAUDE.md and GEMINI.md, plus ready-made rule files for Cursor, Windsurf and Copilot. A skill loads on demand; these are always in context. Windsurf has no skill support, so its rule file is the whole integration there.

The MCP server. integrations/mcp/ holds one config snippet per host. Adding it gives the agent four read-only tools (clean_text, inspect_text, style_report, check_paths) with typed inputs and outputs, which saves a shell round trip and a JSON parse on hosts that prefer structured tools. A fifth tool, clean_files, writes to disk and is registered only under --allow-write.

Take one of the three or all three. They do not conflict.

Differences from the Python reference CLI

The golden fixtures pin clean and inspect to the Python byte for byte across every text fixture and every option combination, and report --json to its exact bytes. Everything below is a deliberate difference, and the tests pin each one.

behaviourPythonghostchars
an abbreviated long option, --punc for --punctacceptedunknown flag, exit 2, with the nearest known flag named
usage text on a parse errorargparse's wrapped blockone line, a suggestion, and a pointer to help --json
a missing or unreadable filetraceback, exit 1{path}: no such file or directory, exit 2
an in-place writetruncate and writetemp file, fsync, mode copied, symlink resolved, rename
.html and .htm in cleancleaned as plain text, markup includedcleaned structurally; --as text restores the old behaviour
an unrecognised extensionan unsupported type errorplain text
the report human rendermiddle dot and en dash separatorsASCII, , and to. Byte parity for report lives in --json
report on an unreadable fileprints the exception, always exits 0prints the message, exits 2
an internal errortraceback, exit 1one line, exit 3, stack under GHOSTCHARS_DEBUG=1
SIGPIPE on stdouta message, exit 120silent exit 0
findings under --nfc or --nfkcseparate code paths, never cross-referencedclean --json scans findings from the original text, so a normalised character can read as found-not-targeted. The cleaned text is authoritative
the NAME column for a codepoint the Unicode database does not name<unnamed>a C0 or C1 control's standard alias (BELL, NEXT LINE) where it has one, U+XXXX otherwise

The last row runs the other way from the rest of the table: it is the one place this tool prints more than the reference, not less. Every codepoint the Unicode Character Database does name is named identically here, look-alike letters included, and a test compares all 1508 nameable look-alikes against unicodedata in both directions, so a missing name and a wrong one both fail.

New here, with no Python counterpart: --bar, --json, --stdin-name, --as, --limit, --fail-on, --fix, the whole check command, and the mcp server.

Two consequences of the atomic write worth knowing: a real write changes the file's inode, so a hard link to it breaks, and a crash mid-write leaves a .ghostchars-*.tmp file beside the original instead of a truncated original. A file that would not change is never opened for writing, so its modification time and inode survive untouched.

What this cannot do

  • Statistical watermarks such as SynthID-Text or KGW. They live in word choice, not in characters. Only rewriting touches them.
  • Stylometric detection. Vocabulary, rhythm and structure. report measures some of the same habits so you can edit them, and it is not a detector: it emits no probability and no score, and it never decides whether a text was generated.
  • PDF, images, .pptx, .xlsx, .rtf, .doc, .epub, .pages and Google Docs. EXIF, XMP and C2PA are a different tool.
  • Removing what hides around the text. Hidden runs, comments and tracked changes are reported, never deleted.
  • Escapes outside check. A U+2014 EM DASH written as a backslash-u escape in a source file is invisible unless a ghostchars.json surface turns decodeEscapes on, and there is no equivalent for extracting the string literals out of a language the tool does not parse.
  • Documents in check or over MCP. A gate walking a repository does not open zip archives.
  • Anything over the network. There is no network code in this package and no telemetry.

来源:cli/README.md,提交 faffeef

工具

0
工具元数据尚未被收录。

版本历史

1
  1. v1.0.2最新Sep 29, 2026