Audit Augmentation
Projects findings from external tools (SARIF) and human auditors (weAudit)
onto Trailmark code graphs as annotations and subgraphs. Trailmark 0.4.0+ can
also import an external binary-analysis graph JSON export via
engine.augment_binary().
When to Use
- Importing Semgrep, CodeQL, or other SARIF-producing tool results into a graph
- Importing weAudit audit annotations into a graph
- Importing binary-analysis graph data into a source graph (Trailmark 0.4.0+)
- Cross-referencing static analysis findings with blast radius or taint data
- Querying which functions have high-severity findings
- Visualizing audit coverage alongside code structure
- Preparing one SARIF or weAudit result for
trailmark-finding-triage
When NOT to Use
- Running static analysis tools (use semgrep/codeql directly, then import)
- Building the code graph itself (use the
trailmarkskill) - Generating diagrams (use the
diagramming-codeskill after augmenting)
Rationalizations to Reject
Installation
MANDATORY: If uv run trailmark fails, install trailmark first:
Version Gate
SARIF and weAudit augmentation are v0.2-safe. Binary graph augmentation is
Trailmark 0.4.0+ only. Before calling engine.augment_binary(), check:
On Trailmark 0.5.0+, known links between source functions and imported binary
or external endpoints can also be declared once in .trailmark/links.toml
(see the main trailmark skill's Repository Links section) instead of being
re-derived per session. Declared external endpoints materialize as
proxy.external:<symbol> nodes on every parse.
Quick Start
CLI
Binary graph augmentation is programmatic in Trailmark 0.4.0+; do not invent a
CLI flag if trailmark augment --help does not show one.
Programmatic API
If auto-detection is wrong for the target, rerun with an explicit language or
comma-separated list such as python,rust.
Workflow
Step 1: Build the graph and run pre-analysis for blast radius and taint context:
If auto-detection is wrong for the target, rerun with an explicit language or
comma-separated list such as python,rust.
Step 2: Locate input files:
- SARIF: Usually output by tools like
semgrep --sarif -o results.sariforcodeql database analyze --format=sarif-latest - weAudit: Stored in
.vscode/<username>.weauditwithin the workspace - Binary graph (v0.4+): External JSON with
artifact,functions, andcallsfields. Trailmark imports this graph; it does not disassemble binaries itself.
Step 3: Run augmentation via engine.augment_sarif() or
engine.augment_weaudit(). For binary graphs, run engine.augment_binary()
only after the Version Gate succeeds. Check unmatched_findings in SARIF and
weAudit results — these are findings whose file/line locations didn't overlap
any parsed code unit.
Step 4: Query findings and subgraphs. Use engine.findings() to list all
annotated nodes. Use engine.subgraph_names() to see available subgraphs.
Step 5: Cross-reference with pre-analysis data to prioritize:
- Findings on tainted nodes: overlap
sarif:errorwithtaintedsubgraph - Findings on high blast radius nodes: overlap with
high_blast_radius - Findings on privilege boundaries: overlap with
privilege_boundary
For one candidate finding that needs a reachability verdict or PoC handoff,
continue with trailmark-finding-triage and use the augmented node as the
bound candidate.
Annotation Format
Findings are stored as standard Trailmark annotations:
- Kind:
finding(tool-generated) oraudit_note(human notes) - Source:
sarif:<tool_name>orweaudit:<author> - Description: Compact single-line:
[SEVERITY] rule-id: message (tool)
Subgraphs Created
How Matching Works
Findings are matched to graph nodes by file path and line range overlap:
- Finding file path is normalized relative to the graph's
root_path - Nodes whose
location.file_pathmatches AND whose line range overlaps are selected - The tightest match (smallest span) is preferred
- If a finding's location doesn't overlap any node, it counts as unmatched
SARIF paths may be relative, absolute, or file:// URIs — all are handled.
weAudit uses 0-indexed lines which are converted to 1-indexed automatically.
Binary graph imports create origin=binary function nodes, origin=proxy
external proxy nodes for unresolved binary calls, and inferred
corresponds_to edges when a binary function maps back to a source node. The
expected JSON shape is intentionally small:
Supporting Documentation
- references/formats.md [blocked] — SARIF 2.1.0 and weAudit file format field reference

