Graph Evolution
Builds Trailmark code graphs at two source snapshots and computes a structural diff. Surfaces security-relevant changes that text-level diffs miss: new attack paths, complexity shifts, blast radius growth, taint propagation changes, and privilege boundary modifications.
When to Use
- Comparing two git refs to understand what structurally changed
- Auditing a range of commits for security-relevant evolution
- Detecting new attack paths created by code changes
- Finding functions whose blast radius or complexity grew silently
- Identifying taint propagation changes across refactors
- Pre-release structural comparison (tag-to-tag or branch-to-branch)
When NOT to Use
- Line-level code review (use
differential-reviewfor text-diff analysis) - Single-snapshot analysis (use the
trailmarkskill directly) - Diagram generation from a single snapshot (use the
diagramming-codeskill) - Mutation testing triage (use the
genotoxicskill)
Rationalizations to Reject
Prerequisites
trailmark must be installed. If uv run trailmark fails, run:
DO NOT fall back to "manual comparison" or reading source files as a substitute for running trailmark. The tool must be installed and used programmatically. If installation fails, report the error.
Quick Start
Decision Tree
Workflow
Phase 1: Create Snapshots
Use git worktrees to get clean copies of each ref without disturbing the working tree.
If comparing two directories instead of git refs, skip this phase and use the directory paths directly in Phase 2.
Phase 2: Build Graphs and Run Pre-Analysis
Build Trailmark graphs for both snapshots and run pre-analysis on each. Pre-analysis computes blast radius, taint propagation, privilege boundaries, and entrypoint enumeration.
Verify both graphs built successfully by checking the summary output.
If either fails, rerun with an explicit language or comma-separated list
instead of auto.
Phase 3: Compute Structural Diff
Run both:
- Trailmark's native structural diff for nodes, edges, and entrypoints
- The plugin's
graph_diff.pyhelper for subgraph membership changes
Use the same work_dir from Phase 2, and pass the same --language value Phase 2
built with. trailmark diff defaults that flag to python, so on any other
target the default exits 0 and writes empty arrays rather than reporting a
mismatch.
If Phase 2 needed an explicit language or a comma-separated list instead of
auto, use that same value here.
If either diff command fails or writes an empty JSON file, stop and report the error instead of continuing to Phase 4.
A trailmark_diff.json whose nodes, edges, and entrypoints arrays are all
empty means either nothing changed structurally or both snapshots parsed to
(near-)empty graphs. Decide which using Phase 2's graph summaries: if either
snapshot's node count is zero or implausibly small for the target, the parse
missed the code — name the language set explicitly (rust, solidity,
python,rust) and re-run. Healthy node counts on both snapshots plus an empty
diff is genuine structural stability.
The native Trailmark diff contains:
The subgraph diff contains:
Phase 4: Interpret Diff and Generate Report
Read both diff JSON files and generate a security-focused markdown report. See references/report-format.md [blocked] for the full template.
Interpretation priorities (highest to lowest):
- New tainted paths — nodes entering the
taintedsubgraph, especially if they also appear in added edges targeting sensitive functions - Privilege boundary changes — new or removed trust transitions from the native entrypoint/edge diff plus the subgraph diff
- Attack surface growth — new entrypoints, especially
untrusted_external, fromtrailmark_diff.json - Blast radius increases — nodes entering
high_blast_radius - Complexity spikes — CC increases > 3 on tainted or entrypoint-reachable nodes
- Structural additions — new nodes and edges (review needed)
- Structural removals — verify removed security functions were replaced
Cross-reference structural changes with git diff {before_ref}..{after_ref}
to add source-level context to findings.
Severity classification:
For detailed metric definitions, see references/evolution-metrics.md [blocked].
Phase 5: Clean Up
Remove git worktrees after the report is written:
Diff Reference
trailmark diff --language defaults to python. On a target in any other
language that default still exits 0, emitting well-formed JSON with empty
nodes, edges, and entrypoints arrays, so always pass the flag: auto
detects and merges every supported language found under the target, and a single
name (rust, solidity) or comma-separated list (python,rust) pins an
explicit set. auto fails loudly with No supported languages detected under <path> when a snapshot holds nothing it can parse, which is the outcome you
want. Confirm the language first; only then can an empty diff count as evidence
that nothing changed.
Use trailmark diff for:
- Node/edge changes
- Added/removed/modified entrypoints
- Human-readable structural diff reports
Use graph_diff.py for:
- Subgraph membership changes derived from
engine.preanalysis() tainted,high_blast_radius,privilege_boundary, and related sets
graph_diff.py input format: Trailmark JSON exports from engine.to_json().
graph_diff.py output: JSON structural diff for nodes, edges, and subgraphs.
Quality Checklist
Before delivering the report:
- Both graphs built successfully (check summaries)
- Pre-analysis ran on both snapshots
- Native Trailmark diff computed (
trailmark_diff.json); if it is empty, both snapshots' Phase 2 node counts were non-zero, so empty means stable - Subgraph diff computed and non-empty (
subgraph_diff.json) - All subgraph changes interpreted (tainted, blast radius, etc.)
- Critical findings include evidence (node IDs, edge diffs)
- Severity levels assigned to all findings
- Source-level context added via git diff cross-reference
- Worktrees cleaned up (or temp dirs removed)
- Report written to
GRAPH_EVOLUTION_*.md
Integration
trailmark skill: Phase 2 uses the trailmark API for graph building and pre-analysis. All trailmark query patterns work on either snapshot's engine.
differential-review skill: Use graph-evolution for structural analysis, differential-review for line-level code review. The two are complementary — graph-evolution finds attack paths that text diffs miss, while differential-review provides git blame context and micro-adversarial analysis.
trailmark-review-gate skill: Use trailmark-review-gate after graph-evolution when a branch, pull request, fix commit, or release diff needs a PASS/WARN/FAIL/UNKNOWN structural review packet. The gate applies deterministic review rules to graph-evolution output; it does not replace human review.
genotoxic skill: If graph-evolution reveals new high-CC tainted nodes, feed them to genotoxic for mutation testing triage.
diagramming-code skill:
Generate before/after diagrams to visualize structural changes.
Use call-graph or data-flow diagrams focused on changed nodes.
Supporting Documentation
- references/evolution-metrics.md [blocked] — What each structural metric means and why it matters for security
- references/report-format.md [blocked] — Report template, severity classification, and example findings

