Audit Augmentation

作者 trailofbits82fe82262526无许可证7.4K 个星标收录于 2026年10月8日更新于 2026年10月8日仓库昨天更新

Augments Trailmark code graphs with external audit findings from SARIF static analysis results, weAudit annotation files, and version-gated Trailmark 0.4.x binary-analysis graph exports. Maps findings to graph nodes by file and line overlap, creates severity-based subgraphs, and enables cross-referencing findings with pre-analysis data (blast radius, taint, etc.). Use when projecting SARIF results onto a code graph, overlaying weAudit annotations, importing binary graph findings, cross-referencing Semgrep, CodeQL, or binary-analysis findings with call graph data, or visualizing audit findings in the context of code structure.

AI 生成的概览

将 SARIF、weAudit 和二进制分析结果作为标注与严重度子图投射到 Trailmark 代码图上。

功能
该技能说明如何用外部审计结果扩充 Trailmark 代码图:SARIF 静态分析结果、weAudit 标注文件,以及受版本限制的 Trailmark 0.4.x 二进制分析图导出。它按文件路径与行范围重叠把结果匹配到图节点,存为标注,并创建按严重度和工具划分的子图。它还涵盖将结果与预分析数据(影响范围、污点、权限边界)交叉引用,以及报告未匹配的结果。
适用场景
适用于把 Semgrep、CodeQL 等 SARIF 输出导入代码图、叠加 weAudit 标注、导入二进制图结果,或查询并可视化哪些函数带有高严重度结果。不适用于运行分析工具本身、构建代码图或生成图表。
运行要求
需要 Trailmark 工具(通过 uv tool install trailmark 安装)和 Trailmark 代码图;二进制图扩充需 Trailmark 0.4.0+,仓库链接需 0.5.0+。输入为 SARIF 文件、.vscode/<用户名>.weaudit 文件或二进制图 JSON 导出。不附带脚本,仅为说明文档。

Audit Augmentation

Projects findings from external tools (SARIF) and human auditors (weAudit) onto Trailmark code graphs as annotations and subgraphs. Trailmark 0.4.0+ can also import an external binary-analysis graph JSON export via engine.augment_binary().

When to Use

  • Importing Semgrep, CodeQL, or other SARIF-producing tool results into a graph
  • Importing weAudit audit annotations into a graph
  • Importing binary-analysis graph data into a source graph (Trailmark 0.4.0+)
  • Cross-referencing static analysis findings with blast radius or taint data
  • Querying which functions have high-severity findings
  • Visualizing audit coverage alongside code structure
  • Preparing one SARIF or weAudit result for trailmark-finding-triage

When NOT to Use

  • Running static analysis tools (use semgrep/codeql directly, then import)
  • Building the code graph itself (use the trailmark skill)
  • Generating diagrams (use the diagramming-code skill after augmenting)

Rationalizations to Reject

RationalizationWhy It's WrongRequired Action
"The user only asked about SARIF, skip pre-analysis"Without pre-analysis, you can't cross-reference findings with blast radius or taintAlways run engine.preanalysis() before augmenting
"Unmatched findings don't matter"Unmatched findings may indicate parsing gaps or out-of-scope filesReport unmatched count and investigate if high
"One severity subgraph is enough"Different severities need different triage workflowsQuery all severity subgraphs, not just error
"SARIF results speak for themselves"Findings without graph context lack blast radius and taint reachabilityCross-reference with pre-analysis subgraphs
"weAudit and SARIF overlap, pick one"Human auditors and tools find different thingsImport both when available
"Tool isn't installed, I'll do it manually"Manual analysis misses what tooling catchesInstall trailmark first

Installation

MANDATORY: If uv run trailmark fails, install trailmark first:

bash
uv tool install trailmark# Python snippets: uv run --with trailmark python -   (a tool env is not importable)

Version Gate

SARIF and weAudit augmentation are v0.2-safe. Binary graph augmentation is Trailmark 0.4.0+ only. Before calling engine.augment_binary(), check:

python
if not hasattr(engine, "augment_binary"):    raise RuntimeError("Binary augmentation requires Trailmark >= 0.4.0")

On Trailmark 0.5.0+, known links between source functions and imported binary or external endpoints can also be declared once in .trailmark/links.toml (see the main trailmark skill's Repository Links section) instead of being re-derived per session. Declared external endpoints materialize as proxy.external:<symbol> nodes on every parse.

Quick Start

CLI

bash
# Augment with SARIFuv run trailmark augment {targetDir} --sarif results.sarif
# Augment with weAudituv run trailmark augment {targetDir} --weaudit .vscode/alice.weaudit
# Both at once, output JSONuv run trailmark augment {targetDir} \    --sarif results.sarif \    --weaudit .vscode/alice.weaudit \    --json

Binary graph augmentation is programmatic in Trailmark 0.4.0+; do not invent a CLI flag if trailmark augment --help does not show one.

Programmatic API

python
from trailmark.query.api import QueryEngine
engine = QueryEngine.from_directory("{targetDir}", language="auto")
# Run pre-analysis first for cross-referencingengine.preanalysis()
# Augment with SARIFresult = engine.augment_sarif("results.sarif")# result: {matched_findings: 12, unmatched_findings: 3, subgraphs_created: [...]}
# Augment with weAuditresult = engine.augment_weaudit(".vscode/alice.weaudit")
# Augment with an external binary graph export (v0.4+)if hasattr(engine, "augment_binary"):    result = engine.augment_binary("binary_graph.json")
# Query findingsengine.findings()                                       # All findingsengine.subgraph("sarif:error")                          # High-severity SARIFengine.subgraph("weaudit:high")                         # High-severity weAuditengine.subgraph("sarif:semgrep")                        # By tool nameengine.annotations_of("function_name")                  # Per-node lookup

If auto-detection is wrong for the target, rerun with an explicit language or comma-separated list such as python,rust.

Workflow

Augmentation Progress:- [ ] Step 1: Build graph and run pre-analysis- [ ] Step 2: Locate SARIF/weAudit/binary graph files- [ ] Step 3: Run augmentation- [ ] Step 4: Inspect results and subgraphs- [ ] Step 5: Cross-reference with pre-analysis

Step 1: Build the graph and run pre-analysis for blast radius and taint context:

python
engine = QueryEngine.from_directory("{targetDir}", language="auto")engine.preanalysis()

If auto-detection is wrong for the target, rerun with an explicit language or comma-separated list such as python,rust.

Step 2: Locate input files:

  • SARIF: Usually output by tools like semgrep --sarif -o results.sarif or codeql database analyze --format=sarif-latest
  • weAudit: Stored in .vscode/<username>.weaudit within the workspace
  • Binary graph (v0.4+): External JSON with artifact, functions, and calls fields. Trailmark imports this graph; it does not disassemble binaries itself.

Step 3: Run augmentation via engine.augment_sarif() or engine.augment_weaudit(). For binary graphs, run engine.augment_binary() only after the Version Gate succeeds. Check unmatched_findings in SARIF and weAudit results — these are findings whose file/line locations didn't overlap any parsed code unit.

Step 4: Query findings and subgraphs. Use engine.findings() to list all annotated nodes. Use engine.subgraph_names() to see available subgraphs.

Step 5: Cross-reference with pre-analysis data to prioritize:

  • Findings on tainted nodes: overlap sarif:error with tainted subgraph
  • Findings on high blast radius nodes: overlap with high_blast_radius
  • Findings on privilege boundaries: overlap with privilege_boundary

For one candidate finding that needs a reachability verdict or PoC handoff, continue with trailmark-finding-triage and use the augmented node as the bound candidate.

Annotation Format

Findings are stored as standard Trailmark annotations:

  • Kind: finding (tool-generated) or audit_note (human notes)
  • Source: sarif:<tool_name> or weaudit:<author>
  • Description: Compact single-line: [SEVERITY] rule-id: message (tool)

Subgraphs Created

SubgraphContents
sarif:errorNodes with SARIF error-level findings
sarif:warningNodes with SARIF warning-level findings
sarif:noteNodes with SARIF note-level findings
sarif:<tool>Nodes flagged by a specific tool
weaudit:highNodes with high-severity weAudit findings
weaudit:mediumNodes with medium-severity weAudit findings
weaudit:lowNodes with low-severity weAudit findings
weaudit:findingsAll weAudit findings (entryType=0)
weaudit:notesAll weAudit notes (entryType=1)
binary:<artifact>Binary function nodes imported from a v0.4+ binary graph

How Matching Works

Findings are matched to graph nodes by file path and line range overlap:

  1. Finding file path is normalized relative to the graph's root_path
  2. Nodes whose location.file_path matches AND whose line range overlaps are selected
  3. The tightest match (smallest span) is preferred
  4. If a finding's location doesn't overlap any node, it counts as unmatched

SARIF paths may be relative, absolute, or file:// URIs — all are handled. weAudit uses 0-indexed lines which are converted to 1-indexed automatically.

Binary graph imports create origin=binary function nodes, origin=proxy external proxy nodes for unresolved binary calls, and inferred corresponds_to edges when a binary function maps back to a source node. The expected JSON shape is intentionally small:

json
{  "artifact": {"name": "libexample", "architecture": "x86_64", "sha256": "..."},  "functions": [    {"symbol": "parse_packet", "address": "0x401000",     "source": {"file": "src/parser.c", "line": 42}}  ],  "calls": [    {"source": "parse_packet", "target": "malloc", "confidence": "inferred"}  ]}

Supporting Documentation

  • references/formats.md [blocked] — SARIF 2.1.0 and weAudit file format field reference

来源与署名

来源:trailofbits/skills位于plugins/trailmark/skills/audit-augmentation提交82fe822

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架