Audit Augmentation

作者 trailofbits82fe82262526無授權條款7.4K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫昨天更新

Augments Trailmark code graphs with external audit findings from SARIF static analysis results, weAudit annotation files, and version-gated Trailmark 0.4.x binary-analysis graph exports. Maps findings to graph nodes by file and line overlap, creates severity-based subgraphs, and enables cross-referencing findings with pre-analysis data (blast radius, taint, etc.). Use when projecting SARIF results onto a code graph, overlaying weAudit annotations, importing binary graph findings, cross-referencing Semgrep, CodeQL, or binary-analysis findings with call graph data, or visualizing audit findings in the context of code structure.

AI 產生的概覽

將 SARIF、weAudit 與二進位分析結果以標註和嚴重度子圖投射到 Trailmark 程式碼圖上。

功能
此技能說明如何用外部稽核結果擴充 Trailmark 程式碼圖:SARIF 靜態分析結果、weAudit 標註檔案,以及受版本限制的 Trailmark 0.4.x 二進位分析圖匯出。它依檔案路徑與行範圍重疊把結果對應到圖節點,存成標註,並建立依嚴重度和工具區分的子圖。它也涵蓋將結果與預分析資料(影響範圍、汙點、權限邊界)交叉參照,以及回報未對應的結果。
適用情境
適用於把 Semgrep、CodeQL 等 SARIF 輸出匯入程式碼圖、疊加 weAudit 標註、匯入二進位圖結果,或查詢並視覺化哪些函式帶有高嚴重度結果。不適用於執行分析工具本身、建置程式碼圖或產生圖表。
執行需求
需要 Trailmark 工具(以 uv tool install trailmark 安裝)和 Trailmark 程式碼圖;二進位圖擴充需 Trailmark 0.4.0+,儲存庫連結需 0.5.0+。輸入為 SARIF 檔案、.vscode/<使用者名稱>.weaudit 檔案或二進位圖 JSON 匯出。不附帶指令碼,僅為說明文件。

Audit Augmentation

Projects findings from external tools (SARIF) and human auditors (weAudit) onto Trailmark code graphs as annotations and subgraphs. Trailmark 0.4.0+ can also import an external binary-analysis graph JSON export via engine.augment_binary().

When to Use

  • Importing Semgrep, CodeQL, or other SARIF-producing tool results into a graph
  • Importing weAudit audit annotations into a graph
  • Importing binary-analysis graph data into a source graph (Trailmark 0.4.0+)
  • Cross-referencing static analysis findings with blast radius or taint data
  • Querying which functions have high-severity findings
  • Visualizing audit coverage alongside code structure
  • Preparing one SARIF or weAudit result for trailmark-finding-triage

When NOT to Use

  • Running static analysis tools (use semgrep/codeql directly, then import)
  • Building the code graph itself (use the trailmark skill)
  • Generating diagrams (use the diagramming-code skill after augmenting)

Rationalizations to Reject

RationalizationWhy It's WrongRequired Action
"The user only asked about SARIF, skip pre-analysis"Without pre-analysis, you can't cross-reference findings with blast radius or taintAlways run engine.preanalysis() before augmenting
"Unmatched findings don't matter"Unmatched findings may indicate parsing gaps or out-of-scope filesReport unmatched count and investigate if high
"One severity subgraph is enough"Different severities need different triage workflowsQuery all severity subgraphs, not just error
"SARIF results speak for themselves"Findings without graph context lack blast radius and taint reachabilityCross-reference with pre-analysis subgraphs
"weAudit and SARIF overlap, pick one"Human auditors and tools find different thingsImport both when available
"Tool isn't installed, I'll do it manually"Manual analysis misses what tooling catchesInstall trailmark first

Installation

MANDATORY: If uv run trailmark fails, install trailmark first:

bash
uv tool install trailmark# Python snippets: uv run --with trailmark python -   (a tool env is not importable)

Version Gate

SARIF and weAudit augmentation are v0.2-safe. Binary graph augmentation is Trailmark 0.4.0+ only. Before calling engine.augment_binary(), check:

python
if not hasattr(engine, "augment_binary"):    raise RuntimeError("Binary augmentation requires Trailmark >= 0.4.0")

On Trailmark 0.5.0+, known links between source functions and imported binary or external endpoints can also be declared once in .trailmark/links.toml (see the main trailmark skill's Repository Links section) instead of being re-derived per session. Declared external endpoints materialize as proxy.external:<symbol> nodes on every parse.

Quick Start

CLI

bash
# Augment with SARIFuv run trailmark augment {targetDir} --sarif results.sarif
# Augment with weAudituv run trailmark augment {targetDir} --weaudit .vscode/alice.weaudit
# Both at once, output JSONuv run trailmark augment {targetDir} \    --sarif results.sarif \    --weaudit .vscode/alice.weaudit \    --json

Binary graph augmentation is programmatic in Trailmark 0.4.0+; do not invent a CLI flag if trailmark augment --help does not show one.

Programmatic API

python
from trailmark.query.api import QueryEngine
engine = QueryEngine.from_directory("{targetDir}", language="auto")
# Run pre-analysis first for cross-referencingengine.preanalysis()
# Augment with SARIFresult = engine.augment_sarif("results.sarif")# result: {matched_findings: 12, unmatched_findings: 3, subgraphs_created: [...]}
# Augment with weAuditresult = engine.augment_weaudit(".vscode/alice.weaudit")
# Augment with an external binary graph export (v0.4+)if hasattr(engine, "augment_binary"):    result = engine.augment_binary("binary_graph.json")
# Query findingsengine.findings()                                       # All findingsengine.subgraph("sarif:error")                          # High-severity SARIFengine.subgraph("weaudit:high")                         # High-severity weAuditengine.subgraph("sarif:semgrep")                        # By tool nameengine.annotations_of("function_name")                  # Per-node lookup

If auto-detection is wrong for the target, rerun with an explicit language or comma-separated list such as python,rust.

Workflow

Augmentation Progress:- [ ] Step 1: Build graph and run pre-analysis- [ ] Step 2: Locate SARIF/weAudit/binary graph files- [ ] Step 3: Run augmentation- [ ] Step 4: Inspect results and subgraphs- [ ] Step 5: Cross-reference with pre-analysis

Step 1: Build the graph and run pre-analysis for blast radius and taint context:

python
engine = QueryEngine.from_directory("{targetDir}", language="auto")engine.preanalysis()

If auto-detection is wrong for the target, rerun with an explicit language or comma-separated list such as python,rust.

Step 2: Locate input files:

  • SARIF: Usually output by tools like semgrep --sarif -o results.sarif or codeql database analyze --format=sarif-latest
  • weAudit: Stored in .vscode/<username>.weaudit within the workspace
  • Binary graph (v0.4+): External JSON with artifact, functions, and calls fields. Trailmark imports this graph; it does not disassemble binaries itself.

Step 3: Run augmentation via engine.augment_sarif() or engine.augment_weaudit(). For binary graphs, run engine.augment_binary() only after the Version Gate succeeds. Check unmatched_findings in SARIF and weAudit results — these are findings whose file/line locations didn't overlap any parsed code unit.

Step 4: Query findings and subgraphs. Use engine.findings() to list all annotated nodes. Use engine.subgraph_names() to see available subgraphs.

Step 5: Cross-reference with pre-analysis data to prioritize:

  • Findings on tainted nodes: overlap sarif:error with tainted subgraph
  • Findings on high blast radius nodes: overlap with high_blast_radius
  • Findings on privilege boundaries: overlap with privilege_boundary

For one candidate finding that needs a reachability verdict or PoC handoff, continue with trailmark-finding-triage and use the augmented node as the bound candidate.

Annotation Format

Findings are stored as standard Trailmark annotations:

  • Kind: finding (tool-generated) or audit_note (human notes)
  • Source: sarif:<tool_name> or weaudit:<author>
  • Description: Compact single-line: [SEVERITY] rule-id: message (tool)

Subgraphs Created

SubgraphContents
sarif:errorNodes with SARIF error-level findings
sarif:warningNodes with SARIF warning-level findings
sarif:noteNodes with SARIF note-level findings
sarif:<tool>Nodes flagged by a specific tool
weaudit:highNodes with high-severity weAudit findings
weaudit:mediumNodes with medium-severity weAudit findings
weaudit:lowNodes with low-severity weAudit findings
weaudit:findingsAll weAudit findings (entryType=0)
weaudit:notesAll weAudit notes (entryType=1)
binary:<artifact>Binary function nodes imported from a v0.4+ binary graph

How Matching Works

Findings are matched to graph nodes by file path and line range overlap:

  1. Finding file path is normalized relative to the graph's root_path
  2. Nodes whose location.file_path matches AND whose line range overlaps are selected
  3. The tightest match (smallest span) is preferred
  4. If a finding's location doesn't overlap any node, it counts as unmatched

SARIF paths may be relative, absolute, or file:// URIs — all are handled. weAudit uses 0-indexed lines which are converted to 1-indexed automatically.

Binary graph imports create origin=binary function nodes, origin=proxy external proxy nodes for unresolved binary calls, and inferred corresponds_to edges when a binary function maps back to a source node. The expected JSON shape is intentionally small:

json
{  "artifact": {"name": "libexample", "architecture": "x86_64", "sha256": "..."},  "functions": [    {"symbol": "parse_packet", "address": "0x401000",     "source": {"file": "src/parser.c", "line": 42}}  ],  "calls": [    {"source": "parse_packet", "target": "malloc", "confidence": "inferred"}  ]}

Supporting Documentation

  • references/formats.md [blocked] — SARIF 2.1.0 and weAudit file format field reference

來源與署名

來源:trailofbits/skills位於plugins/trailmark/skills/audit-augmentation提交82fe822

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架