Vector Forge
Uses mutation testing to systematically identify gaps in test vector coverage, then generates new test vectors that close those gaps. Measures effectiveness by comparing mutation kill rates before and after.
When to Use
- Generating test vectors for cryptographic algorithms or protocols
- Evaluating how well existing test vectors cover an implementation
- Finding implementation code paths that no test vector exercises
- Creating Wycheproof-style cross-implementation test vectors
- Measuring the concrete coverage value of a test vector suite
When NOT to Use
- No implementations exist yet (need code to mutate)
- Single trivial implementation with no edge cases
- Testing application logic rather than algorithm implementations
- The algorithm has no public test vectors to compare against
Prerequisites
- trailmark installed — if
uv run trailmarkfails, run:
Python snippets: uv run --with trailmark python - (a tool env is not importable)
Phase 1: Discovery → Find implementations to test ↓ Phase 2: Harness → Write/adapt test vector harness for each impl ↓ Phase 3: Baseline → Run mutation testing with existing vectors ↓ Phase 4: Escape Analysis → Classify escaped mutants by code path ↓ Phase 5: Vector Gen → Create test vectors targeting escapes ↓ Phase 6: Validation → Re-run mutation testing, compare before/after ↓ Output: Coverage Report + New Test Vectors
Handling Existing Vectors
If the implementation already has test vectors:
- Run mutation testing with ONLY the existing vectors (baseline)
- Run mutation testing with ONLY your new vectors
- Run mutation testing with BOTH combined
- The delta between (1) and (3) shows the new vectors' value
Phase 3: Baseline
Run mutation testing with existing test vectors only.
Framework Selection
See references/mutation-frameworks.md [blocked] for language-specific setup.
Parallelism
Always use parallel execution for large codebases:
cargo mutants -j 8(Rust, 8 parallel workers)gremlins unleash --timeout-coefficient 3(Go, increase timeouts)mutmut run --runner "pytest -x -q"(Python, fail-fast)
Recording Baseline Results
Capture these metrics per implementation:
Save the full mutation log for Phase 4 analysis.
Phase 4: Escape Analysis (Graph-Informed Triage)
Classify each escaped (survived + not covered) mutant using the Trailmark call graph for reachability and blast radius analysis.
This phase MUST use the genotoxic skill's triage methodology. The call graph transforms mutation results from a flat list of survived mutants into an actionable, prioritized set of vector targets.
Step 1: Build the Call Graph
Build a Trailmark code graph for each implementation before triaging mutations:
The graph provides:
- Caller chains — trace from public API entry points to mutated functions to determine reachability
- Cyclomatic complexity — prioritize high-CC functions
- Blast radius — functions with many callers have wider impact if their mutations survive
Step 2: Filter to Relevant Code
Mutation frameworks test the entire package. Filter results to only the files/functions that test vectors should exercise:
Step 3: Graph-Informed Classification
For each escaped mutant, map it to its containing function in the call graph and apply the genotoxic triage criteria:
Step 4: Identify Cross-Package Test Gaps
Critical pitfall: Mutation frameworks often only run tests within the same package as the mutation. For Go (gremlins) and Rust (cargo-mutants), this means:
- A mutation in
hash_to_curve/g2.goonly runs tests in thehash_to_curvepackage, NOT tests in the parentbls12381package that imports it - Functions that are fully exercised by cross-package tests will appear as NOT COVERED — these are false positives
- To confirm: check if the mutated function is called from a test in a different package that wouldn't be run
To resolve cross-package gaps:
- Add a thin test in the sub-package that calls through the same code path as the cross-package test
- Or run gremlins with
--test-pkg ./...(if supported) - Or document as a framework limitation in the report
Step 5: Prioritize by Security Impact
Using the call graph, rank surviving mutants by impact:
Step 6: Group by Vector Strategy
Group escaped mutants by the code path they represent and the type of test vector needed:
Each group becomes a target for new test vectors in Phase 5.
Phase 5: Vector Generation
For each escaped code path group, design test vectors that force execution through that path.
Vector Design Patterns
Single-Fault Negative Vectors
Each negative vector should have exactly one defect with everything else valid — this isolates which validation check is being tested. See references/vector-patterns.md [blocked] for per-flag construction examples.
Fault Simulation (Limb-Width Reimplementation)
When mutation testing only applies local operator swaps, deeper architectural bugs (carry propagation, reduction overflow) go untested. To close this gap, reimplement the target algorithm at reduced limb widths (8, 16, 25, 32 bits) and deliberately inject faults — then generate vectors that catch them.
See references/fault-simulation.md [blocked] for the full methodology: limb-width selection, fault injection catalog, vector extraction, and validation workflow.
Cross-Implementation Verification
Every new test vector MUST be verified against at least two independent implementations before being added to the suite:
- Generate the vector using implementation A
- Verify with implementation B (different codebase, ideally different language)
- If B disagrees, investigate — one implementation has a bug
Vector Format
Use Wycheproof JSON format (algorithm, testGroups[].tests[]
with tcId, comment, result, flags). See
references/vector-patterns.md [blocked]
for the full schema.
Wycheproof contributions: Use Wycheproof's vectorgen tool rather
than formatting vector files directly. Supply the generated changes as an
envelope. The vectorgen tool can add, update, or replace vectors while
handling tcId assignment, test counts, canonical formatting, and schema
validation. Go-based generators can avoid the vectorgen CLI tool and instead
call the programmatic github.com/c2sp/wycheproof/vectorgen API.
See references/lessons-learned.md [blocked] §14 and the upstream vectorgen guide for the current workflow and commands.
Phase 6: Validation
Re-run mutation testing with the new test vectors included.
Tip: Use per-file mutation testing for fast iteration during vector development (see references/lessons-learned.md [blocked] §12). Only run full-crate tests for the final comparison.
Before/After Comparison
Success Criteria
Vectors have both retroactive value (killing mutants in existing code) and proactive value (catching bugs in future implementations). Generate both kinds — boundary-condition vectors may not improve kill rates in mature libraries but will catch bugs in new implementations. See references/lessons-learned.md [blocked] §13.
Retroactive (measurable): previously survived/uncovered mutants become killed, no regressions.
If kill rates don't change: the implementation's own tests likely already cover those paths. The vectors still add cross-implementation verification value. Document which case applies.
Output Format
Write VECTOR_FORGE_REPORT.md covering: target algorithm,
implementations tested, baseline results, escape analysis,
new vectors generated, after results, before/after delta, and
conclusions. See
references/report-template.md [blocked]
for the full template.
Quality Checklist
Before delivering:
- At least one pure implementation mutation-tested (not just FFI wrappers)
- Baseline run completed with existing vectors
- Trailmark call graph built for each implementation
- All escaped mutants triaged using graph-informed classification
- Cross-package false positives identified and documented
- Security-critical mutations (ct_eq, validation, auth) prioritized as P0/P1
- Fault simulation and mutation-derived vectors cross-verified against 2+ implementations
- After run completed with new vectors included
- Before/after delta computed and explained
- Report written to
VECTOR_FORGE_REPORT.md - New test vectors saved in standard format (Wycheproof JSON)
Integration
Supporting Documentation
- references/mutation-frameworks.md [blocked] - Language-specific mutation testing framework setup
- references/vector-patterns.md [blocked] - Common test vector patterns for cryptographic primitives
- references/fault-simulation.md [blocked] - Limb-width reimplementation for carry, reduction, and overflow faults
- references/report-template.md [blocked] - Full markdown template for the Vector Forge report
- references/lessons-learned.md [blocked] - BLS12-381 case study: FFI kill rates, timeout masking, cross-package false positives, bitwise mutation gaps, and security-critical priorities

