Writing Fuzzing Harnesses
A fuzzing harness is the entrypoint function that receives random data from the fuzzer and routes it to your system under test (SUT). The quality of your harness directly determines which code paths get exercised and whether critical bugs are found. A poorly written harness can miss entire subsystems or produce non-reproducible crashes.
Overview
The harness is the bridge between the fuzzer's random byte generation and your application's API. It must parse raw bytes into meaningful inputs, call target functions, and handle edge cases gracefully. The most important part of any fuzzing setup is the harness—if written poorly, critical parts of your application may not be covered.
Key Concepts
When to Apply
Apply this technique when:
- Creating a new fuzz target for the first time
- Fuzz campaign has low code coverage or isn't finding bugs
- Crashes found during fuzzing are not reproducible
- Target API requires complex or structured inputs
- Multiple related functions should be tested together
Skip this technique when:
- Using existing well-tested harnesses from your project
- Tool provides automatic harness generation that meets your needs
- Target already has comprehensive fuzzing infrastructure
Quick Reference
Step-by-Step
Step 1: Identify Entry Points
Find functions in your codebase that:
- Accept external input (parsers, validators, protocol handlers)
- Parse complex data formats (JSON, XML, binary protocols)
- Perform security-critical operations (authentication, cryptography)
- Have high cyclomatic complexity or many branches
Good targets are typically:
- Protocol parsers
- File format parsers
- Serialization/deserialization functions
- Input validation routines
Step 2: Write Minimal Harness
Start with the simplest possible harness that calls your target function:
C/C++:
Rust:
Step 3: Add Input Validation
Reject inputs that are too small or too large to be meaningful:
Rationale: The fuzzer generates random inputs of all sizes. Your harness must handle empty, tiny, huge, or malformed inputs without causing unexpected issues in the harness itself (crashes in the SUT are fine—that's what we're looking for).
Step 4: Structure the Input
For APIs that require typed data (integers, strings, etc.), use casting or helpers like FuzzedDataProvider:
Simple casting:
Using FuzzedDataProvider:
Step 5: Test and Iterate
Run the fuzzer and monitor:
- Code coverage (are all interesting paths reached?)
- Executions per second (is it fast enough?)
- Crash reproducibility (can you reproduce crashes with saved inputs?)
Iterate on the harness to improve these metrics.
Common Patterns
Pattern: Beyond Byte Arrays—Casting to Integers
Use Case: When target expects primitive types like integers or floats
Implementation:
Rust equivalent:
Why it works: Any 8-byte input is valid. The fuzzer learns that inputs must be exactly 8 bytes, and every bit flip produces a new, potentially interesting input.
Pattern: FuzzedDataProvider for Complex Inputs
Use Case: When target requires multiple strings, integers, or variable-length data
Implementation:
Why it helps: FuzzedDataProvider handles the complexity of extracting structured data from a byte stream. It's particularly useful for APIs that need multiple parameters of different types.
Pattern: Interleaved Fuzzing
Use Case: When multiple related operations should be tested in a single harness
Implementation:
Advantages:
- Faster to write one harness than multiple individual harnesses
- Single shared corpus means interesting inputs for one operation may be interesting for others
- Can discover bugs in interactions between operations
When to use:
- Operations share similar input types
- Operations are logically related (e.g., arithmetic operations, CRUD operations)
- Single corpus makes sense across all operations
Pattern: Structure-Aware Fuzzing with Arbitrary (Rust)
Use Case: When fuzzing Rust code that uses custom structs
Implementation:
Harness with arbitrary:
Add to Cargo.toml:
Why it helps: The arbitrary crate automatically handles deserialization of raw bytes into your Rust structs, reducing boilerplate and ensuring valid struct construction.
Limitation: The arbitrary crate doesn't offer reverse serialization, so you can't manually construct byte arrays that map to specific structs. This works best when starting from an empty corpus (fine for libFuzzer, problematic for AFL++).
Advanced Usage
Tips and Tricks
Structure-Aware Fuzzing with Protocol Buffers
For highly structured input formats, consider using Protocol Buffers as an intermediate format with custom mutators:
This approach is more setup but prevents the fuzzer from wasting time on unparseable inputs. See structure-aware fuzzing documentation for details.
Handling Non-Determinism
Problem: Random values or timing dependencies cause non-reproducible crashes.
Solutions:
- Replace
rand()with deterministic PRNG seeded from fuzzer input: - Mock system calls that return time, PIDs, or random data
- Avoid reading from
/dev/randomor/dev/urandom
Resetting Global State
If your SUT uses global state (singletons, static variables), reset it between iterations:
Rationale: Global state can cause crashes after N iterations rather than on a specific input, making bugs non-reproducible.
Practical Harness Rules
Follow these rules to ensure effective fuzzing harnesses:
Note: These guidelines apply not just to harness code, but to the entire SUT. If the SUT violates these rules, consider patching it (see the fuzzing obstacles technique).
Anti-Patterns
Tool-Specific Guidance
libFuzzer
Harness signature:
Compilation:
Integration tips:
- Use
FuzzedDataProvider.hfor structured input extraction - Compile with
-fsanitize=fuzzerto link the fuzzing runtime - Add sanitizers (
-fsanitize=address,undefined) to detect more bugs - Use
-gfor better stack traces when crashes occur - libFuzzer can start with empty corpus—no seed inputs required
Running:
Resources:
AFL++
AFL++ supports multiple harness styles. For best performance, use persistent mode:
Persistent mode harness:
Compilation:
Integration tips:
- Use persistent mode (
__AFL_LOOP) for 10-100x speedup - Consider deferred initialization (
__AFL_INIT()) to skip setup overhead - AFL++ requires at least one seed input in the corpus directory
- Use
AFL_USE_ASAN=1orAFL_USE_UBSAN=1for sanitizer builds
Running:
cargo-fuzz (Rust)
Harness signature:
With structured input (arbitrary crate):
Creating harness:
Integration tips:
- Use
arbitrarycrate for automatic struct deserialization - cargo-fuzz wraps libFuzzer, so all libFuzzer features work
- Compile with sanitizers automatically via cargo-fuzz
- Harnesses go in
fuzz/fuzz_targets/directory
Running:
Resources:
go-fuzz
Harness signature:
Building:
Integration tips:
- Return 1 for inputs that add coverage (optional—fuzzer can detect automatically)
- Return -1 for invalid inputs to deprioritize similar mutations
- go-fuzz handles persistence automatically
Running:
Troubleshooting
Related Skills
Tools That Use This Technique
Related Techniques
Resources
Key External Resources
Split Inputs in libFuzzer - Google Fuzzing Docs Explains techniques for handling multiple input parameters in a single fuzzing harness, including use of magic separators and FuzzedDataProvider.
Structure-Aware Fuzzing with Protocol Buffers Advanced technique using protobuf as intermediate format with custom mutators to ensure fuzzer mutates message contents rather than format encoding.
libFuzzer Documentation Official LLVM documentation covering harness requirements, best practices, and advanced features.
cargo-fuzz Book Comprehensive guide to writing Rust fuzzing harnesses with cargo-fuzz and the arbitrary crate.
Video Resources
- Effective File Format Fuzzing - Conference talk on writing harnesses for file format parsers
- Modern Fuzzing of C/C++ Projects - Tutorial covering harness design patterns

