Skip to content

Measure fuzzing coverage beside test coverage - #652

Merged
brianegge merged 1 commit into
morganstanley:mainfrom
brianegge:fuzz/coverage-compare
Oct 3, 2026
Merged

brianegge merged 1 commit into
morganstanley:mainfrom
brianegge:fuzz/coverage-compare

Conversation

@brianegge

Copy link
Copy Markdown
Contributor

Adds a way to measure which lines of hobbes the fuzz corpora reach, and to compare that line by line with what hobbes-test reaches.

What's added

  • fuzz/coverage.sh <build-dir> [corpora] [out]: runs hobbes-test, replays each harness's corpus once (no mutation), and merges the profiles. It writes one lcov per profile plus a Markdown report.
  • fuzz/coverage-compare.py: sorts every instrumented line into four bins: both, tests only, fuzzing only, or neither. It reports per directory and per file, and counts the lines each harness reaches that nothing else does.
  • FUZZ_STANDALONE CMake option: builds the harnesses as the existing standalone replay runners and doesn't apply -fsanitize=fuzzer-no-link to the rest of the build. Without it, a Clang BUILD_FUZZERS tree can't link hobbes-test, so the tests and the harnesses could never share one instrumented libhobbes. The default libFuzzer configuration is unchanged.
  • standalone_main.C fix: it declared LLVMFuzzerInitialize as a weak undefined symbol, which Apple's linker refuses. On macOS, fregion-reader, type-decode and hog-session therefore never linked as standalone runners. It's now a weak default definition that a harness's own definition overrides on both ELF and Mach-O.
  • New fuzz/README.md section: covers the setup and how to read the numbers.

First results

These come from corpora grown locally for 10 minutes per harness (UBSan libFuzzer, macOS arm64), starting from the shipped seeds. OSS-Fuzz's corpus backups aren't public, so these numbers are a floor for fuzzing.

Lines (of 42,761) Share
Tests 26,933 63.0%
Fuzzing 18,536 43.3%
Either 27,369 64.0%
Tests only 8,833 20.7%
Fuzzing only 436 1.0%
  • Fuzzing rarely reaches lines the tests miss. Almost everything it covers is code the tests also run, but with inputs no test supplies. The 1% it adds is mostly error handling in the parser, lexer and type checker.
  • Overlap is high in the compiler front end and parser: lib/hobbes/lang, read and parse all have similar shares from tests and from fuzzing.
  • The biggest gaps are untrusted-input surfaces that no harness reaches: lib/hobbes/ipc (net.C, prepl.C) and lib/hobbes/db have 8–15% fuzzing coverage, against about 55–69% from the tests. These look like the next harnesses worth writing.
  • Effect of Fuzz code generation as far as LLVM IR, checked by LLVM's verifier #650: with Fuzz code generation as far as LLVM IR, checked by LLVM's verifier #650's code-generation step in the typecheck harness, fuzzing coverage rises to 44.0% (about 300 more lines) on the same corpora.

These numbers use Clang source-based coverage, which counts lines differently from the gcov-based CI coverage report (72.3% on main). Only compare them with each other.

Testing

  • On this branch (based on main): configured the coverage build, built all targets, and ran fuzz/coverage.sh against the grown corpora. Every replay batch finished, and hobbes-test passed under coverage.
  • A normal BUILD_FUZZERS=ON configuration still selects libFuzzer and applies fuzzer-no-link.

🤖 Generated with Claude Code

…gge.md)

Nothing measured what the fuzz corpora reach, so there was no way to say
where fuzzing adds to the tests and where it leaves untrusted-input code
untouched. fuzz/coverage.sh runs hobbes-test and replays each harness's
corpus once under Clang source-based coverage. coverage-compare.py then
sorts every instrumented line into covered by both, by the tests only, by
the fuzzers only, or by neither, per directory and per file, and counts
the lines each harness reaches that nothing else does.

The comparison needs the harnesses and hobbes-test linked against the same
instrumented libhobbes. With Clang that was impossible, because
BUILD_FUZZERS applies -fsanitize=fuzzer-no-link to the whole build, and
hobbes-test then fails to link. The new FUZZ_STANDALONE option builds the
harnesses as the standalone replay runners already used for compilers
without libFuzzer, and leaves the rest of the build alone.

Building those runners on macOS showed that standalone_main.C declared
LLVMFuzzerInitialize as a weak undefined symbol. Apple's linker refuses
that, so the three harnesses that do not define it (fregion-reader,
type-decode, hog-session) never linked as standalone runners there. It is
now a weak default definition, which a harness's own definition overrides
on both ELF and Mach-O.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@brianegge
brianegge merged commit b7e6ad5 into morganstanley:main Oct 3, 2026
37 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants