Fuzzing works, but it doesn't scale without someone babysitting it. I kept hitting the same wall: writing harnesses for unreached code, reading coverage reports, chasing gaps, triaging crashes. All repetitive work that needed a human in the loop.
I built the Fuzzing Taskflow to hand those tasks to an LLM agent instead. Point it at a GitHub repo, and it identifies entrypoints, writes harnesses, runs AFL++, reads coverage, improves targets, and generates vulnerability reports—all autonomous.
The design principle is clean: the agent makes decisions, the tools handle execution. The agent decides what to fuzz and what coverage gap to chase next. It never calls AFL or clang directly—just composes the pipeline from building blocks like run_afl_for or compile_harness.
Each harness builds twice: an AFL binary for fuzzing and a coverage binary for real reports. The agent iterates on coverage with doubling time budgets (30s → 60s → 120s...), stopping when gains plateau below 1%. It can add seeds, enrich AFL dictionaries with magic constants from the source, or skip cold paths.
This started as an internal experiment, but I'm wondering whether there's a real product here. A friction-free fuzzing service for small C/C++ projects, maybe. Point it at your repo, get security findings without the setup overhead.
Anyone else thinking about AI business ideas in security tooling? The gap between "interesting prototype" and "shippable product" is where I'm stuck.
Top comments (0)