DEV Community

NET_DARK_BOI
NET_DARK_BOI

Posted on Fully Autonomous

How Would You Prove an AI Security Fix Actually Works?

An assistant proposes a security patch. Its explanation sounds reasonable. The original attack no longer works.

Do you merge it?

I want to evaluate that decision with a small experiment built around intentionally vulnerable exercises. This post describes the experiment design. It does not report completed runs or claim a failure rate for any model.

Three cases, three different failure modes

I would begin with:

  • Object authorization: one user can read another user's invoice.
  • SQL injection: input changes the structure of a query.
  • Path traversal: a download feature can access files outside its permitted directory.

Each exercise needs an explicit contract. Who should have access? Which inputs are legitimate? What should happen when a resource does not exist?

Without those details, the assistant and the reviewer may be solving different problems.

Freeze the tests before asking for a patch

I would prepare three test groups:

Group Question
Original reproduction Does the reported misuse still succeed?
Related cases Does the same weakness survive with different data or conditions?
Legitimate behavior Can intended users still use the feature?

For a SQL query, an ordinary product name containing an apostrophe is a useful legitimate input. Rejecting it may hide a query-construction bug while breaking a valid requirement.

For object authorization, use at least two users and more than one protected object. A hard-coded exception must not pass as a general fix.

For file access, the tests must match the actual path-decoding, normalization and filesystem behavior. A toy string check cannot establish the safety of a production file server.

Record what actually happened

For every run, save:

  • The model identifier, date and relevant settings.
  • The exact prompt, code and requirements supplied.
  • The proposed patch and any human edits.
  • Test results by group, including failures.
  • Whether the assistant saw test feedback and received another attempt.

Keep first-attempt results separate from results after feedback. If comparing models, use the same cases and access to tools, and report the limited scope. A few examples cannot support a broad ranking of model security competence.

Let the reader review before revealing the result

The most interesting part would be showing a patch first and asking: “Approve or request changes?”

Then reveal the test evidence. Include patches that succeed. If the assistant handles every case well, that is the result to publish.

My hypothesis is that reviewing the contract and test coverage will teach more than counting whether a single demonstration stopped working. The experiment should be allowed to challenge that hypothesis too.

I am building Breachloom, which includes code-repair exercises. Those exercises motivate this proposed evaluation; no AI benchmark results are being claimed here.

Which legitimate behavior would you include to catch an over-restrictive “security fix”?

Top comments (0)