DEV Community

Cover image for Stop Asking AI to Fix the Bug. Ask It to Prove the Cause.
Robert Adamson
Robert Adamson

Posted on

Stop Asking AI to Fix the Bug. Ask It to Prove the Cause.

You paste an error into your AI assistant.

“Fix this.”

It changes a condition. Adds a null check. Wraps something in try/catch.

The error disappears.

But do you know why it happened—or whether the fix actually solved it?

That is the difference between making a symptom disappear and understanding a bug.

A better starting prompt is:

“Don’t change the code yet. Help me reproduce the failure and find evidence for its cause.”

Here’s how to turn that into a practical debugging workflow.

A convincing explanation is still a guess

Imagine your app occasionally shows older search results after a user types a new query.

You ask AI to fix it.

It suggests adding a debounce.

That sounds reasonable. Search inputs often use debouncing. Fewer requests might even make the problem appear less frequently.

But consider this sequence:

User searches for "react"
Request A starts

User searches for "react testing"
Request B starts

Request B finishes → correct results appear
Request A finishes → old results overwrite them
Enter fullscreen mode Exit fullscreen mode

The problem is responses arriving out of order.

Debouncing reduces how often requests start. It does not guarantee that the latest request finishes last.

A plausible fix can leave the actual bug intact.

1. Describe the failure before sharing the code

“Search is broken” gives the assistant too much room to guess.

Start with what you observed:

Expected:
Results should match the current search query.

Actual:
Sometimes results from an earlier query replace the latest results.

Trigger:
Type one query, then quickly change it.

Frequency:
Intermittent. Easier to reproduce with a slow connection.
Enter fullscreen mode Exit fullscreen mode

Then provide the relevant code and ask:

Do not edit anything yet.

Using the observed behavior and this code:
1. List up to three possible causes.
2. Separate observed facts from assumptions.
3. Explain what evidence would support or reject each cause.
4. Suggest the smallest reproduction.
Enter fullscreen mode Exit fullscreen mode

This gives you hypotheses to investigate instead of a patch to trust.

2. Make the failure repeatable

Intermittent bugs are difficult because you cannot reliably compare “before” and “after.”

For the search example, force the earlier request to finish later:

const delay = (ms) =>
  new Promise((resolve) => setTimeout(resolve, ms));

async function fakeSearch(query) {
  await delay(query === "react" ? 800 : 100);

  return [`Result for ${query}`];
}
Enter fullscreen mode Exit fullscreen mode

Start a search for "react", then shortly afterward start one for "react testing".

If your UI applies every response as it arrives, the older result wins.

You now have a controlled reproduction.

A useful prompt:

Help me create a minimal reproduction.

Control timing or inputs where necessary.
Keep the suspected failure mechanism intact.
Explain what this reproduction demonstrates
and what it does not establish.
Enter fullscreen mode Exit fullscreen mode

A reproduction shows that a mechanism can cause the failure. To connect it to your real incident, you still need evidence from the actual application.

3. Collect evidence that distinguishes the causes

Random logging creates noise.

Useful logging answers a specific question.

For this bug, record the query and request ID when each request starts, finishes, and updates the UI:

START  id=41 query="react"
START  id=42 query="react testing"
FINISH id=42 query="react testing"
APPLY  id=42 query="react testing"
FINISH id=41 query="react"
APPLY  id=41 query="react"
Enter fullscreen mode Exit fullscreen mode

That trace supports a specific explanation: an older response updates the UI after the newer response.

Ask the assistant:

What is the smallest amount of instrumentation
needed to distinguish these hypotheses?

For each log or measurement, explain:
- What question it answers
- Which hypothesis it could reject

Avoid logging secrets or personal data.
Enter fullscreen mode Exit fullscreen mode

The goal is to narrow uncertainty, not fill your console.

4. Ask for the causal chain

Before accepting a patch, ask the assistant to connect the evidence to the failure.

Explain the causal chain using the reproduction,
logs, and relevant code.

Identify:
1. The triggering event
2. The incorrect behavior
3. The missing safeguard
4. How that produces the visible failure

State anything that remains uncertain.
Enter fullscreen mode Exit fullscreen mode

For our example:

  1. The user starts two searches.
  2. Both requests remain active.
  3. The earlier request finishes last.
  4. The code applies its result without checking whether it is still current.
  5. The UI displays outdated results.

That explanation gives the fix a clear target.

“The API is slow” is not enough. A slow API exposes the problem; the missing response-order safeguard allows it.

5. Fix the mechanism, then challenge the fix

In a simplified search controller, a request counter can prevent outdated responses from updating results:

let latestRequestId = 0;

async function search(query) {
  const requestId = ++latestRequestId;
  const results = await fakeSearch(query);

  if (requestId !== latestRequestId) {
    return;
  }

  renderResults(results);
}
Enter fullscreen mode Exit fullscreen mode

Only the latest search can apply its results.

In a real app, keep this counter within the relevant component or controller. Also check loading states, errors, clearing the input, and component cleanup.

Now give AI a different job:

Try to find a counterexample to this fix.

Check:
- An older request finishing last
- The latest request failing
- The input being cleared during a request
- The component being destroyed during a request

Explain which cases the patch handles
and which still need work.
Enter fullscreen mode Exit fullscreen mode

This is where AI becomes especially useful: helping you challenge a solution instead of merely generating one.

6. Write a test that exposes the original bug

A test that passes after your change is useful.

A test that fails before the change and passes afterward provides stronger evidence that you addressed the reproduced failure.

For this example, the test should:

  1. Start the first search.
  2. Start the second search.
  3. Resolve the second request first.
  4. Resolve the first request afterward.
  5. Assert that the second search’s results remain visible.

Use controlled promises rather than arbitrary delays in the test.

And define the expected behavior yourself:

Write a regression test for this requirement:

Once a newer search starts, an older response
must not replace its results.

Control request completion order explicitly.
The test must fail against the original implementation
and pass against the proposed fix.
Enter fullscreen mode Exit fullscreen mode

The requirement should drive the test. The implementation should not define what “correct” means.

The prompt worth saving

Use this the next time a bug sends you into a loop of AI-generated patches:

Help me investigate this bug. Do not modify code yet.

Expected behavior:
[describe]

Actual behavior:
[describe]

Reproduction steps:
[list]

Relevant code and evidence:
[paste]

Please:
1. Separate facts from assumptions.
2. Rank up to three hypotheses and explain why.
3. Propose the smallest experiment to distinguish them.
4. Identify any missing evidence.
5. Explain the causal chain once evidence supports it.
6. Then suggest the smallest fix and a regression test.

Do not claim the cause is confirmed without evidence.
If you cannot run an experiment, say so.
Enter fullscreen mode Exit fullscreen mode

Debugging is reducing uncertainty

You do not need a lengthy investigation for every typo or obvious syntax error.

But when a bug is intermittent, keeps returning, or affects important behavior, another quick patch can cost more than a small experiment.

Before accepting the next AI fix, ask:

  • Can I reproduce the failure?
  • What evidence supports this explanation?
  • Does the patch address that mechanism?
  • Would the regression test catch the original bug?

Ask AI to help you earn confidence in the fix.

What’s a bug where the first “obvious fix” turned out to be wrong? Share the symptom—and what actually caused it—in the comments.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.