DEV Community

Agentic Drifter
Agentic Drifter

Posted on

๐—ฃ๐—ผ๐˜€๐˜ ๐Ÿฎ โ€” ๐—ช๐—ต๐˜† ๐— ๐—ฒ๐—ฑ๐—ถ๐˜‚๐—บ ๐—œ๐˜€๐˜€๐˜‚๐—ฒ๐˜€ ๐— ๐—ฎ๐˜๐˜๐—ฒ๐—ฟ (finding medium software engineering issues, within a #codebase)

๐—ฆ๐—ฒ๐—ฟ๐—ถ๐—ฒ๐˜€: ๐—ง๐—ต๐—ฒ ๐Ÿญ.๐Ÿต๐Ÿฒ% ๐—š๐—ฎ๐—ฝ - Medium Issues

A model cannot learn medium-tier reasoning from one prompt, one shot. Here is where #HumanintheLoop comes in.

This is the part most people miss.

Medium issues force the model to confront things it cannot shortcut: the underlying intent of the code, the expected behavior, the conditions that must remain true, and the consequences of a change. These are not โ€œharder bugsโ€ โ€“ they are ๐—ฟ๐—ฒ๐—ฎ๐˜€๐—ผ๐—ป๐—ถ๐—ป๐—ด ๐—ฝ๐—ฟ๐—ผ๐—ฏ๐—น๐—ฒ๐—บ๐˜€.

A model cannot learn medium-tier reasoning from one prompt.

It needs a structured loop:
โ€ขIdentify relevant files
โ€ขDescribe current vs. expected behavior
โ€ขList invariants
โ€ขPropose minimal patch
โ€ขApply patch
โ€ขRun tests
โ€ขDiagnose failures
โ€ขRevise patch

This loop is how you train the model to reason โ€“ and itโ€™s why medium issues require Human-in-the-loop oversight.

Next, putting reasoning under real pressure for the SWEโ€‘bench.

Top comments (0)