๐ง๐ต๐ฒ ๐ญ.๐ต๐ฒ% ๐๐ฎ๐ฝ - ๐ง๐ต๐ฒ ๐๐ถ๐ป๐ฒ ๐ง๐ต๐ฎ๐ ๐๐ฒ๐ณ๐ถ๐ป๐ฒ๐ ๐๐ต๐ฒ ๐๐ฟ๐ผ๐ป๐๐ถ๐ฒ๐ฟ
LLMs look impressive until you ask them to solve something real โ the moment a problem requires reasoning instead of patternโmatching, the 1.96% ceiling shows up.
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
See more: (https://arxiv.org/abs/2310.06770v3). Not my article.
ArXiv finds: โState-of-the-art proprietary models โ and even the fine-tuned SWE-Llama โ can resolve only the simplest issues. Claude 2 tops out at 1.96%.โ
But it is the line that defines the frontier.
๐ญ.๐ต๐ฒ%.
Not 20.
Not 10.
Not 5.
๐ข๐ป๐ฒ ๐ฝ๐ผ๐ถ๐ป๐ ๐ป๐ถ๐ป๐ฒ ๐๐ถ๐
.
That number tells you exactly where the capability gap is: models handle simple issues, and then collapse the moment the problem requires multi-file reasoning, invariant awareness, or any real systems-layer understanding.
That gap โ the space between โsimpleโ and โmediumโ โ is where Iโm building.
Iโm approaching SWE bench with a hybrid persona model designed for that tier, finding medium issues in a #GitHub or similar repository:
โข๐๐๐ฎ๐น๐๐ฎ๐๐ผ๐ฟ ๐ฝ๐ฒ๐ฟ๐๐ผ๐ป๐ฎ โ reproducibility, deterministic patches, invariant preservation, regression avoidance.
โข๐ฆ๐๐๐๐ฒ๐บ๐ ๐ฝ๐ฒ๐ฟ๐๐ผ๐ป๐ฎ โ dependency awareness, side-effect mapping, multi-file reasoning.
โข๐๐ด๐ฒ๐ป๐๐ถ๐ฐ ๐ฝ๐ฒ๐ฟ๐๐ผ๐ป๐ฎ โ multiple patch strategies, comparative reasoning, failure-mode tracing.
๐ง๐ต๐ฒ ๐ด๐ผ๐ฎ๐น ๐ถ๐ ๐๐ผ ๐ผ๐ฝ๐ฒ๐ฟ๐ฎ๐๐ฒ ๐ถ๐ป ๐๐ต๐ฒ ๐๐ถ๐ฒ๐ฟ ๐๐ต๐ฒ๐ฟ๐ฒ ๐๐ต๐ฒ๐ ๐ฐ๐๐ฟ๐ฟ๐ฒ๐ป๐๐น๐ ๐ณ๐ฎ๐ถ๐น โ ๐บ๐ฒ๐ฑ๐ถ๐๐บ ๐ฐ๐ผ๐บ๐ฝ๐น๐ฒ๐ ๐ถ๐๐ ๐ฆ๐ช๐.
This is Part 1.
๐ก๐ฒ๐
๐: ๐ช๐ต๐ ๐ ๐ฒ๐ฑ๐ถ๐๐บ ๐๐๐๐๐ฒ๐ ๐ ๐ฎ๐๐๐ฒ๐ฟ.
Top comments (0)