LLMs donโt fail at hard problems. They fail at the (medium) ones โ the ones that require ๐ฟ๐ฒ๐ฎ๐๐ผ๐ป๐ถ๐ป๐ด, not patternโmatching. Thatโs the ๐ญ.๐ต๐ฒ% Gap.
This week, I saw it directly.
โโโโโโโโโโโโโโโโโโโโโโโโโโโโ
The Medium Issues I Actually Found.
A multiโfile trace probe across a #GitHub repository: models, sessions, and utils, surfaced five medium issues โ not because the code was broken, but because the system had to ๐ฟ๐ฒ๐ฎ๐๐ผ๐ป.
๐๐ฒ๐ฎ๐ฑ๐ฒ๐ฟ ๐บ๐ฒ๐ฟ๐ด๐ฒ ๐ฑ๐ฟ๐ถ๐ณ๐
X-Testandx-teststored separately โ ๐๐ฒ๐บ๐ฎ๐ป๐๐ถ๐ฐ ๐ฑ๐ฟ๐ถ๐ณ๐.๐จ๐ฅ๐ ๐ป๐ผ๐ฟ๐บ๐ฎ๐น๐ถ๐๐ฎ๐๐ถ๐ผ๐ป ๐ฑ๐ฟ๐ถ๐ณ๐
Unicode left unencoded โ ๐ฐ๐ผ๐ป๐๐๐ฟ๐ฎ๐ถ๐ป๐ ๐ฑ๐ฟ๐ถ๐ณ๐.๐๐ผ๐ผ๐ธ๐ถ๐ฒ ๐ฝ๐ฟ๐ผ๐ฝ๐ฎ๐ด๐ฎ๐๐ถ๐ผ๐ป ๐ฑ๐ฟ๐ถ๐ณ๐
Cookies missing after prepare โ ๐๐๐ฎ๐๐ฒ ๐ฑ๐ฟ๐ถ๐ณ๐.๐ฅ๐ฒ๐ฑ๐ถ๐ฟ๐ฒ๐ฐ๐ ๐๐ผ๐ป๐๐ฒ๐ป๐โ๐ง๐๐ฝ๐ฒ ๐ฑ๐ฟ๐ถ๐ณ๐
POST โ GET keptContent-Typeโ ๐๐ฒ๐บ๐ฎ๐ป๐๐ถ๐ฐ + ๐ฐ๐ผ๐ป๐๐๐ฟ๐ฎ๐ถ๐ป๐ ๐ฑ๐ฟ๐ถ๐ณ๐.
-๐ฅ๐ฒ๐ฑ๐ถ๐ฟ๐ฒ๐ฐ๐ ๐ต๐ฒ๐ฎ๐ฑ๐ฒ๐ฟ ๐ฟ๐ฒ๐๐๐ฒ ๐ฑ๐ฟ๐ถ๐ณ๐
POST headers leaked into GET โ ๐ฝ๐ฟ๐ฒ๐บ๐ถ๐๐ฒ ๐ฑ๐ฟ๐ถ๐ณ๐.
None crashed the system. All fractured behavior. Thatโs what ๐บ๐ฒ๐ฑ๐ถ๐๐บ ๐ถ๐๐๐๐ฒ๐ are.
โโโโโโโโโโโโโโโโโโโโโโโโโโโโ
Why They Only Appear Under Reasoning Pressure
Medium issues emerge when you force a model through a ๐ฟ๐ฒ๐ฎ๐๐ผ๐ป๐ถ๐ป๐ด ๐น๐ผ๐ผ๐ฝ โ not just โanswering,โ but:
Interpret โ Extract premises โ Build chain โ Critique โ Revise.
Pattern models break early. Reasoning models break later.
Medium issues live between the steps โ where the ๐ญ.๐ต๐ฒ% becomes visible.
โโโโโโโโโโโโโโโโโโโโโโโโโโโโ
The Breakthrough: A ๐ฟ๐ฒ๐ฎ๐๐ผ๐ป๐ถ๐ป๐ด-๐ณ๐ถ๐ฟ๐๐ ๐ฐ๐ฟ๐ถ๐๐ถ๐ฐ
I didnโt need a critic to find medium issues.
I needed it to ๐ฒ๐
๐ฝ๐น๐ฎ๐ถ๐ป them.
It caught:
โข ๐ฝ๐ฟ๐ฒ๐บ๐ถ๐๐ฒ ๐ฑ๐ฟ๐ถ๐ณ๐
โข ๐๐ฒ๐บ๐ฎ๐ป๐๐ถ๐ฐ ๐ฑ๐ฟ๐ถ๐ณ๐
โข ๐ฐ๐ผ๐ป๐๐๐ฟ๐ฎ๐ถ๐ป๐ ๐ฑ๐ฟ๐ถ๐ณ๐
Example: http, port 80, length 12.
Pattern model: 12 chars โ https โ encrypted.
Reasoning critic: HTTP + port 80 contradict encryption; length irrelevant; HTTPS never introduced.
Thatโs the ๐ญ.๐ต๐ฒ% Gap.
โโโโโโโโโโโโโโโโโโโโโโโโโโโโ
The Takeaway
To see the ๐ญ.๐ต๐ฒ% gap, stop testing patterns.
Force the model to reason.
Then critique the reasoning.
Thatโs where fractures appear.
Thatโs where drift becomes visible.
Thatโs where medium issues live.
Next, soon: Iโm breaking down the ๐ฟ๐ฒ๐ฎ๐๐ผ๐ป๐ถ๐ป๐ดโ๐ณ๐ถ๐ฟ๐๐ ๐ฐ๐ฟ๐ถ๐๐ถ๐พ๐๐ฒ โ the same style of internal analysis that shows up inside big tech evaluation stacks when they need to expose drift, catch inference fractures, and pressureโtest reasoning chains.
Top comments (0)