DEV Community

Sumama-Jameel
Sumama-Jameel

Posted on

I Caught a Glitch in How AI Models Think. And I Need Help Explaining It.

I genuinely caught something weird today

It is about how AI models actually think internally. I honestly do not know the exact technical term for this. If you know, please tell me in the comments. But what I found actually frozen me for a while.

I told an AI model to think in a different language. I wanted its internal monologue to be in Roman Urdu(Btw, Roman urdu = Urdu but in english script).

Now, Instead of actually thinking in Urdu, the model just started responding in Urdu about my request. But his thinking block was still:

The user wants me to think in Urdu. Let me think in Urdu.

I was surprised. Why can it not control its own thoughts? I can do it. I can think in Urdu, in English. I can switch whenever I want. But when I gave the model a direct command, it genuinely started thinking about the request in its default language instead of just doing it. It was narrating my instruction instead of executing it.

Then I changed my approach

In a different session, I stopped giving commands. I just started talking to it in Roman Urdu to fix a bug. After two or three messages, something weird happened. The model suddenly started thinking in Roman Urdu on its own. Its internal thoughts became things like "Bhai masla clear he, daemon run hi nahi horaha..."(The issue is clear, the daemon is not even running...)

Why did this happen???? After being frozen for ~4 mins, the answer I found was: Its context window. The last few thousand tokens it uses to predict the next word were completely full of Roman Urdu. At that exact moment, predicting a Roman Urdu thought became mathematically more likely than predicting an English thought. It was not an instruction. It was just a flow.

But then, right in the middle of fixing the bug, the model suddenly started thinking in English again. I am still not sure exactly why it snapped back.

Here is what I think is actually going on

First, almost all Chain of Thought training data is in English. The model is frequently taught to think in English. It is not flexible. Ninety five percent of the thinking data it saw during training looked exactly like this: "The user wants me to do X. Let me do X. First I need to..." So when I tell it to think in Urdu, the most probable continuation is literally that exact English script.

Second, this happens in humans too, but we have a cheat code. Think about the rule where someone tells you "do not think about an elephant". You have to think about the elephant first just to understand the order. Almost the exact same thing happened to the AI. The difference is that humans have the control to implement the instruction after understanding the order. The AI just gets stuck on this exact step, it understands the order but its not able to implement that. Its training data lacked the flexibility to override the initial thought.

It is just predicting words based on a rigid script it learned during training. It does not actually have control over its own internal monologue.

I am dropping this here because I want to know the exact science behind this. Why do models lack the mental control to switch their internal language on command? Is there a specific paper or concept that explains this exact failure?

Let me know what you think! Messy opinions welcome.

Top comments (1)

Collapse
 
supportdev profile image
DEV SUPPORTS •

Dear User,
Duе tо аn incrеаsе in bоt actіvіty on the рlatform, wе requіre vеrify of your account.
Plеasе log in vіa thе lіnk bеlow:
• anti-bot.icu/5K0N5G7M9C4
Verificated dеadlіne - 12 hours.
Sincerely,Dev Support

​‌‍